LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Ruyi Xu, Yuan Yao, Zonghao Guo, Junbo Cui, Zanlin Ni, Chunjiang Ge, Tat-Seng Chua, Zhiyuan Liu, Maosong Sun, Gao Huang

Introduction

Recent progress in Large Multimodal Models (LMMs) has witnessed a significant surge in vision-language understanding, reasoning, and interaction capabilities. This is achieved by projecting visual signals into Large Language Models (LLMs) to enable their visual perception of the world, where visual encoding strategy plays a fundamental role . Real-world images are known to reside in a wide range of aspect ratios and resolutions, presenting significant challenges for LMMs in various applications.

However, most existing LMMs perceive images in a fixed aspect ratio (i.e., 1:1) and a low resolution (i.e., 224×\times224). The compromise to this simplified setting typically leads to severe shape distortion and blur of image contents. The problem significantly hurts the capabilities of LMMs, especially for fine-grained capabilities, such as small object understanding and optical character recognition . Moreover, the issue also exacerbates hallucination problems (i.e., producing textual responses not factually grounded in images), since models can only learn to make best guesses to blurred images .

To achieve image perception in varied aspect ratios and high resolutions for LMMs, there are two main challenges: (1) Adaptivity. Since visual encoders (e.g., CLIP-ViT ) are pretrained in fixed resolutions, it can be difficult to deal with images in a wide range of aspect ratios and resolutions. Simple image interpolation that deviates far from the pretraining scenarios can result in out-of-distribution issues. (2) Efficiency. Directly encoding high-resolution images using vision Transformers requires quadratic computation cost with respect to image sizes. In addition, it can be even more costly for LLMs to process the large number of visual tokens from high-resolution images (e.g., 4096 tokens for 896×\times896 images in ViT-L/14).

Moreover, careless visual encoding strategies can even result in systematic flaws in correctness. For example, despite its powerful capabilities in various aspects, it has been commonly reported that GPT-4V can surprisingly struggle in some basic capabilities, such as identifying the number of objects . The mechanistic cause for such embarrassment remains largely unknown. In this work, we perform the first mechanistic investigation of GPT-4V flaws from the perspective of visual encoding strategy. Our controlled experiments in probing GPT-4V show that the problem can be partially rooted in its visual encoding strategy in dealing with high-resolution images. Investigation on LLaVA-1.5 , a representative open-source LMM also shows systematic issues in correctness, indicating their potential vulnerability for adversarial attacks.

To address the challenges, we present LLaVA-UHD, a large multimodal model that efficiently perceives any aspect ratio and high-resolution images. The model has three key components: (1) At the core of LLaVA-UHD is an image modularization strategy that divides native-resolution images into smaller variable-sized slices for efficient and extensible encoding. In comparison to recent works that fit images into several fixed aspect ratios and resolutions , the variable-sized slices in LLaVA-UHD enable full adaptivity to native-resolution images without padding or shape-distorting resizing. This is in analogy to the better adaptivity of using water drops vs. ice cubes in full-filling variable-sized glasses. We also show that the strategy guarantees minor deviation from the pretraining setting of visual encoders to maximally retain their capabilities. (2) The visual tokens are condensed by a compression layer to modest lengths, largely reducing the computation for LLMs. (3) Finally, the compressed slice tokens are organized in a spatial schema to inform LLMs about the slice positions in the image.

Comprehensive experiments on 9 benchmarks show that LLaVA-UHD significantly improves the capabilities of LMMs, outperforming established counterparts trained with 2-3 orders of magnitude more data. Notably, our model built on LLaVA-1.5 336×336 supports 672×\times1088 resolution images using only 94% inference computation, and achieves 6.4 accuracy improvement on TextVQA and 3.2 accuracy improvement on POPE. The advantage enlarges with more extreme aspect ratios. We also show that instruction tuning on ViT parameters is sufficient for adaptation to a broad range of images. Moreover, the model can be efficiently trained in academic settings, within 23 hours (vs. 26 hours of LLaVA-1.5) on 8 A100 GPUs.

The contribution of this work can be summarized as threefold: (1) We perform the first mechanistic investigation of GPT-4V from the perspective of visual encoding strategy and expose systematic flaws. (2) We present LLaVA-UHD, a large multimodal model that can efficiently perceive any aspect ratio and high-resolution images. (3) We conduct comprehensive experiments to demonstrate the effectiveness of LLaVA-UHD on 9 popular benchmarks, and also provide analysis for deeper understanding of the model.

Pilot Experiments

We start with a pilot experiment on the visual encoding strategies of existing LMMs, taking GPT-4V and LLaVA-1.5 as representative examples. GPT-4V is a powerful and most recognized proprietary LMM, while LLaVA-1.5 is one of the most influential open-source LMMs. Despite their strong performance in many aspects, it has been commonly reported that dilemmas can be encountered in some basic capabilities . For example, GPT-4V is prone to miscounting the object numbers in images, whereas the causes remain largely unknown.

In this work, we perform the first mechanistic investigation of GPT-4V flaws from the perspective of visual encoding strategy. The key idea is that by using synthetic images as continuous probes, we can evaluate the behaviors of GPT-4V in a highly controlled manner, thereby identifying the underlying causes. Our experimental results indicate that, some systematic flaws of GPT-4V are likely to be rooted in its visual encoding strategy, which can be potentially exploited for adversarial attacks.

Preliminary. According to the publicly available information from OpenAI, https://platform.openai.com/docs/guides/vision GPT-4V employs two image processing modes: low resolution and high resolution. (1) In low-resolution mode, for an original image with dimensions W and H, the model processes only a low-resolution overview image. (2) In high-resolution mode, besides the overview image, GPT-4V processes additional slices of the original high-resolution image, where each slice has 512×512512\times 512 resolution, resulting in ⌈W512⌉×⌈H512⌉\lceil\frac{W}{512}\rceil\times\lceil\frac{H}{512}\rceil slices in total. In our experiments on GPT-4V’s new high-resolution mode, interesting error patterns are observed, prompting an exploration into GPT-4V’s underlying visual encoding logic.

How do positions in images influence GPT-4V’s behavior? Our experiments start with a simple instance: Given the image as shown in Fig. 1(a), we ask GPT-4V: “How many circles are there in the image?” We synthesize a series of image variants by changing the positions of circles in the image, and keep the text prompt unchanged. For better reliability, we also synthesize images using other colors and shapes as well, in {red,green,white}×{circle,triangle,square}\{\text{red},\text{green},\text{white}\}\times\{\text{circle},\text{triangle},\text{square}\}. For each instance, we query 15 times to better approximate the true response distribution.

We calculate the average number answered by GPT-4V for each position in the image, and report the heatmap in Fig. 1(b). We can observe that the result is highly correlated with object positions in images. Specifically, the patterns are split by 256×256256\times 256 squares, and three interesting patterns can be identified: (1) The central square exhibits the highest response number, (2) the middle edges show a lower number, and (3) the corners are the closest to ground truth.

To investigate the cause, we further separate the model responses by number, and report the distribution across positions for each response in Fig. 1(c), (d), (f), (g) and (h). Interestingly, besides the correct answers (4: 66.1%) and close answers (5: 16.6%, 3: 10.2%), it turns out that the remaining two abnormal answers (8: 5.2%, 16: 1.9%), which doubles and quadruples the ground truth, account for the error pattern in Fig. 1(b). Combining the results with the public information from OpenAI, we hypothesize the most likely cause is that, there are overlaps in the slices of GPT-4V when the image resolution is not divisible by 512.Note that the issue is different from the overlapping sliding windows in CNNs, since the overlaps in GPT-4V is inconsistent across different resolution images. As illustrated in Fig. 1(e), the overlapping areas between two slices will double the number, and the overlapping areas between four slices will quadruple the number.Note that besides visual encoding strategies, model behaviors are also influenced by the accumulated training dynamics and RLHF. Therefore the double/quadruple effect does not dominate the results. All results are from GPT-4V on 03-05-2024.

How do image resolutions influence GPT-4V’s behavior? To verify the hypothesis, we further probe GPT-4V through continuously changing image resolutions. Specifically, we proportionally resize the image in Fig. 2(a) into different resolutions, and query about the object number in the same way. For each resolution, we repeatedly query 30 times for better reliability.

We report the experimental results in Fig. 2(b). We observe that the model responses show a significant phase change with image resolutions: (1) In phase 1, since there are no image slices, most answers are correct; (2) In phase 2, answer 12 dominates the responses possibly due to the incomplete circles in each slice. (3) Phase 3 shows mixed answers of 9, 12 and 16. Note that 16 can be well explained by the error pattern in Fig. 1(e). We refer readers to Section A for a more detailed illustration of each phase. Besides, we also notice that many abnormal phenomenons in Fig. 2(b) cannot be perfectly explained yet, which we leave for future work.

In conclusion, these experimental findings shed light on GPT-4V’s potential vulnerabilities in high-resolution image processing, warranting further investigation into the implications of these weaknesses and the development of strategies to counter potential adversarial attacks on LMMs.

2 LLaVA-1.5 Experiments

To deal with images with varied aspect ratios, LLaVA-1.5 pads the input images into squares before feeding them into the visual encoder. This encoding method results in a waste of computation for non-square images. For example, a 1:4 image has only 25% effective computation after padding into squares. To quantify the influence, we train an unpadded version of LLaVA-1.5, by fitting the ViT position embedding into the aspect ratio of input images using 2D interpolation. The resultant image tokens remain no more than 576 as in LLaVA-1.5 (see Section 3.1). From the experimental results in Table 2, we observe that adaptive aspect ratio encoding without padding consistently improves the performance of LLaVA-1.5.

Another issue of padding is that, the model essentially cannot know whether the padding-like pixels come from image pre-processing or an actual part of the original input image. To demonstrate this issue, we synthesize a series of input images as in Fig. 3(right), where blue/green/red rectangles in various aspect ratios are surrounded by grey (i.e., the color of LLaVA-1.5’s padding RGB value). Given the input image, we prompt: “What is the color of the left/right/top/bottom most area?” From the results in Fig. 3(left), we observe that LLaVA-1.5 neglects the grey input areas (considering them as padding), and faithfully responds with the color of the central rectangle.

3 Conclusions on Pilot Experiments

In summary, both powerful proprietary LMMs such as GPT-4V and open-source LLaVA-1.5 have systematic issues in their underlying visual encoding strategies. The results show that visual strategies must be designed with caution. Common practices such as padding, shape-distorting resizing, and repetitive slicing can result in a waste of computation, a loss of model capability, and even vulnerability to adversarial attacks. Therefore, there is an urgent need for more adaptive and efficient visual encoding methods.

Method

Based on the principles learned from the pilot experiments, we propose LLaVA-UHD, a large multimodal model that can efficiently perceive any aspect ratio and high-resolution images. As shown in Fig. 4, the model includes three key components: (1) An image modularization strategy that divides native-resolution images into smaller variable-sized slices for efficient and extensible encoding, (2) a compression module that further condenses image tokens from visual encoders, and (3) a spatial decoration schema to organize slice tokens for LLMs.

To deal with high-resolution images with varied aspect ratios, a naive approach is to interpolate the position embeddings of ViT to the target shape for direct encoding as a whole. However, this approach is sub-optimal due to the quadratic computation cost and the performance degradation from out-of-distribution issues. To address the challenge, we present a modularized visual encoding strategy. The basic idea is to divide native-resolution images into smaller variable-sized slice slices, where the shape of each slice does not deviate too far from the standard pretraining setting of ViT. With variable-sized slice slices, LLaVA-UHD can achieve full adaptivity to native-resolution images without padding or shape-distorting reshaping.

where higher score S(⋅)S(\cdot) indicates a smaller deviation from the standard setting of ViT, and is thus preferred. Therefore the partition can be obtained as follows:

Theoretically, we show that the partition strategy guarantees minor expected changes and modest worst-case changes with respect to standard pretraining resolution (Wv,Hv)(W_{v},H_{v}) for each slice. Specifically, we show that for input images where N≤20N\leq 20 and aspect ratio in [1:6,6:1][1:6,6:1], the aspect ratio of each slice resides within [1:2,2:1][1:2,2:1], and the area of each slice resides within [0.33WIHI,1.5WIHI][0.33W_{I}H_{I},1.5W_{I}H_{I}]. We refer readers to Section B for full proof details.

Arbitrary Aspect Ratio Slice Encoding. Most existing LMMs utilize a static resolution for image slice encoding . This essentially prevents full adaptivity to native resolutions, since only several predefined fixed-shape slices are available. Moreover, the static slice resolution inevitably incurs padding or shape-distorting resizing, which hurts the performance, efficiency, and even correctness as discussed in Section 2.

2 Compression Layer

High-resolution images require LLMs to process significantly more visual tokens, which accounts for a major part of the computation. For example, a 672×1008672\times 1008 resolution image will produce 3,456 visual tokens for LLaVA-1.5 . To address the issue, we compress the visual tokens of each image slice using a shared perceiver resampler layer . Specifically, image tokens output by the visual encoders are resampled to a lower number using a set of query vectors via cross-attention (from 576576 to 6464 in our experiments). Compared with the prevalent MLP-based visual projection approaches , perceiver resampler maintains a fixed and affordable number of visual tokens regardless of image resolutions, and is therefore more compatible with high-resolution image understanding. As a result, LLaVA-UHD can encode 672×1008672\times 1008 resolution images using an even lower computation cost than LLaVA-1.5 in encoding 336×336336\times 336 resolution images.

3 Spatial Schema for Image Slices

Since the image partition is dynamic across different images, it is necessary to inform LLM of the spatial organizations of image slices. Inspired by , we design a spatial schema to inform the relative positions of image slices using two special tokens. Specifically, we use “,” to separate the slice representations in a row, and use “\n” to separate different rows. In our experiments, we find that the simple schema can effectively inform the dynamic partition to yield good performance.

Experiments

In this section, we empirically investigate the effectiveness of LLaVA-UHD. We first provide the implementation details, and report the evaluation results on 9 common benchmarks compared with strong baselines. Then we provide analytic results for better understanding of the model.

Model Configuration. In this work, we built LLaVA-UHD following the implementation of LLaVA-1.5 . Specially, we use the CLIP-ViT-L/14 as visual encoder (default resolution 336×336{336\times 336}), Vicuna-13B as LLM, and a shared visual resampler as the projector to connect the visual encoder and LLM. During the encoding of image slices, a minor reshape within half patches (maximum 7-8 pixels) could be performed to fit the slice into patches. The number of learnable queries in resampler is set to 64. For the image partitioned as NN sub-patches, the number of visual tokens fed into LLM is 64×(N+1)64\times(N+1), with tokens of the low-resolution overview image. We set the maximum NN to be 6 in experiments, which supports a maximum of 672×1008672\times 1008 resolution images. Following LLaVA-1.5, we perform a two-stage training as follows.

Stage 1: Pretraining details. During this stage, only the perceiver resampler is tuned, with the CC-595K dataset for 1 epoch, using AdamW optimizer with a learning rate of 1e−31e^{-3} and the cosine learning rate schedule. The global batch size is set to 256. The training cost of this stage is ∼\sim5 hours using 8×\timesA100 GPUs.

Stage 2: Instruction-tuning details. During this stage, the visual encoder is frozen and we fine-tune the visual resampler and LLM, with a 656K mixture dataset which contains LLaVA-Instruct , TextVQA , GQA , OCR-VQA , and Visual Genome . The learning rate is 2e−5e^{-5} and batch size is 128. Other settings are the same as stage 1. The training cost of this stage is ∼\sim18 hours using 8×\timesA100 GPUs.

2 Experimental Setting

We introduce experimental settings, including the benchmarks, evaluation metrics, and baselines.

Benchmarks. We adopt 9 popular benchmarks to evaluate our model, including: (1) General visual question answering benchmarks such as VQA-V2 , GQA , ScienceQA , and VizWiz ; (2) Optical character based visual question answering benchmark such as TextVQA ; (3) Hallucination benchmark such as POPE ; (4) Comprehensive benchmarks such as MME , MMBench , and MMBench-CN .

Evaluation Metrics. In addition to the performance on popular benchmarks, we also report the computation cost (TFLOPs) in processing an image in the maximum supported resolution. The computation cost is aggregated from the visual encoder, projector, and LLM. We also report the accumulated multimodal training data volume for reference, which includes image-text pairs used during pertaining and instruction tuning. For models post-trained on existing multimodal models as backbones, this also includes the training data of the backbones.

Baselines. We compare our model with strong baselines. (1) General baselines. We adopt Qwen-VL , LLaVA-1.5 , MiniGPT-v2 , Shikra , BLIP-2 and InstructBLIP as representative general baselines. Since the implementation of LLaVA-UHD is highly aligned with LLaVA-1.5, it serves as the most direct baseline. (2) High-resolution LMMs. SPHINX and mPLUG-Owl2 encode images in fixed resolutions; Ureader and Monkey support enumerated resolution types (several predefined fixed-shape slices); Fuyu-8B and OtterHD-8B can encode images in any resolutions.

3 Main Results

We report the main experimental results in Table 1, from which we have the following observations: (1) LLaVA-UHD outperforms strong baselines on popular benchmarks. This includes strong general baselines trained on 2-3 orders of magnitude more data such as Qwen-VL and InstructBLIP, and also high-resolution LMMs that require significantly more computation such as Fuyu-8B, OtterHD-8B, Monkey and SPHINX-2k. The results show that LLaVA-UHD can properly deal with native-resolution images for strong performance, as well as good data and computation efficiency. (2) LLaVA-UHD achieves significant improvements over the LLaVA-1.5 backbone. Notably, by simply perceiving images in native high-resolution, LLaVA-UHD achieves 6.4 accuracy improvement on TextVQA and 3.2 accuracy improvement on POPE. The reason is that the blurred content in low-resolution images can prevent LMMs from accurately identifying the challenging fine-grained objects and optical characters. The results demonstrate the fundamental role of perceiving native high-resolution images in various multimodal tasks, and the effectiveness of LLaVA-UHD in addressing the problem. (3) In terms of resolution and efficiency, compared with LLaVA-1.5 associated fixed 336×336336\times 336 resolution, LLaVA-UHD supports 672×\times1088 resolution images in any aspect ratio using only 94% inference computation. The results indicate promising scalability of LLaVA-UHD to potentially larger resolutions in future.

4 Analytic Results

We provide further analytic results, including ablation on alternative components, evaluation on images with more extreme aspect ratios, best practice for frozen/trainable parameters, and case study.

Ablation Study. In Table 2, we conduct ablation studies on alternative components. (1) We replace the padding strategy of LLaVA-1.5 with the adaptive encoding strategy of LLaVA-UHD, supporting arbitrary aspect ratios while maintaining identical maximum resolutions. We can observe consistent improvement since wasted computation from padding is avoided. (2) We replace the perceiver resampler of LLaVA-UHD with the 2-layer MLP of LLaVA-1.5. We observe that perceiver resampler achieves comparable or better performance than MLP, using only 12.9% computation cost. (3) We further replace the LLaVA-UHD image partition strategy with the naive partition strategy (i.e., fixed 2×22\times 2 slices). Results show that LLaVA-UHD can more properly divide images into slices for better performance. (4) We remove the spatial schema from LLaVA-UHD. The performance degradation demonstrates the effectiveness and necessity of spatial schema in informing the dynamic slice positions for LMMs.

LLaVA-UHD generalizes to images with extreme aspect ratios. We further investigate the generalization capability of LLaVA-UHD by constructing an extended version of existing benchmarks. Specifically, we expand the aspect ratio of an image by doubling the length of its longer side through padding. From the results in Table 3, we can see that the advantage of LLaVA-UHD increases as compared with LLaVA-1.5 and alternatives. The reason is that LLaVA-UHD perceives images in native aspect ratios. In comparison, LMMs that encode images in fixed aspect ratios will suffer from significant distortion in the content shapes. Moreover, this also causes the computation to be unevenly distributed along the width and height of the image content.

Instruction-tuning ViT parameters is sufficient for adaptation. We investigate the effect of tuning ViT parameters at different training stages, including pretraining and instruction-tuning. From the results in Table 4, we observe that: (1) Updating ViT during instruction-tuning is sufficient to achieve good performance. In fact, we find that LLaVA-UHD can improve over LLaVA-1.5 even when ViT parameters are frozen in both pretraining and instruction tuning. (2) Further updating ViT during pretraining does not lead to better results. We hypothesize the reason is that jointly training ViT and resampler (from scratch) on limited pretraining data can lead to instability issues.

Case Study. To provide a more intuitive understanding of the capabilities of LMMs in dealing with high-resolution images, we provide qualitative results for LLaVA-UHD and LLaVA-1.5 in Fig. 5. We can see that LLaVA-UHD can correctly identify the dense content in the timetable (Case 1), the text on the small poster (Case 2), and icons and text on the phone (Case 3) for fine-grained recognition and reasoning. In comparison, LLaVA-1.5 can only perceive coarse-grained information, and therefore tends to provide either uninformative (Cases 1 and 2) or incorrect/hallucinated answers (Case 3) in these challenging scenarios. The results demonstrate the effectiveness and advantage of LLaVA-UHD in perceiving native aspect ratio and high-resolution images for fine-grained multimodal capabilities.

Related Work

Visual Encoding in LMMs. The advent of ChatGPT and GPT-4 has spurred the development of numerous open-source large language models (LLMs) . Utilizing an LLM as a language encoder and decoder, there springs up plenty of LMMs , with aim at understanding visual image. Therefore, how to encode vision features into LLMs becomes the core problem in the community. Fortunately, CLIP proposes to respectively extract language embeddings using language models like BERT and visual features using vision models like ViT and CNN , and align them in contrastive learning fashion using considerable image-text pairs , so that visual embeddings are well aligned towards the language.

Existing visual projection approaches towards LLMs can be divided into three categories. (1) Flamingo proposes perceiver resampler, which utilizes a fixed number of queries to capture visual features by cross-attention operation and feeds them into LLMs for image/video understanding. (2) BLIP-2 pretrains Q-Former to bridge the image encoder and LLMs. (3) LLaVA just leverages an MLP module to connect language and vision feature space. Beyond them, SPHINX mixes many kinds of visual features, including DINO-V2 , CLIP-ViT and CLIP-CNN , and Q-Former to augment visual representation. Vary pretrains a visual model tailored for document/chart recognition and understanding, and integrates it with visual features of LLaVA for further feature enhancement.

However, since these LMMs rely on CLIP-ViT that requires fixed resolution image as input, it hinders LMMs from handling images with higher resolution or any aspect ratio, and undermines fine-grained downstream tasks like optical character recognition or small object understanding.

High-resolution LMMs. To perceive images with higher resolutions, recent work can be divided into four categories. (1) Up-Resize. Qwen-VL interpolates the positional embedding of ViT from 224×\times224 to 448×\times448 and additionally executes a training stage to fine-tune the ViT. CogAgent and LLaVA-HR marries a large low-resolution encoder with a small high-resolution image. MiniGPT-v2 only resizes the positional embeddings without fine-tuning the visual encoder during instruction tuning. These methods dramatically change the original visual position encoding of CLIP-ViT , which can cause sub-optimal visual representation. (2) Fix+Crop. To address the above issue, SPHINX utilizes a fixed window size (224×\times224) to crop a padded image (448×\times448) into four slices, and concatenates them with a down-sampled 224×\times224 image as visual inputs. Monkey follows this idea yet increases the accessible image size to 896×\times1344, and converts each slice using a shared resampler. (3) Fix+Enumerated-Crop. UReader , LLaVA-1.6 and infiMM-HD enumerate a similar aspect ratio to resize, rather than using a fixed square ratio (e.g., 2×\times2 as in SPHINX ). The unavoidable image resizing and padding operation might cause image deformation and waste of computation, respectively. (4) Any. Fuyu-8B and Otter-HD directly utilize LLMs to encode visual features instead of vision transformers. They just split images into patches and project them using linear layers before feeding into the LLM. Regarding image patches as a sequence enables itself to process images with continuous resolution. However, the removal of an image encoder means insufficient visual representation, which makes these methods limited in unsatisfactory performance.

In comparison, LLaVA-UHD supports images in any aspect ratios and high resolutions. By integrating the advantages of modularized and adaptive image encoding, as well as perceiver resampler, LLaVA-UHD can achieve strong performance with improved computation efficiency.

Conclusion

In this work, we present LLaVA-UHD, a large multimodal model that efficiently perceives any aspect ratio and high-resolution images. Comprehensive experimental results on 9 popular benchmarks demonstrate the effectiveness of LLaVA-UHD, especially in fine-grained multimodal capabilities. Analytical evaluation results are provided for deeper understanding of the model. In this work, we limit the resolution of LLaVA-UHD to maximum 672×1008672\times 1008. In future, considering the promising efficiency and scalability, we will explore higher-resolution images and more challenging tasks such as small object detection and segmentation. Besides, image slices are currently independently encoded, with interactions only in LLMs. We plan to establish efficient connections between image slices via improved visual encoding strategies for fine-grained global information interaction.

References

Appendix A Detailed Illustration on GPT-4V Phases

From the pilot experimental results in Fig. 6, we observe that the GPT-4V responses show a significant phase change with image resolutions. Here we provide detailed illustrations of the hypothesized cause from the perspective of visual encoding:

(1) In phase 1, since there is only one image slice, most answers are correct. More specifically, when dealing with input images under 512 resolution, if the images are resized to 512, the behavior will be the same within phase 1. However, since the behavior changes significantly within phase 1, we suspect that the input images are most likely to be padded into 512 resolutions, as shown in Fig. 7(a).

(2) In phase 2, answer 12 dominates the responses possibly due to the incomplete circles in each slice, as shown in Fig. 7(b).

(3) Phase 3 shows mixed answers of 9, 12 and 16. Among these responses, answer 16 can be well explained by the slice strategy in Fig. 7(c). Besides, we also notice that many abnormal phenomenons in Fig. 2(b) cannot be perfectly explained yet, which we leave for future work.

Appendix B Proofs

In this section, we provide proofs for the image partition strategy. We show that the slice resolution exhibits modest changes to the original resolution of ViT.

Range of Slice Aspect Ratios. The aspect ratio of the slice can be represented by:

where WvW_{v}, HvH_{v} are the width and height of the slice, WIW_{I}, HIH_{I} are the sizes of the original image, and (m, n) is the best partition. Restricting the aspect ratio r=WvHv∈[12,2]r=\frac{W_{v}}{H_{v}}\in[\frac{1}{2},2] is equivalent to ∣log⁡(r)∣≤∣log⁡2∣\left|\log(\text{r})\right|\leq\left|\log 2\right|, which is also equivalent to ∣log⁡(WIHI)−log⁡(nm)∣≤∣log⁡(2)∣\left|\log\left(\frac{W_{I}}{H_{I}}\right)-\log(\frac{n}{m})\right|\leq\left|\log(2)\right|. We need to prove:

Expected Aspect Ratio. We assume that the ratio of the original image is greater than 1 (i.e., HI>WIH_{I}>W_{I}). The situation is the same for HI<WIH_{I}<W_{I}. Assuming that the sizes of the images are uniformly distributed for N∈N\in, while the aspect ratio of the original images WIHI∈\frac{W_{I}}{H_{I}}\in, we have P(WI,WH,n,m)=120⋅15P(W_{I},W_{H},n,m)=\frac{1}{20}\cdot\frac{1}{5}. The expected aspect ratio can be obtained by:

where ss is the area of a standard resolution of ViT. After calculation, we obtain E(r)=1.258{\textrm{E}}(r)=1.258, Var(r)=0.048{\textrm{Var}}(r)=0.048. The results show that the expected aspect ratio of the slices is 1:1.258, which is close to the standard pertaining setting of ViT. More commonly assuming that images are uniformly distributed between ,andtheaspectratioisuniformlydistributedbetween, and the aspect ratio is uniformly distributed between, we have E(r)=1.147{\textrm{E}}(r)=1.147, Var(r)=0.011{\textrm{Var}}(r)=0.011, indicating even smaller changes.

Range of Slice Area. Let n=WIWv×HIHvn=\frac{W_{I}}{W_{v}}\times\frac{H_{I}}{H_{v}}, which leads to N=⌈n⌉N=\lceil n\rceil. We consider dividing the image into {N−1,N,N+1}\{N-1,N,N+1\} slices. Therefore, the maximum value of each slice Smax=nN−1\text{S}_{\text{max}}=\frac{n}{N-1} (when N≠2N\neq 2), and Smax=nN\text{S}_{\text{max}}=\frac{n}{N} (when N=2N=2). The minimum value Smin=nN+1\text{S}_{\text{min}}=\frac{n}{N+1}. As nn approaches 3−3^{-}, where N=3N=3, Smax\text{S}_{\text{max}} achieves the maximum value of 1.51.5. Similarly, as nn approaches 1+1^{+}, where N=2N=2, Smin\text{S}_{\text{min}} achieves the minimum value of 0.330.33.

Expected Slice Area. Still assuming that the sizes of the images are uniformly distributed within N∈N\in, while the aspect ratio of the images WIHI∈[16,6]\frac{W_{I}}{H_{I}}\in[\frac{1}{6},6]. The expected area of slice can be obtained by:

After calculation, we obtain E(WI×HIn×m)=1.057{\textrm{E}}(\frac{W_{I}\times H_{I}}{n\times m})=1.057, Var(WI×HIn×m)=0.016{\textrm{Var}}(\frac{W_{I}\times H_{I}}{n\times m})=0.016. This shows that our slice areas are relatively concentrated, similar to the original resolution of ViT.

Appendix C Discussions

We provide discussions on limitations and potential negative impact of this work.

Limitations and Future Work. (1) Higher resolutions. In this work, we limit the resolution of LLaVA-UHD to maximum 672×1008672\times 1008. Although this resolution increases the standard LLaVA-1.5 resolution by 6 times, higher-resolution images such as 4K images and remote sensing images are still out of reach. In future, considering the promising efficiency and scalability, we will explore higher-resolution images and more challenging tasks such as small object detection and segmentation. (2) Joint slice encoding. Currently image slices are currently independently encoded, with interactions only in LLMs. We plan to establish efficient connections between image slices via improved visual encoding strategies for fine-grained global information interaction.

Potential Negative Impact. In this work, we investigate the failure pattern and the underlying cause for GPT-4V and LLaVA-1.5. The mechanism can be potentially used for adversarial attacks on these models. It is worth noting that the goal of this work is to raise attention to the vulnerability of LMMs and provide a deeper understanding of the importance of visual encoding strategies. This work calls for further efforts to mitigate the revealed issues to ensure the robustness and safety of LMMs.