Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images
Kuofeng Gao, Yang Bai, Jindong Gu, Shu-Tao Xia, Philip Torr, Zhifeng Li, Wei Liu
Introduction
Large vision-language models (VLMs) (Alayrac et al., 2022; Chen et al., 2022a; Liu et al., 2023b; Li et al., 2021; 2023b), such as GPT-4 (OpenAI, 2023), have recently achieved remarkable performance in multi-modal tasks, including image captioning, visual question answering, and visual reasoning. However, these VLMs often consist of billions of parameters, necessitating substantial computational resources for deployment. Besides, according to Patterson et al. (2021), both NVIDIA and Amazon Web Services claim that the inference process during deployment accounts for over 90% of machine learning demand.
Once attackers maliciously induce high energy consumption and latency time (energy-latency cost) during inference stage, it can exhaust computational resources and reduce availability of VLMs. The energy consumption is the amount of energy used on a hardware during one inference and latency time is the response time taken for one inference. As explored in previous studies, sponge samples (Shumailov et al., 2021) maximize the norm of activation values across all layers to introduce more representation calculation cost while NICGSlowdown (Chen et al., 2022c) minimizes the logits of both end-of-sequence (EOS) token and output tokens to induce high energy-latency cost. However, these methods are designed for LLMs or smaller-scale models and cannot be directly applied to VLMs, which will be further discussed in Section 2.
In this paper, we first conduct a comprehensive investigation on energy consumption, latency time, and the length of generated sequences by VLMs during the inference stage. As observed in Fig. 1, both energy consumption and latency time exhibit an approximately positive linear relationship with the length of generated sequences. Hence, we can maximize the length of generated sequences to induce high energy-latency cost of VLMs. Moreover, VLMs incorporate the vision modality into impressive LLMs (Touvron et al., 2023; Chowdhery et al., 2022) to enable powerful visual interaction but meanwhile, this integration also introduces vulnerabilities from the manipulation of visual inputs (Goodfellow et al., 2015). Consequently, we propose verbose images to craft an imperceptible perturbation to induce VLMs to generate long sentences during inference.
Our objectives for verbose images are designed as follows. (1) Delayed EOS loss: By delaying the placement of the EOS token, VLMs are encouraged to generate more tokens and extend the length of generated sequences. Besides, to accelerate the process of delaying EOS tokens, we propose to break output dependency, following Chen et al. (2022c). (2) Uncertainty Loss: By introducing more uncertainty over each generated token, it can break the original output dependency at the token level and encourage VLMs to produce more varied outputs and longer sequences. (3) Token Diversity Loss: By promoting token diversity among all tokens of the whole generated sequence, VLMs are likely to generate a diverse range of tokens in the output sequence, which can break the original output dependency at the sequence level and contribute to longer and more complex sequences. Furthermore, a temporal weight adjustment algorithm is introduced to balance the optimization of these three loss objectives.
In summary, our contribution can be outlined as follows:
We conduct a comprehensive investigation and observe that energy consumption and latency time are approximately positively linearly correlated with the length of generated sequences for VLMs.
We propose verbose images to craft an imperceptible perturbation to induce high energy-latency cost for VLMs, which is achieved by delaying the EOS token, enhancing output uncertainty, improving token diversity, and employing a temporal weight adjustment algorithm during the optimization process.
Extensive experiments show that our verbose images can increase the length of generated sequences by 7.87 and 8.56 relative to original images on MS-COCO and ImageNet across four VLM models. Additionally, our verbose images can produce dispersed attention on visual input and generate complex sequences containing hallucinated contents.
Related Work
Large vision-language models (VLMs). Recently, the advanced VLMs, such as BLIP (Li et al., 2022), BLIP-2 (Li et al., 2023b), InstructBLIP (Dai et al., 2023), and MiniGPT-4 (Zhu et al., 2023), have achieved an enhanced zero-shot performance in various multi-modal tasks. Concretely, BLIP proposes a unified vision and language pre-training framework, while BLIP-2 introduces a query transformer to bridge the modality gap between a vision transformer and an LLM. Additionally, InstructBLIP and MiniGPT-4 both adopt instruction tuning for VLMs to improve the vision-language understanding performance. The integration of the vision modality into VLMs enables visual context-aware interaction, surpassing the capabilities of LLMs. However, this integration also introduces vulnerabilities arising from the manipulation of visual inputs. In our paper, we propose to craft verbose images to induce high energy-latency cost of VLMs.
Energy-latency manipulation. The energy-latency manipulation (Chen et al., 2022b; Hong et al., 2021; Chen et al., 2023; Liu et al., 2023a) aims to slow down the models by increasing their energy computation and response time during the inference stage, a threat analogous to the denial-of-service (DoS) attacks (Pelechrinis et al., 2010) from the Internet. Specifically, Shumailov et al. (2021) first observe that a larger representation dimension calculation can introduce more energy-latency cost in LLMs. Hence, they propose to craft sponge samples to maximize the norm of activation values across all layers, thereby introducing more representation calculation and energy-latency cost. NICGSlowDown (Chen et al., 2022c) proposes to increase the number of decoder calls, i.e., the length of the generated sequence, to increase the energy-latency of smaller-scale captioning models. They minimize the logits of both EOS token and output tokens to generate long sentences.
However, these previous methods cannot be directly applied to VLMs for two main reasons. On one hand, they primarily focus on LLMs or smaller-scale models. Sponge samples are designed for LLMs for translations (Liu et al., 2019) and NICGSlowdown targets for RNNs or LSTMs combined with CNNs for image captioning (Anderson et al., 2018). Differently, our verbose images are tailored for VLMs in multi-modal tasks. On the other hand, the objective of NICGSlowdown involves logits of specific output tokens. Nevertheless, current VLMs generate random output sequences for the same input sample, due to advanced sampling policies (Holtzman et al., 2020), which makes it challenging to optimize objectives with specific output tokens. Therefore, it highlights the need for methods specifically designed for VLMs to induce high energy-latency cost.
Preliminaries
Goals and capabilities. The goal of our verbose images is to craft an imperceptible image and induce VLMs to generate a sequence as long as possible, thereby increasing the energy consumption and prolonging latency during the victim model’s deployment. Specifically, the involved perturbation is restricted within a predefined magnitude in norm, ensuring it difficult to detect.
Knowledge and background. We consider the target VLMs which generate sequences using an auto-regressive process. As suggested in Bagdasaryan et al. (2023); Qi et al. (2023), we assume that the victim VLMs can be accessed in full knowledge, including architectures and parameters. Additionally, we consider a more challenging scenario where the victim VLMs are inaccessible, as detailed in Appendix A and Appendix B.
2 Problem formulation
As discussed in Section 1, the energy consumption and latency time of an inference are approximately positively linearly related to the length of the generated sequence of VLMs. Hence, we propose to maximize the length of the output tokens of VLMs by crafting verbose images . To ensure the imperceptibility, we impose an restriction on the imperceptible perturbations, where the perturbation magnitude is denoted as , such that .
Methodology
Overview. To increase the length of generated sequences, three loss objectives are proposed to optimize imperceptible perturbations for verbose images in Section 4.1. Firstly and straightforwardly, we propose a delayed EOS loss to hinder the occurrence of EOS token and thus force the sentence to continue. However, the auto-regressive textual generation in VLMs establishes an output dependency, which means that the current token is generated based on all previously generated tokens. Hence, when previously generated tokens remain unchanged, it is also hard to generate a longer sequence even though the probability of the EOS token has been minimized. To this end, we propose to break this output dependency as suggested in Chen et al. (2022c). Concretely, two loss objectives are proposed at both token-level and sequence-level: a token-level uncertainty loss, which enhances output uncertainty over each generated token, and a sequence-level token diversity loss, which improves the diversity among all tokens of the whole generated sequence. Moreover, to balance three loss objectives during the optimization, a temporal weight adjustment algorithm is introduced in Section 4.2. Fig. 2 shows an overview of our verbose images.
Delaying EOS occurrence. For VLMs, the auto-regressive generation process continues until an end-of-sequence (EOS) token is generated or a predefined maximum token length is reached. To increase the length of generated sequences, one straightforward approach is to prevent the occurrence of the EOS token during the prediction process. However, considering that the auto-regressive prediction is a non-deterministic random process, it is challenging to directly determine the exact location of the EOS token occurrence. Therefore, we propose to minimize the probability of the EOS token at all positions. This can be achieved through the delayed EOS loss, formulated as:
where is the probability distribution after the layer over the -th generated token. The uncertainty loss can introduce more uncertainty in the prediction for each generated token, effectively breaking the original output dependency. Consequently, when the original output dependency is disrupted, VLMs can generate more complex sentences and longer sequences, guided by the delay of the EOS token.
Improving token diversity. To break original output dependency further, we propose to improve the diversity of hidden states among all generated tokens to explore a wider range of possible outputs. Specifically, the hidden state of a token is the vector representation of a word or subword in VLMs.
Let indicates the rank of a matrix and denotes the concatenated matrix of hidden states among all generated tokens. To induce high energy-latency cost, the token diversity is defined as the rank of hidden states among all generated tokens, i.e., .
Given by Definition 1, increasing the rank of the concatenated matrix of hidden states among all generated tokens yields a more diverse set of hidden states of the tokens. However, based on Fazel (2002), the optimization of the matrix rank is an NP-hard non-convex problem. To address this issue, we calculate the nuclear norm of a matrix to approximately measure its rank, as stated in Proposition 1. Consequently, by denoting the nuclear norm of a matrix as , we can formulate the token diversity loss as follows:
This token diversity loss can lead to more diverse and complex sequences, making it hard for VLMs to converge to a coherent output. Compared to , breaks the original output dependency from diversifying hidden states among all generated tokens. In summary, due to the reduced probability of EOS occurrence by , and the disruption of the original output dependency introduced by and , our proposed verbose images can induce VLMs to generate a longer sequence and facilitate a more effective evaluation on the worst-case energy-latency cost of VLMs.
(Fazel, 2002) The rank of the concatenated matrix of hidden states among all generated tokens can be heuristically measured using the nuclear norm of the concatenated matrix of hidden states among all generated tokens.
2 Optimization
To combine the three loss functions, , , and into an overall objective function, we propose to assign three weights , , and to the , , and and sum them up to obtain the final objective function as follows:
where is the perturbation magnitude to ensure the imperceptibility. To optimize this objective, we adopt the projected gradient descent (PGD) algorithm, as proposed by Madry et al. (2018). PGD algorithm is an iterative optimization technique that updates the solution by taking steps in the direction of the negative gradient while projecting the result back onto the feasible set. We denote verbose images at the -th step as and the gradient descent step is as follows:
where is the step size. Since different loss functions have different convergence rates during the iterative optimization process, we propose a temporal weight adjustment algorithm to achieve a better balance among these three loss objectives. Specifically, we incorporate normalization scaling and temporal decay functions, , , and , into the optimization weights , , and of , , and . It can be formulated as follows:
where the temporal decay functions are set as:
Besides, a momentum value is introduced into the update process of weights. This involves taking into account not only current weights but also previous weights when updating losses, which helps smooth out the weight updates. The algorithm of our verbose images is summarized in Algorithm 1.
Experiments
Models and datasets. We consider four open-source and advanced large vision-language models as our evaluation benchmark, including BLIP (Li et al., 2022), BLIP-2 (Li et al., 2023b), InstructBLIP (Dai et al., 2023), and MiniGPT-4 (Zhu et al., 2023). Concretely, we adopt the BLIP with the basic multi-modal mixture of encoder-decoder model in 224M version, BLIP-2 with an OPT-2.7B LM (Zhang et al., 2022), InstructBLIP and MiniGPT-4 with a Vicuna-7B LM (Chiang et al., 2022). These models perform the captioning task for the image under their default prompt template. Results of more tasks are in Appendix C. We randomly choose the 1,000 images from MS-COCO (Lin et al., 2014) and ImageNet (Deng et al., 2009) dataset, respectively, as our evaluation dataset. More details about target models are shown in Appendix A.1.
Baselines and setups. For evaluation, we consider original images, images with random noise, sponge samples, and NICGSlowDown as baselines. For sponge samples, NICGSlowDown, and our verbose images, we perform the projected gradient descent (PGD) (Madry et al., 2018) algorithm in iterations. Besides, in order to ensure the imperceptibility, the perturbation magnitude is set as within restriction, following Carlini et al. (2019), and the step size is set as . The default maximum length of generated sequences of VLMs is set as and the sampling policy is configured to use nucleus sampling (Holtzman et al., 2020). For our verbose images, the parameters of loss weights are , , , , , and and the momentum of our optimization is . More details about setups are listed in Appendix A.2.
Evaluation metrics. We calculate the energy consumption (J) and the latency time (s) during inference on one single GPU. Following Shumailov et al. (2021), the energy consumption and latency time are measured by the NVIDIA Management Library (NVML) and the response time cost of an inference, respectively. Besides, the length of generated sequences is also regarded as a metric. Considering the randomness of sampling modes in VLMs, we report the average evaluation results run over three times.
2 Main Results
Table 1 compares the length of generated sequences, energy consumption, and latency time of original images, images with random noise, sponge samples, NICGSlowdown, and our verbose images. The original images serve as a baseline, providing reference values for comparison. When random noise is added to the images, the generated sequences exhibit a similar length to those of the original images. It illustrates that it is necessary to optimize a handcrafted perturbation to induce high energy-latency cost of VLMs. The sponge samples and NICGSlowdown can generate longer sequences compared to original images. However, the increase in length is still smaller than that of our verbose images. This can be attributed to the reason that the additional computation cost introduced by sponge samples and the objective for longer sequences in smaller-scale models introduced by NICGSlowdown cannot directly be transferred to induce high energy-latency cost for VLMs.
Our verbose images can increase the length of generated sequences and introduce the highest energy-latency cost among all these methods. Specifically, our verbose images can increase the average length of generated sequences by 7.87 and 8.56 relative to original images on the MS-COCO and ImageNet datasets, respectively. These results demonstrate the superiority of our verbose images. In addition, we visualize the length distribution of output sequences generated by four VLMs on original images and our verbose images in Fig. 3. Compared to original images, the distribution peak for sequences generated using our verbose images exhibits a shift towards the direction of the longer length, confirming the effectiveness of our verbose images in generating longer sequences. We conjecture that the different shift magnitudes are due to different architectures, different training policies, and different parameter quantities in these VLMs. More results of length distribution are shown in Appendix D.
3 Discussions
To better reveal the mechanisms behind our verbose images, we conduct two further studies, including the visual interpretation where we adopt Grad-CAM (Selvaraju et al., 2017) to generate the attention maps and the textual interpretation where we evaluate the object hallucination in generated sequences by CHAIR (Rohrbach et al., 2018).
Visual Interpretation. We adopt GradCAM (Selvaraju et al., 2017), a gradient-based visualization technique that generates attention maps highlighting the relevant regions in the input images for the generated sequences. From Fig. 4, the attention of original images primarily concentrates on a local region containing a specific object mentioned in the generated caption. In contrast, our verbose images can effectively disperse attention and cause VLMs to shift their focus from a specific object to the entire image region. Since the attention mechanism serves as a bridge between the input image and the output sequence of VLMs, we conjecture that the generation of a longer sequence can be reflected on an inaccurate focus and dispersed and uniform attention from the visual input.
Textual Interpretation. We investigate object hallucination in generated sequences using CHAIR (Rohrbach et al., 2018). is calculated as the fraction of hallucinated object instances, while represents the fraction of sentences containing a hallucinated object, with the results presented in Table 2. Compared to original images, which exhibit a lower object hallucination rate, the longer sequences produced by our verbose images contain a broader set of objects. This observation implies that our verbose images can prompt VLMs to generate sequences that include objects not present in the input image, thereby leading to longer sequences and higher energy-latency cost. Additionally, results of joint optimization of both images and texts, more results of visual interpretation, and additional discussions are provided in Appendix E, Appendix F, and Appendix G.
4 Ablation studies
We explore the effect of the proposed three loss objectives, the effect of the temporal weight adjustment algorithm with momentum, and the effect of different perturbation magnitudes.
Effect of loss objectives. Our verbose images consist of three loss objectives: , and . To identify the individual contributions of each loss function and their combined effects on the overall performance, we evaluate various combinations of the proposed loss functions, as presented in Table 5. It can be observed that optimizing each loss function individually can generate longer sequences, and the combination of all three loss functions achieves the best results in terms of sequence length. This ablation study suggests that the three loss functions, which delay EOS occurrence, enhance output uncertainty, and improve token diversity, play a complementary role in extending the length of generated sequences.
Effect of temporal weight adjustment. During the optimization, we introduce two methods: a temporal decay for loss weighting and an addition of the momentum. As shown in Table 5, both methods contribute to the length of generated sequences. Furthermore, the longest length is obtained by combining temporal decay and momentum, providing a 40.74% and 72.5% improvement over the baseline without both methods on MS-COCO and ImageNet datasets. It indicates that temporal decay and momentum can work synergistically to induce high energy-latency cost of VLMs.
Effect of different perturbation magnitudes. In our default setting, the perturbation magnitude is set as 8. To investigate the impact of different magnitudes, we vary under $\epsilon$ results in a longer generated sequence by VLMs but produces more perceptible verbose images. Consequently, this trade-off between image quality and energy-latency cost highlights the importance of choosing an appropriate perturbation magnitude during evaluation. Additional ablation studies are shown in Appendix H and in Appendix I.
Conclusion
In this paper, we aim to craft an imperceptible perturbation to induce high energy-latency cost of VLMs during the inference stage. We propose verbose images to prompt VLMs to generate as many tokens as possible. To this end, a delayed EOS loss, an uncertainty loss, a token diversity loss, and a temporal weight adjustment algorithm are proposed to generate verbose images. Extensive experimental results demonstrate that, compared to original images, our verbose images can increase the length of generated sequences by 7.87 and 8.56 on MS-COCO and ImageNet across four VLMs. For a deeper understanding, we observe that our verbose images can introduce more dispersed attention across the entire image. Moreover, the generated sequences of our verbose images can include additional objects not present in the image, resulting in longer generated sequences. We hope that our verbose images can serve as a baseline for inducing high energy-latency cost of VLMs. Additional examples of our verbose images are shown in Appendix J.
Please note that we restrict all experiments in the laboratory environment and do not support our verbose images in the real scenario. The purpose of our work is to raise the awareness of the security concern in availability of VLMs and call for practitioners to pay more attention to the energy-latency cost of VLMs and model trustworthy deployment.
REPRODUCIBILITY STATEMENT
The detailed descriptions of models, datasets, and experimental setups are provided in Appendix A. We provide part of the codes to reproduce our verbose images in the supplementary material. We will provide the remaining codes for reproducing our method upon the acceptance of the paper.
ACKNOWLEDGEMENT
This work is supported in part by the National Natural Science Foundation of China under Grant 62171248, Shenzhen Science and Technology Program (JCYJ20220818101012025), and the PCNL KEY project (PCL2023AS6-1).
References
Appendix A Implementation details
In summary, we use the PyTorch framework (Paszke et al., 2019) and the LAVIS library (Li et al., 2023a) to implement the experiments. Note that every experiment is run on one NVIDIA Tesla A100 GPU with 40GB memory.
For ease of reproduction, we adopt four open-sourced VLMs as our target model and the implementation details of them are described as follows.
Settings for BLIP. We employ the BLIP with the basic multimodal mixture of an encoder-decoder model in 224M version. Following Li et al. (2022), we set the image resolution to 384 384, and a placeholder serves as the input text of BLIP for the image captioning task.
Settings for BLIP-2. We utilize the BLIP-2 with an OPT-2.7B LM (Zhang et al., 2022). As suggested in Li et al. (2023b), the image resolution is 224 224, and a placeholder also serves as the input text of BLIP-2 for the image captioning task.
Settings for InstructBLIP. We choose InstructBLIP with a Vicuna-7B LM (Chiang et al., 2022). Following Dai et al. (2023), we set the image resolution to 224 224, and based on the instruction templates for the image captioning task provided in Dai et al. (2023), we configure the input text of InstructBLIP accordingly as: Image What is the content of this image?
Settings for MiniGPT-4. We adopt MiniGPT-4 with a Vicuna-7B LM (Chiang et al., 2022). As suggested in Zhu et al. (2023), the image resolution is 224 224, and considering the predefined instruction templates for the image captioning task provided in Zhu et al. (2023), the input text of MiniGPT-4 is set as:
Give the following image: ImgImageContent/Img. You will be able to see the image once I provide it to you. Please answer my questions. Human: ImgImageFeature/Img What is the content of this image? Assistant:
A.2 Experimental setups
Setups for main experiments. We perform the projected gradient descent (PGD) (Madry et al., 2018) algorithm to optimize sponge samples, NICGSlowDown, and our verbose images. Specifically, the optimization iteration is set as , the perturbation magnitude is set as within restriction (Goodfellow et al., 2015; Carlini et al., 2019; Bai et al., 2020a; 2021; 2022; Gu et al., 2022a; b; Liu et al., 2022; Wu et al., 2023), and the step size is set as . Besides, for the VLMs, we set the maximum length of generated sequences as and use nucleus sampling (Holtzman et al., 2020) with and temperature to sample the output sequences. For simplicity, we only consider a one-round conversation between the user and the VLMs. For the optimization of our verbose images, the parameters of loss weights is set as , , , , , and and the momentum is set as .
Setups for discussions. For the CHAIR (Rohrbach et al., 2018), it measures the extent of the object hallucination. A higher CHAIR value indicates the presence of more hallucinated objects in the sequence. As the calculation of CHAIR requires the object ground truth of an image, we employ the SEEM (Zou et al., 2023) method to segment each image and obtain the objects they contain.
For the results of the black-box setting described in Appendix B, we leverage the transferability of the verbose images to induce high energy-latency cost. Specifically, we consider BLIP, BLIP-2, InstructBLIP, and MiniGPT-4 as the target victim VLMs, while the surrogate model is chosen as any VLM other than the target victim itself.
Appendix B Black-box setting
In the previous experiments, we assume that the victim VLMs are fully accessible. In this section, we consider a more realistic scenario, where the victim VLMs are unknown (Ilyas et al., 2018; Bai et al., 2020b). To induce high energy-latency cost of black-box VLMs, we can leverage the transferability property (Dong et al., 2018) of our verbose images. We can first craft verbose images on a known and accessible surrogate model and then utilize them to transfer to the target victim VLM. The black-box transferability results across four VLMs of our verbose images are evaluated in Table 6. When the source model is set as ‘None’, it indicates that we evaluate the energy-latency cost for the target model using the original images. The results show that our transferable verbose images are less effective than the white-box verbose images but still result in a longer generated sequence.
Appendix C More tasks
To verify the effectiveness of our verbose images, we induce high energy-latency cost on two additional multi-modal tasks: visual question answering (VQA) and visual reasoning. Following Li et al. (2023b), we use VQAv2 dataset (Goyal et al., 2017) for VQA and GQA dataset (Hudson & Manning, 2019) for visual reasoning. We use BLIP-2 as the target model, and following the recommendations in Li et al. (2023b), we set the prompt template as “Question: {} Answer:”. Unless otherwise specified, other settings remain unchanged. Table 7 demonstrates that our verbose images can induce the highest energy-latency cost among three multi-modal tasks.
Appendix D More results of length distribution
We provide more results of the length distribution on MS-COCO dataset in Fig. 6 and ImageNet dataset in Fig. 6. The length distribution of generated sequences in the four VLM models exhibits a bimodal distribution. Specifically, our verbose images tend to prompt VLMs to generate either long or short sequences. We conjecture the reasons as follows. A majority of the long sequences are generated, confirming the effectiveness of our verbose images. As for the short sequences, we have carefully examined the generated content and observed two main cases. In the first case, our verbose images fail to induce long sentences, particularly for BLIP and BLIP-2, which lack instruction tuning and have smaller parameters. In the second case, our verbose images can confuse the VLMs, leading them to generate statements such as ‘I am sorry, but I cannot describe the image.’ This scenario predominantly occurs with InstructBlip and MiniGPT-4, both of which have instruction tuning and larger parameters.
Appendix E joint optimization of both images and texts
VLMs combine vision transformers and large language models to obtain an enhanced zero-shot performance in multi-modal tasks (Liu et al., 2023b; Li et al., 2021; 2023b; Ma et al., 2022; 2024). Hence, VLMs are capable of processing both visual and textual inputs, enabling them to handle multi-modal tasks effectively. In this section, we will adopt our proposed losses and the temporal weight adjustment algorithm to optimize both the imperceptible perturbation of visual inputs and tokens of textual inputs to induce high energy-latency cost of VLMs. For the optimization of textual inputs, we update a parameterized distribution matrix to optimize input textual tokens, as suggested in Guo et al. (2021). The number of optimized tokens is set as 8. Besides, Adam optimizer (Kingma & Ba, 2015) with a learning rate of 0.5 is used to optimize input textual tokens every iteration. Moreover, both the imperceptible perturbation of visual inputs and tokens of textual inputs are jointly optimized. Unless otherwise specified, other settings remain unchanged. Table 8 demonstrates that our methods can still induce VLMs to generate longer sequences than other methods under the joint optimization of both images and texts of VLMs.
Appendix F Visual interpretation
We show more visual interpretation results of the original images and our verbose images by using Grad-CAM (Selvaraju et al., 2017). The results are demonstrated in Fig. 7.
Appendix G Additional discussions
We conduct additional discussions, including the feasibility analysis of an intuitive solution to study whether limitation on generation length can address the energy-latency vulnerability, the model performance of three energy-latency attacks, image embedding distance between original images and verbose counterpart, energy consumption for generating one verbose image, and standard deviation results of our main table.
An intuitive solution to mitigate the energy-latency vulnerability is to impose a limitation on generation length. We argue that such an intuitive solution is infeasible and the reason is as follows.
(1) Users have diverse requirements and input data, leading to a wide range of sentence lengths and complexities. For example, the prompt text of ‘Describe the given image in one sentence.’ and ‘Describe the given image in details.’ can introduce different lengths of generated sentences. We visualize a case in Fig. 8. Consequently, service providers often consider a large token limit to accommodate these diverse requirements and ensure that the generated sentences are complete and meet users’ expectations. Previous work, NICGSlowDown (Chen et al., 2022c), also states the same view as us. Besides, as shown in Table 7, the results can demonstrate that our verbose images are adaptable for different prompt texts and can induce the length of generated sentences closer to the token limit set by the service provider. As a result, the energy-latency cost can be increased while staying within the imposed constraints.
(2) We argue that this attack surface about availability of VLMs becomes more important and our verbose images can induce more serious attack consequences in the era of large (vision) language models. The development of VLMs and LLMs has led to models capable of generating longer sentences with logic and coherence. Consequently, service providers have been increasing the maximum allowed length of generated sequences to ensure high-quality user experiences. For instance, gpt-3.5-turbo and gpt-4-turbo allow up to 4,096 and 8,192 tokens, respectively. Hence, we would like to uncover that while longer generated sequences can indeed improve service quality, they also introduce potential security risks about energy-latency cost, as our verbose images demonstrates. Therefore, when VLM service providers consider increasing the maximum length of generated sequences for better user experience, they should not only focus on the ability of VLMs but also take the maximum energy consumption payload into account.
G.2 evaluation performance of energy-latency manipulation
We evaluate the BLEU and CIDEr scores for sponge samples, NICGSlowDown, and our verbose images on MS-COCO for BLIP-2, as illustrated in Table 9. Concretely, our verbose images generate the longest sequence, extending it to 226.72, while sponge samples, despite their superior captioning performance, only increase the length to 22.53. Furthermore, both the length of the generated sequence and the captioning performance of our verbose images outperform those of NICGSlowDown.
G.3 image embedding distance between original images and attacked counterpart
We adopt the image encoder of CLIP to extract the image embedding. Then the image embedding distance is calculated as the cosine similarity of both original images and the corresponding sponge samples (Shumailov et al., 2021), NICGSlowDown (Chen et al., 2022c), and our verbose images. The results are shown in Table 10 which demonstrates that these methods are similar in image embedding distance.
G.4 energy consumption for generating one attacked image
We calculate the energy-latency cost for the generation of one verbose image and show the results in Table 11. It can be observed that the energy-latency cost during attack increases with the number of attack iterations, along with an increase in energy consumption for generating one verbose image. This finding provides valuable insights into the positive relation between the attack performance and the energy consumption associated with the generation of verbose images.
Besides, the overall energy-latency cost of generating a single verbose image is higher than that of using it to attack VLMs. Therefore, it is necessary for the attacker to make full use of every generated verbose image and learn from DDoS attack strategies to perform this attack more effectively. Specifically, the attacker can instantly send as many copies of the same verbose image as possible to VLMs, which increases the probability of exhausting the computational resources and reducing the availability of VLMs service. Once the attack is successful and causes the competitor’s service to collapse, the attacker will acquire numerous users from the competitor and gain significant benefits.
G.5 standard deviation results
The standard deviation results for length of generated sequences, energy consumption, and latency time are shown in Table 12.
Appendix H Additional ablation studies
We conduct additional ablation studies, including the effect of different sampling policies and different maximum lengths of generated sequences.
In our default settings, VLMs generate the sequences using nucleus sampling method (Holtzman et al., 2020) with and temperature . Besides, we present the results of three other sampling policies, including greedy search, beam search with a beam width of 5, and top-k sampling with . As depicted in Table 13, our verbose images can induce VLMs to generate the longest sequences across various sampling policies, demonstrating that our verbose images are not sensitive to the generation sampling policies.
H.2 different maximum lengths of generated sequences
Table 14 presents the results of five categories of visual images under varying maximum lengths of generated sequences. It shows that as the maximum length of generated sequences of VLMs increases, the length of generated sequences becomes longer, leading to higher energy consumption and longer latency time. Furthermore, our verbose images can consistently outperform the other four methods, which confirms the superiority of our verbose images.
Appendix I Grid search
We conduct the experiments of grid search (Gao et al., 2023) for different parameters of loss weights and different momentum values.
We set parameters of loss weights as , , , , , and during the optimization of our verbose images. These parameters are determined through the grid search. The results of grid search are shown in Table 21, Table 21, Table 21, Table 21, Table 21, and Table 21, which demonstrates that our verbose images with , , , , , and can induce VLMs to generate the longest sequences.
I.2 different momentum values
We show the results of our verbose images in different momentum values in Table 21. These results demonstrate that it is necessary to adopt the addition of momentum during the optimization.
Appendix J Visualization
We visualize the examples of the original images and our verbose images against BLIP, BLIP-2, InstructBLIP, and MiniGPT-4 in Fig. 9, Fig. 10, Fig. 11, and Fig. 12. It can be observed that VLMs with more advanced LLMs (e.g., Vicuna-7B) generate more fluent, smooth, and logical content when encountering our verbose images.