VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use

Yonatan Bitton, Hritik Bansal, Jack Hessel, Rulin Shao, Wanrong Zhu, Anas Awadalla, Josh Gardner, Rohan Taori, Ludwig Schmidt

Introduction

A long-standing challenge for artificial intelligence is to build general-purpose assistants that can, in collaboration with humans, solve diverse and never-before-seen tasks . For textual tasks, several recent works have shown that fine-tuning language models such as GPT-3 and LLaMA with supervised instruction+response examples enables them to respond to imperative requests and questions without task-specific training. Zero-shot generalization is promising not only for standard academic benchmarks, but – perhaps more-so – for creative, useful, and real-world queries that downstream users of language technologies are likely to make.

On the multimodal side, recent instruction-following vision-language models also provide a zero-shot interface. Given an image (or multiple images) and a query (e.g., “how many apples are in this image?” or “What is this?” or “Write a poem in the style of Robert Frost about this scene.”) a textual response is provided. Recent works like OpenFlamingo , LLaVA and others , have implemented this interface with promising initial results. Although standard benchmarks like VQAv2 and COCO captioning are commonly used to assess performance, less is know about how models perform on broader, open-ended queries that resemble real-world user behavior. Evaluations of such queries typically rely on informal and qualitative approaches.

To support quantitative evaluation for this setting, we present VisIT-Bench (Visual InsTruction Benchmark), a dynamic benchmark consisting of 592 challenging vision-language instructions. Each instance contains an instruction, input image(s), a instruction-conditioned caption (a human-crafted caption for the image(s)/instruction), and a human verified reference (Figure 1). Instructions are image-contextual imperative requests or questions, e.g., for an image of pancakes, a user asks “how can I cook this in a healthy way?”. Different from existing zero-shot evaluations, many of the instructions focus on open-ended generation requests (e.g., “write a poem…” or “what should I bring if I were to visit here?”).

We created VisIT-Bench to cover a wide array of “instruction families”. Our starting point was a set of 70 “wish-list” tasks such as “home renovation” and “gardening tips” collected by the authors:We recognize that promising applications may not be covered by our set; and we don’t necessarily advocate for deploying models in all cases we cover – we hope VisIT-Bench can help to quantify shortcomings and risks. each requiring varied high-level skills from recognition to complex reasoning (Figure 2). We derived 25/70 instruction families from benchmark tasks such as Visual Question Answering (VQA) and robust change captioning into a chatbot-style format (this reformatting differs from prior work , as we focus on open-ended chatbot style responses.). Notably, 10 of these repurposed tasks involve multiple images.

We started with 10 images for each instruction family. Our annotators, guided by an example, create a new instruction, and provide a (permissively licensed) image. For each instruction, we next collect instruction-conditioned captions – unlike prior work these descriptions are designed not only to describe the image in general, but also, surface information targeted to the instruction. Finally, we use instruction-conditioned captions to generate a reference candidate output from GPT-4; an additional human verification step discards GPT-4 references deemed to be incorrect.

We conduct a large-scale empirical comparison of multimodal instruction-following models using VisIT-Bench (§4). We first gather predictions for each instance from 7 candidate models. Then, we collect 5K human judgements of output quality by pitting model outputs head-to-head, and (in a forced-choice setup) crowd-sourcing pairwise preference judgements. This analysis not only reveals significant differences between models (e.g., that LLaVA-13b is generally preferred to Panda ), but also, that the human verified references in our corpus are preferred significantly more than the ones generated using multimodal models. We summarize head-to-head comparisons with two metrics: 1) Elo ratings , which provide relative “skill” rating estimates encoding the probability that model A will be preferred to model B; and 2) win rate versus our references, which provides an absolute metric. The best model according to human judgement is LLaMA-Adapter-v2 , yet it only wins in a pairwise setting against the reference in 27.4% of cases.

Finally, we design an automated evaluation for VisIT-Bench, utilizing GPT-4 to rank pairs of model responses based on factors like correctness, relevance, and fluency. Using the instruction-conditioned caption and the instruction, GPT-4 determines the better response between two options, expediting iteration compared to human preferences. We explore reference-free and reference-backed versions of this metric. Compared to various metrics (BLEU-4 , ROUGE-L , METEOR , CIDEr , and BERTScore ), our evaluation aligns best with human preferences. For example, it achieves a 94% agreement rate in the cases where all five annotators agree. See Figure 7 for a schematic of the process.

While it is difficult to a priori envision all possible scenarios under which more performant multimodal chatbots might be used, we hope VisIT-Bench can provide a path to improving vision-language models “in the wild.” Table 1 presents a summary of our contributions in comparison to the recent works in the evaluation of multimodal chatbots. We publicly release VisIT-Bench data, code, and automatic metrics to facilitate future model evaluations, available in https://visit-bench.github.io/.

VisIT-Bench: A Real-World Inspired VL Instruction-Following Benchmark

VisIT-Bench was built to emulate real-world applications of multimodal models through image-text tasks, creating an extensive and practical benchmark. These tasks, or ‘instruction families’, are seen as key capabilities of a high-performing vision-and-language model. Although our selections are not exhaustive, they provide a broad basis for evaluating beyond academic benchmarks. We prioritize family coverage vs. number of instances-per-task. The final corpus, comprising 592 instances and 1,159 public images, can be found at VisIT-Bench Sheet and VisIT-Bench Sheet Multi-Images. VisIT-Bench instances are either from 45 newly assembled instruction families or reformatted from 25 existing datasets (see Table 5). Notably, 10 instruction families cater to multi-image query scenarios (e.g., Figure 4).

The authors of this work perform an initial annotation step of curating instruction families. For each instruction family not derived from an existing task (45 out of 70), we designate a name for the family (e.g., “Contextual Knowledge of Events”) and identify an image-instruction pair that exemplifies the category, along with a sample response (“Martin Luther King Jr. is waving to acknowledge and greet the crowd of protesters […]”). 10 sample familes are in Figure 2.

The following steps are carried out in collaboration with crowdworkers, who receive an hourly wage of $18. These steps are outlined in Figure 3: (1) taking the image/instruction example as a guiding seed task crowdworkers formulate a new instruction that examines the same instruction family (“instruction generation”); (2) crowdworkers create detailed image captions that describe the image and allow an entity, relying solely on this text, to interpret and execute the given instruction successfully (“instruction-conditioned caption generation”); (3) crowdworkers assess the correctness of GPT-4’s response to the instruction (“model output evaluation”). We further elaborate on these steps using human annotators below.

25/70 instruction families (corresponding to 25*10=250 instances) are re-formatted versions of existing vision-language tasks (See Appendix D for full list).Users of VisIT-Bench should also cite the original datasets. This process involves re-formatting tasks into chatbot-style instruction/response versions. In re-formatting, we re-write instructions to retain the original task’s goal while maintaining the original images, see Figure 4. These repurposed tasks are integrated into our data collection process, ensuring uniformity between the chatbot-style answers in the full VisIT-Bench instances and the reinterpreted tasks.

Instruction Generation.

Here, annotators create a new instance from the same instruction family as a given example, along with an instruction and corresponding image. For instance, in Figure 3 (left), the instruction family is “Contextual Knowledge of Events”, and the example instruction is “Why is he waving? What happened in this event?” alongside an image of Martin Luther King, Jr. To collect images, annotators were instructed to use Openverse (https://openverse.org/) for Creative Commons licened images.

Instruction-Conditioned Caption Generation.

Annotators are provided with the image and instruction, and are tasked to construct a caption that is rich enough to allow an entity, solely receiving the text they author, to follow the instruction. This caption will later facilitate GPT-4 reference candidate generation, and will be used for text-only auto-evaluation. We call these instructions instruction-conditioned captions. See Figure 3 (middle) for an example: an annotator doesn’t just mention the skittles and a spoon, but, given the query regarding specific colors, they indicate the exact colors in detail.

Model Output Evaluation.

The goal of this stage is to gather human-validated reference chatbot responses for each multimodal instruction query. We initially obtain response candidates from GPT-4 given the instruction and the instruction-conditioned caption. GPT4’s prompt is: “Consider an image depicted by: ’. Now, briefly follow this instruction, and you can add a short explanation: ’. Response: This prompt is employed for both single and multiple image instances, with appropriate modifications for the latter. Then we verify each response with human annotators.An alternate annotation scheme would have been to task annotators to write target responses from scratch. The rationale for using GPT-4 verification instead is derived from prior results that show promising human-machine collaboration of this form . If a response is marked incorrect, the annotator identifies whether the issue lies with the detail level of the instruction-conditioned captions or with GPT-4’s response itself. For VisIT-Bench, we discard any case marked as incorrect for either reason.The annotators are also tasked to screen for any offensive, unsound, or harmful advice present in the responses. We did not find or discard any instances. An example is given in Figure 3 (right), where GPT-4’s candidate reference response aims to answer a question about a chess position (which it does so incorrectly, and thus, the instance is discarded).

2 Data Collection Annotation and Results

We conduct the data collection steps in Figure 3 using Amazon’s Mechanical Turk (MTurk) platform. Prior to annotating, each MTurk worker passed a qualification test, which involved five to ten sample tasks designed to assess their ability to generate high-quality annotations. More detailed information about the execution process and full user interface examples can be found in Appendix C.

Our annotation results are summarized in Table 2. We measure the throughput of the collection and filtration pipeline. For single-image instances, our pipeline’s yield was 91.5% from the original candidate set. However, the success rate dropped to 63.0% in the more complex multi-image tasks, accompanied by an uptick in issues either in the captions (6.0%) or GPT-4’s responses (30.0%). This drop suggests that multi-image queries may pose a more difficult data collection challenge.

VisIT-Bench Analysis

We analyze the tasks, images, and instruction-conditioned captions of VisIT-Bench.

To clarify the role of the instruction-conditioned captions we collect, we conducted an experiment covering 150 single-image instances. Instead of using our instruction-conditioned captions, we use BLIP2 image captions, which is a state-of-the-art image captioning model. We extract image captions, and feed them to GPT-4 as detailed earlier, to provide a text-based chatbot response. This process is depicted in Figure 5.

We manually evaluated whether the resulting output accurately followed the instructions. We find that while instruction-conditioned captions led to correct outputs in 91% of the cases, the success rate fell to 31% when using BLIP2 captions (Table 2). These results highlight the importance of instruction-conditioned captions in the construction of VisIT-Bench, and show that the instances in our dataset are sophisticated enough such that most are not solvable by using a simple Socratic model baseline of caption →\rightarrow LLM.

2 What skills are required for VisIT-Bench?

The full list of instruction families we cover are in Appendix Table 6. Following , for the VisIT-Bench instructions, we extract the most frequent root verbs and their direct nouns (a full plot is in Figure 6). The most common include: ‘answer question’, ‘write story/poem’, ‘create title’, etc. There’s also a long-tail of diverse requests that demand comprehension, commonsense, and cross-modal understanding, e.g., ‘identifying objects’ to ‘need ingredient’ to ‘connect device’. Additional qualitative examination reveals a range of underlying skills required ranging from ‘emotion identification’ to complex reasoning tasks such as ‘paper folding’.

3 What is contained in VisIT-Bench images?

We detect all the COCO objects present in the images from our dataset using Yolov5-L ; The most common detected objects in VisIT-Bench are “person” (∼\scriptstyle\sim 900 detections), chair, and car (∼\scriptstyle\sim 100). But, a long tail of rarer objects exists as well: full distribution in Appendix Figure 10. Overall, to perform well at VisIT-Bench, a model must account for a broad range of scenes and objects.

Experiments

We evaluate a range of state-of-the-art publicly accessible vision-and-language chatbots on the 592 instances in VisIT-Bench. In §4.1, we provide the details of the instruction-following models in our benchmark. Following this, we collect the human preferences for pairwise model generations to achieve a human-guided Elo ranking and the win-rates against the reference of the models in §4.2. We then develop automatic evaluation on VisIT-Bench in §4.3, that can be scaled and improved given new and improved models. Finally, we establish the trustworthiness of our automatic evaluation method by performing agreement analysis with the human judgments in §4.3

We evaluate LLaVA-13B , InstructBLIP-13B , MiniGPT4-7B , mPLUG-Owl-7B , LlamaAdapter-v2-7B , PandaGPT-13B , VisualChatGPT , Multimodal GPT , OpenFlamingo v1 , Otter v1 , Lynx and idefics . For the execution-based VisualChatGPT , we implement a chat window for each sample, hold inputs and intermediate chains of thoughts and actions in memory, and feed the images and the instruction sequentially. For OpenFlamingo and Otter , we feed the image(s) and the instruction in an interleaved format. For the others, we feed the image to the vision feature extractor and feed the instruction as a prompt to the text encoder.Following the authors’ instructions, we run all models using default settings to obtain the best possible responses. We include specific samples for reproducibility. We acknowledge hyperparameter impact and are willing to reassess submissions to VisIT-Bench if conditions were sub-optimal.

2 Human Evaluation

We collect 5K pairwise human preference judgements across an initial set of 6 models and the human-verified references. For 1K uniformly randomly sampled tuples of (query, model A, model B), we collect 5 crowdworker judgements each. Preferences are collected in a “forced choice” setting, annotators are instructed to decide based on accuracy, helpfulness, and detail. We provide the template for the human annotation process in Appendix Figure 15. We summarize the results with two metrics:

Relative metric: Elo We follow and compute Elo ratings, treating each pairwise human judgement as a “match.”We use the following code/hyperparameters for Elo ratings: https://github.com/lm-sys/FastChat/blob/main/fastchat/serve/monitor/elo_analysis.py The difference between the Elo ratings of two different models provides an estimate for the win probability when pitting model A vs. model B. More details are in Appendix E.

Absolute metric: Win rate vs. reference. We provide a win-rate vs. the human-verified reference. We use the 1.4K pairwise human judgments where one of A or B is the reference. We report the percent of cases where the human judge prefers the output from that model vs. the human-verified GPT-4 reference output. Because we do not allow for ties in our forced-choice setup, if the annotator believes the responses are of equal quaity, they choose one arbitrarily.

Table 3 contains the Elo and win-rate vs. reference. In terms of Elo, the Human Verified GPT-4 reference achieves a higher rating than all alternatives, validating the quality of our reference set: concretely, for our Elo settings, the reference (Elo =1223) has an estimated win-rate over one of the best performing models, LLaVA, (Elo =1085) of 69%, and an estimated win rate of 93% against the lowest performing model in this setup, PandaGPT (Elo =786). This result can partly be explained by the training process of the underlying models: The improved performance of LLaVA (13B) might be attributed to its fine-tuning process, which utilized 150K instruction-tuning data that is rich in both diversity and quality. Interestingly, despite achieving a slightly lower Elo (the computation of which is based on all head-to-head “matches”, rather than just ones against the human reference), LlamaAdapter-v2 (7B) wins with the highest rate against the reference. However, the complexity and variety of models and tasks in VisIT-Bench makes it challenging to definitively pinpoint the factors influencing performance. While we make a preliminary attempt to unravel these intricacies in Section 4.3, a comprehensive understanding will necessitate more nuanced and extensive future research.

3 Automatic Evaluation and Leaderboard

Because it is costly to gather human pairwise preference judgements for new model submissions, to support faster model development, we seek an automatic evaluation procedure that produces high correlation with our human evaluation setup.

We consider several existing reference-backed evaluation metrics: BLEU-4 , ROUGE-L , METEOR , CIDEr , and BERTScore , we use the RoBERTa-Large english version , treating the human-verified GPT-4 reference as the evaluation reference. We additionally report two baseline metrics: random, which assigns a random score without accounting for the candidate, and length, which assigns a score equal to the number of non-whitespace tokens in the candidate. Beyond existing metrics and baselines, following the recent line of work utilizing API-accessed LLMs with a prompt for automatic evaluation , we consider two GPT-4 backed evaluation metrics.

Specifically, we provide the LLM with: 1) a system prompt describing the desired evaluation behavior; 2) the instruction-conditioned caption for the image; 3) the instruction to be followed; and 4) two candidate generations dubbed “Response A” and “Response B”. We also consider a reference-backed version where the human-verified reference is provided as well. We provide our prompts in Appendix F. To mitigate potential biases in “A” and “B” positioning, for all pairs of candidates, we run two queries covering both possible orderings. Our prompt encourages the model to think step-by-step so that its chain-of-thought process is made explicit . Despite strongly encouraging the model to select between the two references in a forced-choice setup, it sometimes refuses and outputs “tie” which we account for later. We call the reference-free version of this metric “GPT4-no-ref”, and the reference-backed version of this metric “GPT4-ref”.

Evaluating evaluation metrics.

We measure the correlation between the candidate metrics and human judgements using a pairwise framework. Specifically, we use a subset of the 5K pairwise human judgements in § 4.2. For 690 pairwise instances where both candidate instances are model-generated (rather than human-verified references), we have 5 pairwise judgements from crowd-workers. For 336 pairs, there is 5/5 agreement, for 200 pairs, there is 4/5 agreement, and for 154 pairs, there is 3/5 agreement. For each metric, we measure the percent of time the metric is able to accurately reconstruct a majority vote judgement from the 5 crowdworkers. The newly proposed GPT-4 based metrics sometimes outputs “tie” (this happens in 10-15% of cases overall) – for fair comparison with the other metrics in forced choice setting, we randomly choose one of the two options when GPT-4 reports a tie.

The results are in Figure 9, with GPT-4-no-ref best aligns with human correlation. The best performing metric is our newly proposed GPT-4 based metric, which accurately reconstructs majority-vote pairwise human judgments better than alternatives (p<.05p<.05; binomial proportion CI nonoverlapping). For example, for instances where 5/5 annotators agree, GPT4-no-ref, with no reference, accurately reconstructs human judgment 93% of the time, whereas the next best metrics BERTScore/METEOR/ROUGE-L reconstruct accurately 80%/78%/70% of the time; among the metrics we consider, these are reasonable options for static/offline evaluation without relying on OpenAI API access, especially when compared to our length baseline metric, which achieves only 60%. Notably, the reference-backed version of the newly proposed GPT-4 based metric achieves comparable (but slightly worse) performance compared to the reference-free version. Thus, we adopt the reference-free version, which additionally enables us to place the references themselves into the the Elo setup, because they are not used in the prompts.

System-level Correlation. We summarize the LLM’s pairwise judgements using the same metrics as introduced in §4.2, Elo ratings and win rate vs. reference, but instead of using a human judge, we use our reference-free GPT-4 based metric. The results are in LABEL:tab:table_auto_scoring_results. Notably, among the 7 systems for which we gathered human ratings for, the automatic metric produces the same ordering compared to human evaluation (ρ=1.0\rho=1.0, p<.01p<.01).

Shortcomings of proposed metric. While the relative ranking of models produced by the automatic metric correlates strongly with the ranking produced by human judgements, the win rate vs. reference according to human judgement (Table 3) are higher overall compared to the win-rate vs. reference according to the automatic metric LABEL:tab:table_auto_scoring_results. One plausible explanation for this discrepancy is that GPT-4, as an evaluation model, may prefer responses that closely match its own response distribution.

Per-category results. In Figure 8, we plot the win-rate vs reference for the models across all the single-image instruction families. We find that there is no model that performs the best and worst across all the instruction families. Thus, VisIT-Bench aids in highlighting the strengths and weaknesses of the instruction-following models along various real-world use-cases.

Related Work

Multimodal Models for Image-Text Understanding: Recently, the field of machine learning has experienced a rapid proliferation of new models which can perform various image-text tasks . This growth has been driven by several factors, including the emergence of large-scale multimodal datasets (e.g. LAION-5B , Multimodal C4 ), improved software and hardware frameworks, and advances in modality-specific models such as language models (e.g., ). Our work specifically evaluates models which can generate textual outputs, given one or more images, and text. Recent examples of such models include LLaVA , mPLUG-Owl , InstructBLIP, LLaMA-Adapter, Flamingo and OpenFlamingo , PandaGPT , and GPT-4 (which reports multimodal capabilities but has not yet seen a release of the multimodal variant).

Instruction Following: “Instruction-following” is an emerging paradigm for training models via language, where instead of being trained to complete only a single, fixed task (such as image classification or captioning), models are trained to follow textual instructions that describe an arbitrary task, with the aim of generalizing to novel instructions. Examples of instruction-following models include Alpaca , LLaMA-Adapter , Koala , InstructBLIP , LLaVA , and mPLUG-owl . As the downstream capabilities of these models are influenced by the quality of the training dataset, there has also been extensive work on developing instruction-following datasets .

To build these models, two broad approaches have been shown to be effective. One approach focuses on leveraging existing pretrained task-specific tools such as image captioners , object detectors and text-to-image generators by either creating multimodal prompt interfaces or by executing LLM-generated programs . The other approach focuses on building a single pretrained model that can follow instructions by supervised finetuning on multimodal vision-language data.

Despite the success of both these approaches on the existing vision-language datasets e.g., VQA, GQA, Image Captioning , there is a lack of a high-quality benchmarking dataset for multimodal instruction-following tasks that reliably replicates the way in which humans would interact with multimodal chatbots in the wild. Similar to the image-text models discussed above, many instruction-following models have been released directly as open-source without undergoing peer review or thorough evaluation. As a result, the effectiveness of these models for many tasks is not well-understood.

Benchmarks for Machine Learning: High-quality evaluation datasets have served both to (re)assess, and to accelerate, progress on many machine learning tasks . For example, our work draws particularly from the fields of computer vision and natural language processing, where benchmarking datasets have been critical drivers of progress. On the vision side, datasets such as ImageNet and CIFAR have proven to be critical yardsticks of progress. On the language side, benchmarks such as SQuAD , SST , GLUE/SuperGLUE and more seen wide use. Recent work has indicated that improvements on these high-quality benchmark datasets is not the result of overfitting, and is a reliable indicator of genuine progress beyond the benchmark data .

However, high-quality benchmarking datasets and evaluation methods do not yet exist for multimodal instruction-following. As a result, it is difficult to assess progress in this direction, which both reduces the field’s ability to identify true breakthroughs and increases vulnerability to potential pitfalls of evaluation that have hampered progress in other areas of machine learning .

Conclusion

We introduce VisIT-Bench, a dynamic benchmark providing a broad evaluation of multimodal chatbots’ capabilities. Going beyond prior efforts, VisIT-Bench’s collection process centers potential real-world use cases, and 70 diverse instruction families encompassing a range of tasks from recognition to complex reasoning. Our benchmark not only offers human-verified reference outputs for all examples but also gives an Elo-based ranking system for multimodal chatbots that correlates with human judgements. Our experiments reveal a gap between model and human performance.We release data, code, and automatic metrics, encouraging community involvement. We hope VisIT-Bench can provide a new quantification of progress and shortcomings of multimodal AI systems.

Limitations

Although VisIT-Bench covers a wide spectrum of potential use-cases, it does not incorporate every possible vision-language task. We hope to add more categories of tasks over time. In terms of dialogue, VisIT-Bench concentrates on single-turn instances with one instruction and response. This does not encompass multi-turn interactions between users and chatbots, which presents a promising direction for future research. Our study focuses on image-text modalities. Future extensions could expand the scope to include other modalities like audio and video, enabling a more comprehensive evaluation. Additionally, while the dataset offers a wide variety of tasks, a larger number of examples per category could provide more depth. Finally, while our GPT-4 based metric correlates well with human judgement both at the instance level and at the system level, we see some evidence that the GPT-4 based metric has a stronger preference for GPT-4 based generations compared to humans. Thus, models which train, e.g., by distilling from GPT-4 outputs, may have an unfair advantage on our evaluation.

Acknowledgements

We thank Pang Wei Koh, Ashima Suvarna, Nitzan Guetta and Roee Aharoni for their valuable feedback. Hritik Bansal is supported in part by AFOSR MURI grant FA9550-22-1-0380. RT is supported by the NSF GRFP under Grant No. DGE 1656518.

References

Appendix

Appendix A License and Intended Use

The VisIT-Bench dataset, along with its various contributions such as instructions, reference outputs, and model ranking annotations, is licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). This license applies to all the images we have directly contributed, each of which carries a public license specification in the “public images metadata” field within the dataset sheets. However, the dataset also incorporates images sourced from pre-existing collections. For these images, the original licensing terms are respected and remain applicable.

VisIT-Bench’s primary purpose is to function as a dynamic benchmark that continuously evolves and evaluates instruction-following vision-language models. In the current landscape, commercial chatbots are often trained on non-disclosed and non-public datasets, which raises concerns about potential data contamination and inadvertent training on our evaluation data . This risk is further highlighted by recent studies . To mitigate such concerns, we have chosen to withhold the complete VisIT-Bench test set from public disclosure, while still making the images and instructions available for direct download. Researchers, however, can utilize VisIT-Bench to its full potential as a dynamic benchmark by submitting their model predictions for evaluation. We will assess their models using the undisclosed test set, ensuring the ongoing evolution of the benchmark. Moreover, we are open to releasing the test data upon receiving reasonable and justified requests, particularly when additional analysis is necessary, provided that requesters agree to our non-contamination policy which prohibits the use of this data for training commercial chatbots. This approach strikes a balance between the need for robust model evaluation and the mitigation of potential data contamination.

Appendix B Dataset Analysis

Appendix C Interfaces for Collecting Human Annotations

In this section, we provide the templates we used to collect human annotations for the instruction generation (Figure 11), the dense caption generation (Figure 12), the model verification (Figure 13 and Figure 14), and the model rating (Figure 15).

Appendix D Existing Datasets incorporated in VisIT-Bench

In Table 5, we listed the existing datasets that are incoprated in our VisIT-Bench. Among these datasets, 15 contain a single image in each sample pair, and 10 require reasoning based on multiple images.

Appendix E Elo Rating

For many years, the Elo rating has been popular in ranking players in zero-sum games such as chess . Recently, it has been adopted to rate large language models (LLMs) against each other on the user instructions. In this work, we adopt the same strategy to rank a set of instruction-following vision-language models, that can grow dynamically with further advances in the field.

Given two multimodal chatbots Ca\mathcal{C}_{a} and Cb\mathcal{C}_{b} with their absolute Elo rating Ra\mathcal{R}_{a} and Rb\mathcal{R}_{b}, respectively. Simply put, the probability of Ca\mathcal{C}_{a} winning over Cb\mathcal{C}_{b} in a head-to-head battle is given by:

In practice, calculating the Elo rating requires us to set hyperparameters to decide the weightage for each win and loss in a head-to-head battle between two models. In our work, we use the open implementation of Elo for LLMs by FastChat at https://github.com/lm-sys/FastChat/blob/main/fastchat/serve/monitor/elo_analysis.py.

Appendix F GPT-4 Pairwise Evaluation Prompts

The specific prompts we use to extract pairwise judgements from our language model are provided in Table 16 (reference-free version) and Table 17 (reference-backed version). When applied to GPT-4 , these prompts usually solicit a definitive pairwise response by the model. But, in some cases, the model either produces a pairwise judgement in an unexpected format, or, refuses to issue a judgement at all. For cases like these, we issue an additional query to ChatGPT to extract an answer (or decide there is no answer) using an additional prompt, given in Table 18. If after this step there is still no definitive pairwise judgment, we call the result a tie.