The Role of Chain-of-Thought in Complex Vision-Language Reasoning Task

Yifan Wu, Pengchuan Zhang, Wenhan Xiong, Barlas Oguz, James C. Gee, Yixin Nie

Introduction

Large language models (LLMs) have shown impressive performance in many language tasks, fostering the development of general AI assistants. An emerging trend in AI research aims to expand the potential of LLMs beyond text perception, incorporating visual data for more comprehensive models Zhu et al. (2023); Liu et al. (2023); Jun et al. (2023). GPT-4V(ision) was recently released and has garnered significant interest for its exceptional abilities in multimodal perception and reasoning.

However, in complex vision-language tasks, although GPT-4V significantly outperforms the current state-of-the-art, it still lags behind human performance, as demonstrated in Section 3. We try to understand this challenge. Visual understanding extends beyond mere perception Zellers et al. (2019). Complex visual-language tasks demand recognition-level perception, such as localizing and classifying objects and their attributes, as well as cognition-level reasoning, such as inferring intents, goals, and temporal and social dynamics. Humans can seamlessly integrate these two stages, but this remains challenging for LLMs with vision modality (vision-LLMs).

The Chain-of-Thought strategy, known for its effectiveness in language tasks by breaking them into sub-tasks with intermediate steps, has been studied extensively Wei et al. (2022); Yao et al. (2023); Lyu et al. (2023), In this research, we investigate whether this method can enhance vision-language tasks, particularly those requiring complex reasoning. However, when confronting complex visiolinguistic tasks, current LLMs struggle to figure out the proper reasoning paradigm by themselves. Unlike math problems, which have a specific task-driven reasoning approach, these tasks may not have a clear and unique reasoning path leading to the final output.

Research on human cognition provides clues to a proper reasoning mode. Visual information propagates through two streams (Figure 1 (a)). The ventral stream (or the ‘what pathway’) is involved with object identification, while the dorsal stream (or the ‘where pathway’) processes objects’ spatial location. These two streams decouple recognition into local processing modules. The cognition part, i.e., the reasoning and decision-making function, is mainly executed in the frontal lobe. Such modularization is a significant characteristic of brain structure and function Gu et al. (2015); Bassett and Sporns (2017). Inspired by how humans process signals, we design prompts with specialized modules for steps with different emphases.

The latest initial attempt at Chain-of-Thought on vision-language tasks focuses on the recognition task Yang et al. (2023a), while another only provides qualitative analysis Yang et al. (2023b). In this work, we analyze the prompting strategy for complex vision-language tasks analogous to the brain’s information processing. Our Description (information-extracting) then Decision (decision-making) strategy consistently improves the performance across different experimental settings.

Probing Task

The ideal probing task should present sufficient challenges in both visual recognition and text comprehension. Furthermore, it should demand intricate reasoning, implying that there is a logical connection between the image and text facts leading to the final output. Consequently, we selected Winoground Thrush et al. (2022) as a case study for our experiments. Winoground is both a dataset and a task designed specifically to evaluate visio-linguistic compositional reasoning. This task involves being given two images and two captions, with the goal of correctly matching each image with its corresponding caption. Notably, both captions utilize the exact same set of words, but they are arranged in different orders so that each caption primarily describes one of the two images.

The Winoground task is notably challenging Diwan et al. (2022). It requires not only a robust visual recognition ability to identify small or blurred objects and differentiate attributes and actions among proximate objects but also sophisticated visio-linguistic compositional reasoning. For instance, solving Winoground sometimes necessitates interpreting images non-literally due to the idiomatic language usage in a caption (e.g., “it starts with Z and ends with A” might describe an image of a zebra). In other instances, one might need to infer past or future events from the scenes depicted in the image (e.g., "the cup on the left was filled first, and the cup on the right was filled second"). Examples of Winoground can be found in Fig 2.

The original Winoground task consists of two experimental setups: the text score and the image score. The text score evaluates the model’s ability to select the correct caption from two given captions when provided an image. Conversely, the image score assesses the model’s ability to choose the appropriate image from two available options given a caption. The original Winoground was tested based on the feature embedding similarities between captions and images in vision-language models, such as CLIP Radford et al. (2021). To assess recent large vision-language models like GPT-4V, we reformulate the Winoground as a choice-based visual question-answering task. Formally, given images I0I_{0} and I1I_{1} and captions C0C_{0} and C1C_{1}, the text score for a data point (C0C_{0}, I0I_{0}, C1C_{1}, I1I_{1}) is computed as follows:

, where f(⋅)f(\cdot) is the large language model that provides answers through a generation process. For a data point to be classified as correct, both images in a pair must align with their textual descriptions.

Similarly, for the image choice task, the score is determined as follows:

The Winoground dataset comprises 400 pairs of images and their corresponding captions.

Evaluation

We first present the Winoground benchmark in Table 1 to evaluate the performance of GPT-4V in comparison to others. Note that while the original task evaluation is based on the encoding similarities between images and text, the setups for large vision-language models may vary slightly among methods. For example, TIFA Hu et al. (2023) and VQ2 Yarom et al. (2023) ask a series of questions given one image and one caption, then accumulate the scores. MMICL Zhao et al. (2023) is given two images and two captions in each prompt. While in our set-up, we are given one image with two captions, or two images with one caption, as formulated in Eqn 1 and Eqn 2. However, these variations will not impact the primary objective of our study, which is to analyze the effects of various prompt configurations.

We analyze whether using a chain-of-thought prompt strategy, by decomposing the visual-language complex reasoning task into recognition and reasoning steps, can be beneficial compared to directly asking the model for the answer. The quantitative results are reported in Table 1, and the qualitative examples are shown in Figure 2 (a), with the prompt configurations detailed accordingly.In the following, we present the prompt we use to evaluate GPT-4V’s text score (Text) and image score (Image), without and with CoT.

GPT-4V (Text): [‘image-0’ or ‘image-1’] Does this image present (A) [‘caption-0’], or (B) [‘caption-1’]? Note, you must choose one of the two options.

GPT-4V CoT (Text): [‘image-0’ or ‘image-1’] Does this image present (A) [‘caption-0’], or (B) [‘caption-1’]? First, describe the image information relevant to the question. Then, provide your answer. Note you must choose one of the two options.

GPT-4V (Image): [‘image-0’], [‘image-1’] Which image better aligns with the description [‘caption-0’ or ‘caption-1’]? The first image or the second image? Note you must choose one of two options.

GPT-4V CoT (Image): [‘image-0’], [‘image-1’] Which image better aligns with the description [‘caption-0’ or ‘caption-1’]? The first image or the second image? First, describe the image information relevant to the question. Then, provide your answer. Note you must choose one of two options.

As shown in Table 1, there are consistent improvements in both the Text and Image score settings, leading to a 50%50\% improvement (from 39.25 to 58.75) in the Group score. The latter represents the percentage of data points that have both Text and Image probes answered correctly. The “Description then Decision" strategy is particularly beneficial for the Image score setting, which improves from 46.25 to 68.75. We observe that for GPT-4V, the image score is significantly lower than the text score. This could be attributed to the use of multiple images as input. The Chain of Thought (CoT) strategy significantly improves the image score and largely closes the gap with the text score.

The difference in generative processes between these two prompts is as follows: for GPT-4V, the answer is conditioned on the image and question:

While for GPT-4V CoT, the answer is generated based on the image, question, and description:

Although generating descriptions from images does not introduce new information, our results show that this step simplifies the reasoning or decision-making for the model by translating visual signals into textual ones. To assess generalizability, we conducted the same experiments on other vision-based large language models, LLaVA Liu et al. (2023) and InstructBLIP Dai et al. (2023), as shown in Table 2 (a). The “description then decision" strategy consistently improves performance.

More qualitative results demonstrating the effectiveness of our “Description then Decision” prompt strategy are shown in Figure 4 and Figure 5.

2 The Effect of Two-turns Prompt

While examining the output of “GPT-4V CoT", we observed instances where correct descriptions were followed by incorrect answers. To simplify this generation process, we conducted experiments to divide the recognition and reasoning into two turns, on the text setting, as illustrated in Figure 2 (b).

To examine the quality of the first turn, i.e., the question-relevant image description, we conducted ablation studies employing GPT-4 in the second turn, which generates responses without image access. In the following, we present the prompts we used for different settings of the second turn.

GPT-4 QA: [‘image description’] Based on this image description, does this image depict (A) [‘caption-0’], or (B) [‘caption-1’]? Note, you must choose one of the two options.

GPT-4 CoT: [‘image description’] Based on this image description, does this image depict (A) [‘caption-0’], or (B) [‘caption-1’]? First, analyze the two options, then provide your answer. Note, you must choose one of the two options.

GPT-4V QA: [‘image description’] Does this image depict (A) [‘caption-0’], or (B) [‘caption-1’]? Note, you must choose one of the two options.

GPT-4V CoT: [‘image description’] Does this image depict (A) [‘caption-0’], or (B) [‘caption-1’]? First, analyze the two options, then provide your answer. Note, you must choose one of the two options.

The results are summarized in Table 2 (b). Our observations are threefold: 1) The two-turn prompt notably improves results, from 75.25% to 79.50%. 2) The QA performance of GPT-4 reflects the quality of the image descriptions generated by GPT-4V. The “GPT-4V Desp + GPT-4 CoT (2-turns)" experiment achieved a 78.75% performance rate, validating the preciseness of GPT-4V’s image descriptions. 3) GPT-4V performs slightly better than GPT-4, indicating that while the text format makes reasoning easier, fully enumerating all related information presented in an image remains challenging.

Error Analysis

The performance of GPT-4V on the Winoground is remarkable, yet it’s essential to pinpoint its limitations for a comprehensive understanding. Our error analysis, which adapts Winoground’s tagging categorization, examines GPT-4V’s performance on vision-language understanding tasks that require various recognition and reasoning skills.

The experiment selected for this analysis is "GPT-4V Desp + GPT-4V CoT (2-turns)," which achieves 80% text score across the dataset. As per Table 3, GPT-4V is adept at interpreting symbolic categories, which represent symbolic representations, e.g., a child’s drawing. The "Series" tag indicates that a pair of images comes from the same photographic series, which might include identical individuals and scenes. While humans quickly understand semantic differences, visual similarities can pose challenges to the model, leading to lower accuracy compared to the baseline. The "Pragmatics" tag is for images that require non-literal interpretation, such as understanding idiomatic language or clarifying syntactic ambiguities in captions. Challenges such as discerning metaphors or the specific semantic relationships in prepositional phrases are reflected in the lower accuracy rates for both humans and models compared to the baseline.

In differentiating attributes, objective ones based on visual facts, such as color and shape, are easier to discern. In contrast, more abstract attributes, like size or amount, which require a reference, or weight, which necessitates additional knowledge, are more challenging to identify.

The "Determiner Numeral" category, which usually involves counting, is handled competently. Object-centric spatial tasks require identifying the correct reference frame and understanding the spatial relationships from an object’s perspective in the image, which demands an interpretation of the 3D world from a 2D representation; this remains challenging and is generally more difficult than discerning camera-view spatial relationships. Temporal dynamics, which involve inferring past, present, and future events, are still difficult for the model.

We present the error analysis in Figure 3 for all experimental settings shown in Table 2. The results consistently show that the categories tagged with ’Series, Pragmatics, Size/Amount, Weight, Object-Centric Spatial, Temporal’ are comparatively more difficult. Error examples are presented in Figure 6.

Conclusion

In this study, we introduce a “description then decision" strategy for vision-language tasks. From a neuroscience perspective, humans conduct recognition and reasoning in distinct modules and through multiple steps. From a model training perspective, large language models are proficiently trained on linguistic tasks, and vision encoders have increasingly been aligned with these language models through image captioning. Given a vision-language task, the “description then decision" approach transforms the task into two well-trained tasks. Although straightforward, our prompt strategy has demonstrated consistent improvements across various models, paving the way for future research into reasoning paradigms for vision-language tasks.

References