MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines
Dongzhi Jiang, Renrui Zhang, Ziyu Guo, Yanmin Wu, Jiayi Lei, Pengshuo Qiu, Pan Lu, Zehui Chen, Chaoyou Fu, Guanglu Song, Peng Gao, Yu Liu, Chunyuan Li, Hongsheng Li
Introduction
Search engines (Brin & Page, 1998) have been the main tools for humans to navigate through the overwhelming quantity of online resources. Recently, Large Language Models (LLMs) (OpenAI, 2023a; b, Touvron et al., 2023a) have demonstrated impressive performance on various zero-shot downstream applications. On top of this, AI search engine (OpenAI, 2024c), which integrates LLMs with traditional search engines, stands among one of the most promising ones. It points the direction of the next-generation interaction paradigm of human and Internet. Combining the language understanding ability of LLMs and up-to-date information from the Internet, AI search engines could better grasp the user’s intention and summarize contextual-aligned answers from the raw web information. These systems can only process textual queries and interpret textual web content, significantly constraining user query scenarios and information-seeking methods (Barbany et al., 2024, Xie et al., 2024). This limitation impacts both the range of input queries and the accuracy of results (Jiang et al., 2024a, Chen et al., 2021, Lù et al., 2024), particularly given the complexity and interleaved nature of modern websites (Liu et al., 2024b). For example, consider a scenario where you possess numerous medals belonging to your grandfather but are unaware of their specific names. A multimodal AI search engine could match photographs of these medals with an interleaved table of images and text retrieved from the Internet, thereby identifying each medal. In contrast, text-only search engines can neither take photographs for searching nor understand the interleaved table. Hence, a multimodal AI search engine is crucial for advancing information retrieval and analysis.
On the other hand, with the recent rapid advancements, Large Multimodal Models (LMMs) (Liu et al., 2023a, Lin et al., 2023, OpenAI, 2023c, Gao et al., 2024, Zhang et al., 2024b) have showcased significant abilities across diverse scenarios, including general image understanding (Fu et al., 2023, Liu et al., 2023b, Yu et al., 2023), expert image reasoning (Zhang et al., 2024d, Gao et al., 2023a, Zhang et al., 2024c, Guo et al., 2024a), multi-image perception (Li et al., 2024a, Wang et al., 2024, Jiang et al., 2024b, Li et al., 2024c), and spatial environment perception (Guo et al., 2023, Yang et al., 2023, Han et al., 2023). Despite these developments, a framework for LMMs to function as multimodal AI search engines remains largely unexplored. Consequently, the potential of LMMs in multimodal searching also remains a significant open question.
To bridge this gap, we first present MMSearch-Engine, a multimodal AI search engine pipeline, empowering any LMMs with advanced search capabilities. MMSearch-Engine maximizes the utilization of LMMs’ multimodal information comprehension abilities, incorporating both visual and textual website content as information sources. On top of this, we introduce \dataset, a multimodal AI search engine benchmark to comprehensively evaluate LMMs’ searching performance. The design of MMSearch-Engine facilitates the zero-shot evaluation of any LMMs within the context of AI search engine. Our experiment covers state-of-the-art closed-source (OpenAI, 2023c, Anthropic, 2024, Gemini Team, 2023) and open-source LMMs (Li et al., 2024b, Qwen Team, 2024, Chen et al., 2024d, Ye et al., 2024). Our efforts are summarized as follows:
MMSearch-Engine, a multimodal AI search engine pipeline for LMMs, empowering large models for multimodal searching. In contrast with the conventional text-only AI search engines, MMSearch-Engine fully integrates multimodal information in two ways: (i) for queries containing images, we conduct web searches across both textual and visual modalities. We utilize Google Lens (len, ) to identify critical visual information from the input image; (ii) all search results are presented in both textual and visual formats, ensuring a comprehensive understanding of the interleaved website content. The working flow of MMSearch-Engine contains multi-round interaction between LMM and the Internet. The LMM needs to first requery the user question into a search-engine-friendly format. Then, the LMM reranks the retrieved websites based on its helpfulness. Finally, the LMM is required to summarize the answer based on the most informative webpage content selected from the rerank. Thanks to the design of the pipeline, we propose a step-wise evaluation strategy on the three core tasks within the searching process: requery, rerank, and summarization. The final score is weighted by the end-to-end evaluation results and scores of the three core tasks.
, a comprehensive benchmark for multimodal AI search engines, which, to our best knowledge, serves as the first evaluation dataset to measure LMMs’ multimodal searching capabilities. Our benchmark categorizes searching queries into two primary areas: News and Knowledge, as shown in Fig. 1. We employed different strategies for these two areas to ensure the challenging nature of the benchmark. News area covers the latest news at the time of data collection (August, 2024). This is to guarantee the answers to the queries will not be present in the training data of LMMs. As for the area of Knowledge, we collect queries requiring rare knowledge and then select the queries unable to be answered by current SoTA LMMs such as GPT-4o (OpenAI, 2024b) or Claude-3.5 (Anthropic, 2024). The two areas sum up to 14 subfields. In total, \dataset encompasses 300 meticulously collected queries, with 2901 unique images.
Extensive experiments and error analysis for future development direction. We evaluate popular closed-source models and open-source LMMs on \dataset. GPT-4o achieves the best overall performance across different tasks. Surprisingly, our MMSearch-Engine equipeed with SoTA LMMs, such as GPT-4o and Claude 3.5 Sonnet, even surpasses the prominent commercial AI search engine Perplexity Pro (Perplexity, ) in the end-to-end task. Our thorough error analysis reveals that current LMMs still struggle to generalize to multimodal search-specific tasks. Their poor requery and rerank capabilities significantly limit their ability to correctly identify useful websites and extract relevant answers. Additionally, we identify five error types for requery and summarization tasks, respectively. We find that current LMMs cannot fully understand the requery task and do not know how to query the search engine. As for the summarization task, LMMs often have difficulty in extracting useful information, either from text or images. These capabilities are essential for LMMs to function as robust multimodal search engines and require further development. We also conduct a preliminary ablation study to explore the potential of scaling test-time computation versus scaling model size (OpenAI, 2024a). Initial results indicate that scaling test-time computation demonstrates superior performance in this task.
\dataset
In Section 2.1, we first detail the design of our multimodal AI search engine pipeline, which serves as both data collection and evaluation tools. Then, in Section 2.2, we detail the data composition and collection of the curated multimodal search benchmark \dataset. Then, in Section 2.3, we elaborate on our step-wise evaluation strategy. Finally, we detail the dynamic nature of our benchmark in Section 2.4.
The searching process is a complex action including multi-round interactions between LMMs and conventional search engines. We develop a delicate pipeline that queries LMMs multiple times to accomplish this task. Leveraging the image comprehension capabilities of LMMs, we incorporate two types of visual data. First, we incorporate Google Lens (len, ) to search for information from the image. The second type of visual data is the screenshot of the retrieved websites, in the purpose of preserving the original format of website content. Our framework is shown in Fig. 2. Below we detail how an LMM works with this pipeline, which comprises three sequential phases:
Requery. The query direct from users may contain references to certain information in the image, e.g., the News-Finance example shown in Fig. 1. Since a conventional search engine only accept text-only input, it is necessary for LMM to translate the image content and combine it with the query to ask a valid question to it. In addition, the raw user query may be ambiguous or inefficient sometimes (Chan et al., 2024, Ma et al., 2023), reformulating the query to be more clear is also a must for LMM. If the user query contains an image, we incorporate the screenshot of the image search result from the google lens (len, ). We treat the user query, user image, and the image search screenshot as basic information of the query. This information will be input to LMM in every round in the pipeline. For the requery round, we prompt LMM to output a requery to a conventional search engine.
Rerank. The requery is sent to a search engine API, e.g., DuckDuckGo, to retrieve top relevant websites. Depending on the requery quality, not all retrieved websites are necessarily relevant for query answering. Hence, we prompt LMM to select one most informative website for answer summarization. Due to the LMM’s context length limitations and the extensive content of websites, we provide only essential information of each website, which we term brief results. These brief results include the title, the snippet, and a screenshot of the webpage’s top section, which serves as the input for LMM’s reranking. The inclusion of the screenshot serves two purposes. First, the screenshot offers a visual cue to assess the web’s credibility, as a well-organized website often appears more trustworthy than one cluttered with advertisements (Fogg et al., 2001, Sillence et al., 2004). Additionally, the screenshot may contain essential visual information. For instance, it might include images similar or identical to query images, as shown in the Website 2 in Fig. 2.
Summarization. We start by crawling the selected website to gather all the available information. We parse the HTML to obtain the raw textual content and capture a full-page screenshot of the website. However, there are two issues: the raw content tends to be extensively lengthy and disorganized, while substantial areas in the full-page screenshot are blank due to the ad blocks on the website. These two issues lead to a large number of input tokens filled with irrelevant information. To enhance data efficiency, we slim the screenshot and retrieve the relevant content before inputting them to LMM. For the full-page screenshot, we identify the blank areas and remove them iteratively, detailed in Appendix B. As for the text content, we apply a text embedding model (Chen et al., 2024a) to retrieve a maximum of 2K tokens relevant to the requery from the raw content. We define the slimmed screenshot and the retrieved content as full website content. Finally, we input the full website content, website title, and website snippet, along with the query information, to LMM for summarizing the answer.
2 Data Composition and Collection
To thoroughly assess multimodal search proficiency, we compile a comprehensive problem set covering a broad spectrum of news topics, specialized knowledge domains, and query image patterns. This widespread collection for \dataset aims to simulate diverse user searching scenarios, ensuring a robust evaluation of LMMs’ capabilities in multimodal search.
Data Composition and Categorization. Our benchmark aims to isolate LMMs’ inherent knowledge and assess their actual search capabilities. We focus on two primary areas: News and Knowledge. For the News area, the queries are related to the latest news at the time of data collection (August 2024). This guarantees no overlap between the current LMMs’ training data and questions in our benchmark. All questions in this area are recorded with their occurrence time. For fairness, LMMs with recently updated knowledge should be tested on queries that occurred after their latest data update. Due to its time-sensitive nature, the News area serves as a dynamic part of our benchmark. Please refer to Section 2.4 for details. As for the Knowledege area, we focus on rare knowledge in targeted domains. Each question proposed by an annotator is verified to be beyond the capabilities of state-of-the-art Large Language Models (LLMs) such as GPT-4o (OpenAI, 2024b) or Claude 3.5 Sonnet (Anthropic, 2024). The Knowledge area serves as a static component of our benchmark and remains constant over time. We collect a total of 300 queries across the 2 primary areas and 14 subfields. Detailed statistics for data composition and categorization are presented in Table LABEL:table:statistics and Fig. LABEL:fig:pie. Definitions of each subfield are in Appendix C.1.
Data Collection and Review Process. Thanks to the design of our pipeline, the data collection process follows similar procedure introduced in the pipeline. An annotator is first required to propose a query and provide its answer, either sourced from the latest news or rare knowledge. The annotator then formulates a requery based on the query information. After websites are retrieved from the search engine, the annotator is required to divide all websites into three sets based on the brief results: valid (likely to contain the answer), unsure (relevance is difficult to determine), and invalid (entirely irrelevant to the question). We mandate that at least one website must be classified as valid; if this criterion is not met, the annotator is required to adjust the requery to obtain new search results. Finally, we randomly pick one website from the valid set and obtain its full content. To ensure the question is answerable, another annotator is employed to give an answer to the query based on the full content. If the answer is incorrect, the question needs to be revised until it is answerable.
3 Evaluation Protocol
In contrast with previous LMM benchmarks, the multimodal search process of LMM contains multiple rounds. Only the end-to-end evaluation of the final answer is inadequate to reveal the models’ deficiency in each core searching step. For example, the errors made by the model may occur during the summarization process, but it might also stem from choosing an incorrect website during the reranking stage. To this end, we propose a step-wise strategy to evaluate the LMMs’ capability on the three core searching steps, in addition to the end-to-end evaluation.
End-to-end score (): We compute the F1 score between the predicted answer and the ground truth to judge if the answer is correct.
Requery score (): We apply the average of ROUGE-L and BLEU-1 scores to measure the similarity between the model’s requery and human-annotated requery.
Rerank score (): The rerank score is derived from the LMM’s selection among pre-defined websites. The score values is 1.0 for valid set, 0.5 for unsure set, and 0 for invalid set or incorrect format.
Summarization score (): Again, we compute the F1 score of LMM’s answer based on a pre-defined website content against ground truth.
The input, output, and ground truth of the four tasks are visualized in Fig. 4. The final score is weighted by these four scores. We assign the highest weight (75%) to the end-to-end task, as it reflects the real-world multimodal search capability. The remaining 25% is distributed among the intermediate steps: 10% each for the rerank and summarization tasks, and 5% for the requery task. The lower weight for the requery task accounts for the inherent uncertainty in this process. The scoring process can be formulated as:
4 Benchmark Evolution
In Fig. 5, we showcase the statistics of data timestamp distribution in the News area. Our dataset spans from 1st May 2024 to 31th August 2024. By the time of evaluation, we inspect the knowledge cutoff dates of the closed-source models. Claude 3.5 Sonnet reports a knowledge cutoff of April 2024, while both GPT-4V and GPT-4o state they lack information from 2024. For open-source models, we examine their release dates and training data, confirming that none possess knowledge beyond May 2024. This temporal gap ensures the fairness of our evaluation, as the models’ performance solely reflects their multimodal search capabilities rather than pre-existing knowledge. We will update the News area if a new LMM’s training data may overlap with our collection period.
Experiment
In this section, we conduct a systematic evaluation of existing LMMs on \dataset. We first introduce the experimental setup in Section 3.1. Then, we detail the quantitative results in Section 3.2 and narrate the error analysis in Section 3.3. Finally, we explore scaling test-time compute versus scaling model size in Section 3.4.
Evaluation Models We examine the performance of foundation models across three distinct categories on \dataset: (a) Commercial AI Search Engines, represented by Perplexity (Perplexity, ). We test the pro version of Perplexity, which takes only the user query and image as input. Since SearchGPT (OpenAI, 2024c) has not been public yet, we do not test on it. (b) Closed-source LMMs, represented by models like GPT-4V (OpenAI, 2023c), GPT-4o (OpenAI, 2024b), and Claude 3.5 Sonnet (Anthropic, 2024), and (c) Open-source LMMs, featuring models such as LLaVA-OneVision-7B (Li et al., 2024b) (Qwen2-7B (Yang et al., 2024a)), LLaVA-OneVision-72B (Li et al., 2024b) (Qwen2-72B (Yang et al., 2024a)), LLaVA-NeXT-Interleave (Li et al., 2024c) (Qwen1.5-7B (Yang et al., 2024a)), InternVL2 (Chen et al., 2024d) (InternLM2.5-7B-Chat (Cai et al., 2024)), InternLM-XC2.5 (Zhang et al., 2024a) (InternLM2-7B (Cai et al., 2024)), Qwen2-VL-7B (Qwen Team, 2024) (Qwen2-7B (Yang et al., 2024a)), Qwen2-VL-72B (Qwen Team, 2024) (Qwen2-72B (Yang et al., 2024a)), mPlug-Owl3 (Ye et al., 2024) (Qwen2-7B (Yang et al., 2024a)), Idefics3 (Laurençon et al., 2024) (LLaMA3.1-7B-Instruct (AI@Meta, 2024)), and Mantis (Jiang et al., 2024b) (LLaMA3-7B (AI@Meta, 2024)). Note that the open-source LMMs’ sizes are 7B unless otherwise specified.
We set the number of retrieved websites as 8. All our experiments of open-source models are conducted without any fine-tuning on search data or tasks. As for the prompts, the requery prompt contains 3 examples to better guide LMMs to output a valid requery. While prompts for other tasks are all in a zero-shot setting. We prompt the LMM to output as few words as possible for a better match with the ground truth. We employ the metric introduced in Section 2.3. Besides, we recruit eight qualified college students and ask them to solve the problems in \datasetindependently, following the same pipeline of MMSearch-Engine. This score serves as a baseline for human performance. We conduct all experiments on NVIDIA A100 GPUs.
The input image dimensions for the webpage’s top section screenshot were set to pixels. For the full-page screenshot, we set the initial webpage width to pixels, although the actual width of a small portion of webpages may vary due to its layout settings. Furthermore, considering that a full-page screenshot can be extremely lengthy, directly inputting it as a single image into an LLM would result in excessive downsizing, making the content too vague for accurate identification. To address this, we segmented the full-page screenshot into multiple images, starting from the top, with each segment measuring pixels in height. Because of the context length limitations of LMMs, the maximum number of full-page screenshot segments is therefore restricted to ten.
For the default settings, the longest edge of the input image is resized to match the largest resolution of the vision encoder of LMM. This ensures the image not to be cropped to multiple images and will only take up the minimum of tokens for image input. For any resolution settings, we input the image without resizing.
2 Experimental Analysis
To thoroughly investigate the multimodal searching capabilities, we present the evaluation results of different models on \dataset following the proposed step-wise evaluation strategy in Table 2 and fourteen subfields in Table 3. We now provide a detailed discussion of notable findings and their implications for multimodal search capabilities.
Any-resolution input only provides slight or no improvement. Of the tested LMMs, four models, which are InternLM-XC2.5, InternVL2, mPlug-Owl3, and Idefic3, all support both low-resolution (LowRes) and any-resolution input (AnyRes). As one would expect, AnyRes input enables better OCR and perception of the image. However, we only observe slight or even no enhancement comparing the difference between the LowRes performance and its AnyRes counterpart. Take mPlug-Owl3 as an example, AnyRes input surpasses LowRes input on overall score by 1.8%, end-to-end score by 2.7%, and rerank on 0.2%. While it falls behind LowRes on requery by 0.8% and summarization by 1.7%. This suggests that the OCR and perception quality do not bottleneck the search performance. Rather, the suboptimal performance appears to stem from the LMMs’ inherent lack of robust search capabilities.
Current LMMs still have significant shortcomings in requery and rerank. Comparing the average score of the end-to-end task with that of the summarization task, we find that the summarization score consistently surpasses the end-to-end task by a large margin, both in the closed-source and open-source models. The minimum margin is 2.7% for GPT-4o, while the maximum is 23.9% for LLaVA-OneVision-7B. This discrepancy can be attributed to the differences in the tasks’ input quality. While the summarization task input always contains the answer, the end-to-end task’s third-round input quality depends on the model’s requery and rerank quality in previous rounds. The magnitude of this performance gap reflects the disparity between a model’s summarization ability and its capacity for requery and rerank tasks. The larger the difference, the larger the capability gap. Observing the result, we find that this gap of most open-sourced models exceeds 14%, while the closed-sourced models are all below 10%. This suggests all current LMMs needs improvement of their requery and rerank ability, especially for open-source models. Mantis is one exception of open-source models with a margin of only 3.4%. This means its poor summarization capability bottlenecks its end-to-end performance. Qwen2-VL-72B’s 10.5% gap, also falling below 14%, highlights its superiority among other open-source LMMs.
Closed-source LMMs are better-performed than open-sourced LMMs on overall performance. For the final score, closed-source LMMs consistently outperform the open-source LMMs. GPT-4o achieves the highest overall score of 62.3%, demonstrating superior zero-shot multimodal search capabilities. While Qwen2-VL-72B leads among open-source models, it still lags behind GPT-4o by 9.6%. The performance gap widens to 11.3% on the most challenging end-to-end task and further expands to 20.1% for 7B open-source LMMs. These significant disparities highlight substantial room for improvement in open-source models.
SoTA LMMs with our MMSearch-Engine surpass commercial AI search engines in the end-to-end task. We also evaluate the pro version of Perplexity (Perplexity, ), a prominent commercial AI search engine that accepts both image and text queries, on our dataset. Perplexity pro can accept both image and text in the user query. Surprisingly, although Perplexity also leverages SoTA LMMs like GPT-4o and Claude 3.5 Sonnet, it largely underperforms MMSearch-Engine equipped with the same model in the end-to-end task. Even more remarkably, MMSearch-Engine can even surpass Perplexity with Qwen2-VL-72B, an open-source LMM. This suggests that our MMSearch-Engine provides a better open-source plan for multimodal AI search engine. The performance gap validates MMSearch-Engine’s design effectiveness and highlights the value of testing various LMMs within our pipeline, since the pipeline can indeed achieve remarkable performance when using powerful LMMs. Upon investigating Perplexity’s sub-optimal performance, we discovered that it appears to utilize only a rudimentary image search algorithm, if any. This limitation leads to poor identification of the key objects in the image and failure to retrieve relevant information. Our findings underscore the effectiveness of MMSearch-Engine’s design, particularly the incorporation of a robust image search step, which plays a crucial role in accurately recognizing important information from the input image.
3 Error Analysis
To investigate the limitations of current LMM search capabilities, we conducted a comprehensive analysis of error types observed in our evaluation. Our proposed step-wise evaluation strategy enables analysis of failure modes for each core search step, complementing the end-to-end assessment. This analysis encompasses the entire benchmark. We first examine the end-to-end error types for both the best-performing closed-source model (GPT-4o) and open-source model (Qwen2-VL-7B). To better understand the failure cases, we then identify distinct error types in the requery and summarization task, which requires open-ended generation. We quantify these error types for a systematic understanding of current LMM limitations and point out critical areas for improvement.
In this section, we are trying to answer the question: Which step does LMM make a mistake in the end-to-end evaluation? In Fig. 8, we showcase the statistics of different error types occurring in GPT-4o and Qwen2-VL-7B. We define the following four error categories: (i) requery, where the model requery is incorrect, and leads to all retrieved websites being invalid; (ii) rerank, where the model selects a website without a correct answer; (iii) summarization, where the full website content contains the information of correct answer, but the model fails to extract it; (iv) informal, the output format deviates from the prompt specifications. As shown in the figure, GPT-4o’s primary error sources are rerank and summarization errors, while requery and informal errors account for approximately half the frequency of the main error causes. This suggests that GPT-4o’s limitations lie primarily in information source ranking and multimodal information integration. As for Qwen2-VL, all four error types occur with similar frequency. The rise of the informal error portion may be attributed to the model’s inferior instruction-following ability. Besides, it should be noted that the requery task demands advanced comprehension and key image information extraction ability. This task seldom appears in the training data of current LMMs. The prevalence of this error type in Qwen2-VL may indicate that it fails to generalize to adequately address this complex task.
3.2 Error Analysis of Reuqery and Summarization Task
To better understand how open-source LMM makes the mistake, we dive into the requery and summarization task to find out the error patterns of Qwen2-VL-7B. We particularly select the two tasks requiring open-ended generation, which provides more information to identify the error.
As for the requery task, we categorize five types of errors:
Lacking Specifility, where the model fails to include all the specific information in the requery and therefore leads to sub-optimal search results. For example, the query is asking the release date of Vision Pro in China. However, the model omits the condition of China and directly asks about the release date of Vision Pro.
Inefficient Query, where the model does not consider the real scenario and the requery is inefficient for the search engine to find the answer. For example, the query is asking whether the Van Gogh’s Sunflowers and Antoni Clavé’s Grand Collage are both oil paintings. Clearly, it is a commonsense that Van Gogh’s Sunflowers is an oil painting and Antoni Clavé’s Grand Collage is much less well-known. An efficient query should be asking about the images of Antoni Clavé’s Grand Collage and further determine if it is also an oil painting by directly looking at it. However, the model directly asks the original query to the search engine. There is very little chance that an exact same question has ever been raised so probably this requery will bring very little helpful information.
Excluding Image Search Results, where the model totally ignores the information in the screenshot of the image search results and therefore lacks important specific information in the requery. For example, the query is ‘When did this football player obtain the gold medal?’ and provides an image of the player. The model is supposed to find out the player’s name by viewing the image search result and raise a requery like ‘[PLAYER NAME] obtained the gold medal time’. However, the model fails to incorporate the player’s name in the requery and definitely the retrieved websites will not include any helpful information.
No Change, where the model just uses the question as the query input to the search engine.
Irrelevant, where the model either matches wrong information from the image search result or mistakenly understands the query and outputs an irrelevant requery.
These error types of requery suggest that LMM often fails to fully understand the requery task and fails to aggregate all available information. Besides, the error type of inefficient query indicates that LMM has no clue of the real working scenario and query principles of search engines.
As for the summarization task, we also identify five types of errors:
Text Reasoning Error, where the model fails to extract the answer from the website textual information.
Image-text Aggregation Error, where obtaining the answer needs combining the information from both images and texts. The model fails to do so.
Image reasoning Error, where the model fails to extract the answer from the image, and the answer can only be obtained from the image.
Hallucination (Huang et al., 2023), where the model provides an unfaithful answer that cannot be grounded in the given content.
Informal, the output format does not follow the prompt specifications, the same error type in the end-to-end task.
The occurrence of the five types of summarization errors reflects that current LMMs still cannot correctly extract the given multimodal information to answer the query. The ability of content understanding still requires further enhancement for current LMMs.
4 Scaling Test-Time Compute vs Scaling Model Size
Recent works such as OpenAI o1 (OpenAI, 2024a) and Li et al. (2024d) have highlighted the critical role of scaling test-time computation in enhancing model performance. Our end-to-end task, which requires multiple Internet interactions, presents an opportunity to investigate the potential of scaling test-time computation compared to scaling model size. To explore this, we conduct experiments using LLaVA-OneVision-7B (Li et al., 2024b), focusing on scaling test-time computation, and compare against LLaVA-OneVision-72B scaling in model size, which aims to provide insights into the relative benefits of increased inference computation versus increased model parameters.
For scaling up the test-time computation, we adopt a multi-modal search strategy similar to best-of-N solution, where ‘N’ denotes 25 in our settings. Specifically, for LLaVA-OneVision-7B, we first prompt the model to generate a requery 5 times, from which we selected the one with the highest requery score . This requery is then used to retrieve brief results from 8 websites from a search engine. The model is again prompted 5 times to select the most informative website. After removing duplicates from the selected websites, we extract the full website content from the remaining ones and prompt the model to answer 5 times, obtaining 25 end-to-end outputs in total. We compute the F1 score for each answer against the ground truth and take the maximum as the model’s end-to-end score for the query. Table 4 shows that LLaVA-OneVision-7B (TTC) achieves the score of 55.2% in the end-to-end task, significantly enhancing the original score of 29.6%, which surpasses LLaVA-OneVision-72B’s 44.9% and GPT-4V’s 52.1%. This result reveals the substantial potential of scaling test-time computation, validating the effectiveness of this technique as introduced by OpenAI o1. Our findings provide valuable insights for future research in this domain, suggesting that increased inference computation may offer comparable or superior performance improvements to increased model size not only in math and code tasks, but also in multimodal search tasks.
Conclusion
In this paper, we investigate the potential of LMMs as multimodal AI search engines. We first design MMSearch-Engine, a streamlined pipeline, enabling zero-shot LMMs to perform multimodal searches. To comprehensively assess the search capabilities, we introduce \dataset, a benchmark comprising 300 queries across 14 subfields. Our evaluation methodology analyzes LMM search abilities step-by-step, facilitating a deeper understanding of their limitations. Using MMSearch-Engine, we evaluate various closed-source and open-source LMMs, revealing that current models still fall short of human-level search proficiency. Through thorough error analysis, we identify specific patterns of failure in key search process steps, providing valuable insights for future improvements in LMM search ability.
References
Appendix Overview
Section B: Additional experimental details.
Appendix A Related work
Recently, multimodal models (Radford et al., 2021, Li et al., 2022, OpenAI, 2023c, Rombach et al., 2022, Jiang et al., 2024c) has gained unparalleled attention. Building on the success of Large Language Models (LLMs) (Touvron et al., 2023a; b) and large-scale vision models (Radford et al., 2021), Large Multimodal Models (LMMs) are gaining prominence across diverse domains. These models extend LLMs to handle tasks involving various modalities, including mainstream 2D image processing (Liu et al., 2023a, Zhu et al., 2023, Lin et al., 2023, Gao et al., 2023b), as well as 3D point clouds (Xu et al., 2023, Guo et al., 2023; 2024b), and videos (Li et al., 2023, Chen et al., 2023a, Zhang et al., 2023, Fu et al., 2024). Among these LMMs, OpenAI’s GPT-4o (OpenAI, 2024b) and Anthropic’s Claude 3.5 Sonnet (Anthropic, 2024) demonstrate outstanding visual reasoning and comprehension capability, setting new standards in multi-modal performance. However, their closed-source nature limits broader adoption and development. In contrast, another research trajectory focuses on open-source LMMs for the community. Pioneering works like LLaVA (Liu et al., 2023a; 2024a, Li et al., 2024c; b), LLaMA-Adapter (Zhang et al., 2024b, Gao et al., 2023b), and MiniGPT-4 (Zhu et al., 2023, Chen et al., 2023b) incorporate a frozen CLIP (Radford et al., 2021) model for image encoding and integrate visual information into LLM for multi-modal instruction tuning. Later, works such as mPLUG-Owl (Ye et al., 2023a; b; 2024), SPHINX (Gao et al., 2024, Lin et al., 2023), and InternLM-XComposer (Dong et al., 2024) further advanced the field by incorporating diverse visual instruction tuning data and generalizing to more scenarios. More recent developments in the field have taken diverse directions. For example, several studies (Zong et al., 2024, Tong et al., 2024) explore multiple vision encoders design. Meanwhile, other works (Liu et al., 2024a, Chen et al., 2024d, Qwen Team, 2024) incorporate high-resolution image input. Multi-image instruction data (Li et al., 2024c, Jiang et al., 2024b) is also integrated to enable perception across multiple images. While various benchmarks, both in the general (Fu et al., 2023, Liu et al., 2023b, Yu et al., 2023) and expert (Zhang et al., 2024c, Lu et al., 2023; 2022) domain, has been proposed, the potential of LMM to function as a multimodal search engine remains largely unexplored. To this end, we introduce the \datasetbenchmark, which evaluates LMMs’ zero-shot abilities of multimodal search, offering valuable insights for future research.
RAG (Retrieval-Augmented Generation) is an effective strategy for enhancing model knowledge by retrieving relevant information from external sources (Fan et al., 2024). RAG has been leveraged in various scenarios including knowledge-intensive question answering (Borgeaud et al., 2022, Guu et al., 2020), machine translation (He et al., 2021), and hallucination elimination (Béchard & Ayala, 2024). Current works has focused on improving specific aspects of RAG. RG-RAG (Chan et al., 2024) proposes to refine the query for retrieval by decomposition and disambiguation. Self-RAG (Asai et al., 2023) incorporates the self-reflection of LLM to enhance the generation quality. The AI search engine could be viewed as a form of RAG with the Internet serving as the external knowledge source. Recently, MindSearch (Chen et al., 2024c) proposes an AI search engine framework to simulate the human minds in web information seeking. Meanwhile, multiple benchmarks of RAG (Yang et al., 2024b, Chen et al., 2024b) have been introduced to comprehensively evaluate a RAG system. However, both the current AI search engine and RAG benchmark are limited to the text-only setting, leaving the multimodal search engine and evaluation largely unexplored. To bridge this gap, we introduce MMSearch-Engine and \dataset, a multimodal AI search engine pipeline and dataset designed to evaluate various multimodal scenarios.
Appendix B Additional experimental details
Model Sources. For different LMMs, we select their latest models with size around 7B for evaluation to fully reveal their multimodal search proficiency. Table 5 presents the release time and model sources of LMMs used in \dataset.
Full-page Screenshot Slimming. For the full-page screenshot, we compute the Sobel gradients (Kanopoulos et al., 1988) to detect the edges and generate a gradient magnitude image. We iteratively remove the areas with gradients below a threshold, which represent the blank areas. This approach effectively reduces image size while maintaining critical document content.
Input Prompts of LMM for Response Generation. We showcase the input prompts of LMM for the three tasks respectively in Table 6-8. We adopt two types of prompts for queries with an image and without images. For query with an image, we specifically require the LMM to leverage the image search result to solve the task.
Appendix C More data details
News area encompasses a vast spectrum of information, ranging from everyday events to engaging entertainment content and specialized fields such as scientific discoveries and financial analysis. This comprehensive coverage serves as a rigorous assessment of the model’s ability to process information in diverse domains. We divide this expansive area into eight distinct subfields:
Traditional Sports: Data concerning traditional athletic competitions, team performances, player statistics, and sporting events. This includes scores, league standings, player transfers, and analysis of various professional sports across different leagues and countries.
e-Sports: Information about competitive video gaming, including tournament results, player rankings, and league information. This covers various game titles, team formations, streaming viewership statistics, and tournament information.
Technology: Information about technological innovations, gadgets, software developments, and tech industry news. This includes product launches, software updates, cybersecurity issues, and artificial intelligence advancements.
Paper: Content related to academic papers, research publications, and scholarly articles in various artificial intelligence fields. The queries include method explanation, figure understanding, and experiment settings.
Entertainment: Data about movies, music, television, celebrities, and other forms of popular entertainment. It also includes data concerning video games.
Finance: Information on financial markets, economic indicators, business news, and monetary policies. This covers stock prices, company earnings reports, company financial statements, and regulatory news regarding finance.
General News: Broad coverage of various news topics not specific to any particular subfield. This includes a mix of local and global events, human interest stories, lifestyle articles, climate news, and general interest content that doesn’t fit neatly into other specialized news subfields.
False Premise: Data related to misinformation or incorrect assumptions in the query. This subfield focuses on fact-checking capabilities. All the answers to the queries of this subfield are ‘invalid question’.
Knowledge area represents broad subfields of information and data related to general knowledge across various disciplines. This area concentrates on rare knowledge that most LMMs fail to answer. We categorize this area into five subfields:
Architecture: Information about building design, architectural styles, building information, and construction projects. This includes city landmarks, the comparison of architectural styles, and multi-view architecture matchings.
Arts: Data concerning visual arts, drawings, sculptures, badges, and other forms of creative expression. This covers artwork details, artist profiles, artwork history, and artwork style comparisons.
Fashion: Content related to clothing trends, fashion brands, and designer collections. This includes retail price, clothing style, release date, and brand information.
Astronomy: Information about celestial objects, space exploration, astronomical phenomena, and related research. This covers observational data from telescopes and image results from space missions. The questions focus on the background information of these celestial objects presented in the query image.
Anime: Data about Japanese animation, including series storylines and character information. This encompasses character background, character appearance, voice actor information, and chapter information.
Auto: Content related to automobiles, including vehicle specifications, industry trends, and automotive technology. This covers new car models, performance test results, coefficients of cars, and release date.