GPT-4V(ision) is a Generalist Web Agent, if Grounded
Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, Yu Su
Introduction
Large multimodal models (LMMs; ), especially recent ones such as GPT-4V(ision) and Gemini , have shown a remarkable capability on standard vision-and-language understanding and reasoning benchmarks . While web content has been a primary source of training data, a largely overlooked part of the web is the websites themselves—every website is designed to be rendered visually for easy consumption by human users. This poses a new challenge and a new opportunity for LMMs. On the one hand, screenshots of rendered websites, which could contain thousands of elements with rich relations, are more complex than the images in most existing benchmarks, which are usually object- or scene-centric. On the other hand, if LMMs can accurately comprehend websites, it will open the door for numerous applications on the web.
In this work, we aim to investigate the potential of LMMs as generalist web agents . A generalist web agent, as defined in Mind2Web , is expected to follow natural language instructions and complete tasks on any given real-world website (e.g., Figure 1). The tasks can be fairly diverse and complex, with one task possibly taking + actions across multiple dynamically rendered webpages. Existing work primarily uses large language models (LLMs) such as GPT-4 on the raw HTML input. However, HTML code is noisier than the rendered visuals and has a lower information density. For example, the screenshot in Figure 1 contains HTML elements that would require textual tokens with the GPT-2 Tokenizer, while requiring only visual tokens using GPT-4V’s visual tokenizer. Furthermore, HTML alone provides incomplete information and misses critical semantics from, e.g., embedded images.
To this end, we propose SeeAct, a generalist web agent that harnesses the power of LMMs for integrated visual understanding and acting on the web. We will focus on GPT-4V, the most advanced LMM publicly available to date, and compare it with smaller LMMs such as BLIP-2 . We find that GPT-4V exhibits a strong capability in visually understanding rendered webpages and generate the right plans in textual forms across a wide range of websites and tasks. However, grounding , i.e., converting the textual plan into precise actions on the website, remains a major challenge. It involves selecting the right HTML element to interact with as well as the right operation (e.g., Click, Type, or Select). We propose multiple grounding methods, including superpositioning bounding boxes and index labels onto the image, similar to set-of-mark prompting that has been shown effective on object- or scene-centric images. However, we find that on complex images with rich semantic and spatial relationships like webpage screenshots, severe hallucination is observed from GPT-4V. The most effective grounding strategy leverages the known correspondence between HTML element and their visual rendering, a unique property for websites compared to natural images.
We evaluate SeeAct on the Mind2Web dataset and compare it with text-only large language models (LLMs) like GPT-4 as well as smaller models (FLAN-T5 and BLIP-2 ) specifically fine-tuned for web agents. In addition to the standard offline evaluation setting on cached websites, we further establish an online evaluation setting by developing a new tool that allows for running web agents on live websites. The major findings from our exploration are summarized below:
SeeAct with GPT-4V is a strong generalist web agent, if oracle grounding is provided. In online evaluation, it can successfully complete % of tasks on different websites, substantially outperforming existing methods like GPT-4 (%) or FLAN-T5 (%). This strongly demonstrates the potential of LMMs like GPT-4V for web agents.
However, grounding is still a major challenge. The best grounding strategy still has a -% gap with oracle grounding. Among the various grounding strategies, the best one organically leverages both HTML text and visuals, substantially outperforming image annotation strategies by up to %.
In-context learning with large models (both LMMs and LLMs) show better generalization to unseen websites, while supervised fine-tuning still has an edge on websites seen during training.
There is a non-negligible discrepancy between online and offline evaluation because there can often be multiple viable plans for completing the same task. Online evaluation is more indicative of a model’s true performance.
SeeAct
In this section, we first explain the problem formulation of web agent and then introduce our developed SeeAct, a generalist web agent based on GPT-4V.
Specifically, given a web-based task (e.g., “Rent a truck with the lowest rate” in the car rental website), we examine two essential capabilities of GPT-4V as a generalist web agent: (i) Action Generation to produce an action description at each step (e.g., “Move the cursor over the ‘Find Your Truck’ button and perform a click”) towards completing the task, and (ii) Element Grounding to identify an HTML element (e.g., “[button] Find Your Truck”) at the current step on the webpage.
Given a website (e.g., a car rental website) and a task (e.g., “Rent a truck with the lowest rate”), the web agent should generate a sequence of executable actions to complete the task. Specifically, at time step , the agent should generate an action based on the current environment observation , the previous actions , and the task :
The environment observation comprises an HTML document and a screenshot image . LLMs can only be grounded on the HTML document, while LMMs can be grounded on both the HTML document and the screenshot image. The website status is updated accordingly after each action:
For simplicity, in subsequent step-wise formulations, the time step notation is omitted.
An action corresponds to a browser event provided by the website environment. Therefore, we formulate an action as a triplet of three necessary variables for a browser event . identifies the target webpage element to operate on, such as the "Find Your Truck" button in Figure 2.
represents the set of webpage elements within the environment . The operation is the action to be performed on the target element, with encompassing all possible operations in (e.g., Click, Type). The variable denotes the additional value needed for a certain operation (e.g., the date 12/10/2023 for a Type operation).
2 Action Generation
3 Action Grounding
To address this challenge, we explore three approaches using different types of information: Grounding via Element Attributes, Grounding via Textual Choices, and Grounding via Image Annotation, as depicted in Figure 2. The prompting details of action generation and grounding are included in Appendix A.
Grounding via Element Attributes. This approach involves prompting the model to generate as detailed attributes of the target element as possible, thereby providing more information to precisely match with the target HTML element.
Grounding via Textual Choices. The above approach demands precise and sufficient attribute descriptions from GPT-4V and accurate matching by the heuristic search, which can be highly demanding. For instance, many elements may have no textual content or have textual information in a nearby element instead of itself.
Grounding via Image Annotation. Textual representations alone are sometimes insufficient to distinguish similar or identical elements, as illustrated in Appendix E. Therefore, in this approach, we propose to overlay a bounding box for each candidate element selected by the ranker as well as an alphabet label on the bottom left of the bounding boxWe use the Supervision library for image annotation: https://supervision.roboflow.com/.The model is expected to generate the label corresponding to the target element.
Experiments
We evaluate our methods on Mind2Web , a comprehensive dataset encompassing over complex web tasks with annotated actions. This dataset spans websites across low-level domains, categorized into high-level domains. It supports three primary operations: Click, Type, and Select, with Hover and Press Enter operations integrated into Click to avoid ambiguity.
The dataset’s test sets aim to measure the generalization of web agents across different tasks, websites, and domains. Specifically, the Cross-Task setting focuses on evaluating agents on tasks that are new to the training data but within included domains and websites. The Cross-Website setting evaluates agents with tasks across new websites for each of the top-level domains in the training data. The Cross-Domain setting assesses agent performance on tasks in two top-level domains held out from the training data.
We align each HTML document in the dataset with its corresponding webpage screenshot image from the Mind2Web raw dump, which undergoes human verification to confirm correct rendering and element visibility. Given the constraints of the GPT-4V inference quota, our experiments focus on a random subset of 30 tasks from each test split. Detailed statistics of these subsets are presented in Table 1.
2 Methods
SeeAct. In grounding via image annotation and textual choices, we first employ the DeBERTa-base cross-encoder from MindAct to rank the top 50 elements for better comparison with its text-only counterparts. Then, we cluster elements into groups of options for inference. In grounding via element attributes, no candidate element is provided. We access GPT-4V through ChatGPT WebUI, which only allows a zero-shot setting.
To compare with SeeAct, we also implement methods based on text-only LLMs and BLIP2 following the two-stage strategy of MindAct . Firstly, we employ the ranker above to pick the top 50 elements. Subsequently, the action generation problem is formulated as a multi-choice question answering problem, with the candidate elements as options, including a "None" option if the target element is absent. During inference, elements are clustered into groups of 5 elements, with iterative refinement, until a single choice is made or all options are discarded. We evaluate supervised fine-tuning (SFT) methods using FLAN-T5 and BLIP2-T5 and in-context learning (ICL) methods using GPT-3.5, GPT-4.
FLAN-T5. We fine-tune FLAN-T5 using a left-to-right language modeling objective with the target sequence of ground-truth actions. The fine-tuned FLAN-T5 then serves as the backbone for inference, enabling action generation in the format defined in Equation 1.
BLIP2-T5. The BLIP-2 model combines a vision encoder and an LLM with a bridging component for modality connection. We jointly fine-tune the LLM and the bridge module on Mind2Web training data while keeping the vision encoder frozen. For the vision encoder, we leverage the ViT-L/14 pre-trained from CLIP with an image resolution of . To ensure a fair comparison with the FLAN-T5-based text-only model, we choose FLAN-T5 as the language model and initialize it with the parameters fine-tuned on Mind2Web.
GPT-3.5 and GPT-4. We also conduct experiments with text-only LLMs, specifically GPT3.5-turbo-0613 and GPT4-turbo-1106-preview, using in-context learning in 3-shot settings. We use the same multiple-choice formulation and include three demonstration examples for in-context learning as specified in MindAct.
3 Offline Evaluation
We adopt the evaluation metrics utilized in Mind2Web. Element Accuracy (Ele. Acc) compares the predicted element with the ground-truth elements. Operation F1 (Op. F1) calculates the token-level F1 score for the predicted operation comprised of action and input value. Step Success Rate (Step SR) measures the success of each action step. A step is successful only if the selected element and the predicted operation are correct. We report macro averages across tasks for these step-wise metrics. Success Rate (SR) measures the success of an entire task. A task is regarded successful only if all steps have succeeded. This metric is stringent without allowing the model any space for exploration and error correction. Therefore, for offline evaluation, we focus on the first three metrics. However, we also conduct online evaluation on live websites to get a better evaluation on the whole task success rate, as detailed below.
4 Online Evaluation
We develop a new online evaluation tool using Playwrighthttps://playwright.dev/python/docs/api/class-playwright to evaluate web agents on live websites (instead of cached websites in offline evaluation). Our tool can convert the predicted action into a browser event. To adhere to ethical standards, our experiments are restricted to non-login tasks in compliance with user agreements, and we closely monitor agent activities during online evaluation to prevent any actions that have potentially harmful impact, like placing an order or modifying the user profile.
For a fair comparison between offline and online evaluations, we only re-write time-sensitive tasks to ensure they are still valid when the evaluation is conducted. For instance, we update the dates for flight-related tasks. Finally, we conduct the online evaluation on a subset of total tasks from thethree test splits and report the mean performance across these tasks.
Results and Analysis
GPT-4V can be a Generalist Web Agent with Oracle Action Grounding. Given an effective action grounding method, GPT-4V has the potential to serve as a generalist web agent. Specifically, as described in subsection 2.3, we provide GPT-4V with an oracle action grounding method (SeeActOracle) through human annotation, the model achieves a step success rate of %, %, and % across three test splits, respectively. As shown in Table 2, this method substantially outperforms other models under all metrics across three test splits. Specifically, it achieves a % step success rate improvement over the second-best method in the Cross Task setting. The performance advantage is more pronounced under the Cross-Website and Cross-Domain settings, where it leads by 28.3% and 21.2% step success rates, demonstrating its generality compared with supervised fine-tuning. This observation is further corroborated with in the online evaluation (Table 3).
Element Grounding Method Comparison. However, there is a noticeable gap between oracle grounding and all the three proposed grounding methods. This demonstrates that grounding, especially element grounding, is a major bottleneck. Element grounding via textual choice (SeeActChoice) demonstrates the best performance under all metrics across all settings, comparable to supervised fine-tuning and showing a substantial improvement over text-only LLMs. Grounding via image annotation (SeeActAnnotation) offers an intuitive approach and shows promising results in recent work that focuses on object- or scene-centric images . However, we find that on complex images with rich semantic and spacial relationships like webpage screenshots, severe hallucination is observed from GPT-4V. Specifically, it often fails to correctly map its generated element description (which is often correct according to oracle grounding) to the right bounding boxe and index label in the image, leading to a low element accuracy. This limitation primarily arises from GPT-4V’s weakness in understanding image details and relative spatial location, a topic that we will further delve into in subsection 4.3.
Grounding via element attributes (SeeActAttribute) also demonstrates inferior performance. This method’s effectiveness is primarily limited by its heuristic-based element localization strategy, which depends on textual and locality characteristics. This becomes problematic as not all webpage elements contain text, and sometimes the relevant text is associated with a nearby but distinct element.
LMMs vs. LLMs. The SeeActChoices model demonstrates a substantial performance advantage over the text-only GPT-4 under all three metrics across all three test splits. Specifically, it outperforms GPT-4 in step success rate of 8.9%, 14.7%, and 11.2% on three settings, respectively. Interestingly, fine-tuned BLIP2-T5 does not show a noticeable gain over FLAN-T5, despite having additional visual input. Several factors may contribute to this. First, the CLIP model used as the image encoder may not be sufficiently adept at image details, as explored by Shen et al. . This limitation is particularly relevant for our web navigation task, which demands a high level of image detail comprehension. Second, BLIP2-T5 utilizes an off-the-shelf CLIP model that may not be optimal for webpage screenshots. Finally, although the screenshots in the test splits are error-free, some of the examples in the training set might contain issues such as rendering failures or inaccuracies in the timing of screenshot capture by annotators.
SFT vs. ICL. We compare SFT and ICL methods to offer insights for developing web agents in different scenarios. ICL (with SeeAct) demonstrate consistent and robust performance across three test splits. ICL is particularly advantageous in scenarios lacking annotations or requiring strong generalization capabilities for new domains and websites. As grounding methods improve towards oracle grounding, ICL is poised to show even stronger performance. On the other hand, SFT methods show better generalization across tasks on websites already seen during training. Considering the high cost of data annotation for web agents and the billions of websites on the internect, ICL offers a more compelling solution for generalist web agents. However, if one only needs to develop a strong web agent for a certain website, SFT is still a competitive solution.
2 Online Evaluation Results
In online evaluation, we pair a web agent with a human annotator, where the human was tasked to monitor agent actions that may change real-world states and determine whether each task was successfully completed. For comparative analysis, we include success rates from offline evaluation, denoted as Offline0 (allowing zero wrong action) and Offline1 (allowing one wrong action).
Table 3 shows that the whole task success rate in online evaluation substantially exceeds that of offline evaluation (Offline0). This finding suggests that the whole task success rate is likely underestimated in the offline evaluation due to the variability in actions and plans. In other words, there may be multiple viable plans for a task, but the reference plan in offline evaluation only captures one of them.
Across all three settings, SeeActChoice outperforms both GPT-4 and FLAN-T5-XL. Using oracle grounding further improves the performance substantially, reaching a remarkable whole task success rate of %. Although GPT-4 shows much worse performance than FLAN-T5-XL in step success rate under offline evaluation (Table 2), they exhibit similar success rates in online evaluation. These results further confirm the potential of large models for generalist web agents compared with fine-tuned small models.
Online Success Rate by Task Difficulty. We investigate the performance of web agents on tasks across different difficulty levels. We estimate the task difficulty based on the number of actions taken by annotators during action trace annotation. As shown in Figure 3, the whole task success rate is negatively correlated with the number of actions—it decreases as the number of actions increases across all four methods. SeeActOracle consistently outperforms other methods across all difficulty levels. Interestingly, the gap between SeeActOracle and SeeActChoice enlarges on longer-horizon tasks. This is understandable because grounding errors cascade to later steps; nonetheless, it further shows the challenge of grounding for GPT-4V and the need for better grounding methods.
3 Further Analysis and Discussion
Error Analysis in Grounding via Image Annotation. Set-of-mark prompting uses a similar method as grounding via image annotation and has been shown effective on object- or scene-centric images . However, this grounding method is suboptimal on webpage screenshot images that are complex and contain rich semantic and spatial relationships.
To analyze the reasons behind the failures, we randomly sample predicted actions and observe the major types of errors as: (1) Wrong action generation; (2) Making up bounding box & label; (3) Failure to link bounding boxes with the correct labels. Illustrative examples are included in Appendix C.
Our analysis reveals that % of the errors originate from wrong action generation wherein the model fail to predict the target element within the action description. This error rate is % higher than images without annotation, which exhibits a % error rate because image annotation often obscures critical content within webpages. Overlaying image annotations on webpages without obscuring important content poses a significant challenge. This is primarily due to the dense layout of webpages, which often feature small elements designed for precise cursor interactions. In contrast, the arrangement of objects in mobile UIs and natural scenes tend to be sparser.
Then % of the errors can be attributed to GPT-4V’s tendency of visual illusion , where the model misinterprets and fabricates content over the image. Specifically, the target element described in action generation does not have a bounding box or a label on the bottom-left, where the model is supposed to generate "None". However, the model falsely assumes the presence of a bounding box and makes up a label as the answer. Another % of errors are caused by GPT-4V’s limitation in recognizing the relative position within an image. Specifically, the model is capable of identifying both the target element within the bounding box and the labels. However, it struggles to link the bounding box with its corresponding label.
Powerful Web Action Planner. GPT-4V exhibits promising capabilities, ranging from long-range action planning, webpage content reasoning, and error correction to surpassing the limitations of superficial textual similarity matching inherent in fine-tuned, text-only models. GPT-4V exhibits capability in long-range planning involving a series of interconnected low-level actions. As illustrated in Appendix D, besides generating the immediate next step action, the model also anticipates a sequence of subsequent actions in the future to complete the given task.
Error Correction Awareness. GPT-4V also exhibits the awareness of error correction in the previous actions. In the example in Appendix G, GPT-4V is capable of discovering the mobile phone number is invalid due to wrong format and generate the description about the action to correct this error. This highlights the model’s potential for adaptation in online settings, where actions may not always follow pre-defined, ideal paths as in offline evaluations. This error-correction capability paves the way for adding robustness and reasonable dynamic planning.
Advantages of Knowledge. GPT-4V demonstrates substantial advantages in tasks that requires certain knowledge over fine-tuned models at a smaller scale. As shown in Appendix F, GPT-4V is able to identify the IATA code of the airport in Los Cabos as SJD. In contrast, models are typically weaker at knowledge intensive tasks and also likely to lose knowledge during the fine-tuning process due to catastrophic forgetting.
Related Work
The development of LMMs has been characterized by large-scale image-text pretraining to achieve in-context learning capabilities . Further work focuses on improving instruction following capabilities with curated multimodal instruction-following data and alignment with reinforcement learning from human feedback . GPT-4V and Gemini represent significant progress in LMMs. Several studies have highlighted their remarkable multimodal capabilities, emphasizing the advanced and versatile integration of visual and language reasoning abilities. Their performance on a series of benchmarks also showcases remarkable capabilities on vision-and-language understanding and reasoning. Although open-sourced models still exhibit a performance gap with GPT-4V, they have the advantages of controllability and ease of fine-tuning for various applications. For example, in CogAgent , LMMs are fine-tuned on HTML and screenshot image pairs to enhance webpage understanding ability and further enhanced with an image encoder for high-resolution image details. Ferret is finetuned to allow visual referring and grounding.
2 Web Agent
Driven by the vision of achieving a seamless human-web interaction, considerable initiatives have been invested in web agents that can understand language instruction and complete tasks on the web. Significant efforts have been invested in building benchmarks for web agents with simplified environment simulation . WebArena creates website simulations in the sandbox from four popular categories with functionality and data mimicking their real-world equivalents. Mind2Web instead provides environments in diverse domains, websites, and various tasks on real-world websites by dumping webpages from live websites along with action trajectory annotation. Similar benchmarks have also been proposed for mobile UI automation . To improve web agent performance, WebAgent pretrains a T5 model on HTML data and performs instruction tuning to enable long HTML document summarization and action planning. WebGUM enables web agents to observe both HTML and the rendered screenshot by fine-tuning LMM with a large multimodal corpus for web agents. CogAgent further improves with a high-resolution image encoder to enable the agent to recognize small webpage elements and texts. In a concurrent work , GPT-4V exhibits strong performance on mobile UI understanding, which is much simpler than the desktop websites we study.
3 Visual Grounding
While LMMs, especially GPT-4V and Gemini, have achieved remarkable vision-language understanding capabilities, they still face challenges in fine-grained visual grounding. Various visual prompting methods have been proposed to augment GPT-4V’s image detail grounding ability. Typically, these methods involve overlaying visual marks onto images, facilitating the model’s reference and grounding to specific regions. Various types of visual marks have been explored, such as red circles , mask areas , and combinations of circles, arrows, and hand drawings . SoM involves segmenting the image into semantically meaningful regions and overlaying an array of visual marks like numbers, alphabets, masks, or bounding boxes. Fine-tuning vision-language models with image-annotated data has been shown to be effective. Kosmos-2 represents bounding box locations through textual location tokens. BuboGPT extract entities and find corresponding masks for objects in the image. Shikra handles image detail referring and grounding by applying spatial coordinates as text tokens in inputs and outputs, respectively. Ferret represents regions with both discrete coordinates and continuous features along with a spatial-aware visual sampler to handle diverse spatial characteristics across various shapes.
Limitations and Potential Societal Impact
Web agents demonstrate a great promise in automating routine web tasks through natural language commands, thereby enhancing web accessibility, especially for people less familiar with information technology. However, there are still potential concerns and limitations regarding our experiment scale and safety concerns for deployment in the real world.
Limited Experiment Scale. We can only run experiments on a subset of Mind2Web due to the limited GPT-4V inference quota. We will seek to expand the scale of the experiments.
Agent Safety. While a generalist web agent holds the potential to automate routine web tasks, enhance user experiences and promote web accessibility, safety concerns related to their real-world deployment are also critical. These concerns span privacy issues, such as access to users’ personal profiles, and sensitive operations, such as financial transactions or application form submissions. During online evaluation, we noticed the possibility for these web agents to generate harmful actions on the web, and we manually validate the safety of all the actions before execution. It is critical for further research to thoroughly assess and mitigate the safety risks associated with web agents, ensuring they are safeguarded against producing and executing harmful actions.
Conclusion
In this work, we developed SeeAct, a generalist web agent that harnesses the power of large multimodal models (LMMs) like GPT-4V to integrate visual understanding and acting on the web. We showed that LMMs present a great promise for generalist web agents, with a success rate of % on live websites given an oracle grounding method. GPT-4V also exhibits impressive capabilities, such as error correction and long-range planning. However, fine-grained visual grounding is still a major challenge. The most effective grounding strategies we explored in this paper still exhibit a -% performance gap compared to oracle grounding. Future work should better leverage the unique properties of the Web, e.g., the known correspondence between HTML and visual elements, for improving grounding and reducing hallucinations from LMMs.
Furthermore, we show a significant discrepancy between online and offline evaluations, emphasizing the importance of online evaluation for an accurate assessment of a model’s capabilities. This discrepancy is largely due to the variability in potential plans for completing the same task, pointing to the dynamic nature of web interactions.
Acknowledgments
The authors would like to thank colleagues from the OSU NLP group for their thoughtful comments. This research was supported in part by ARL W911NF2220144 and Cisco.
References
Appendix A Offline Experiment Details
The prompt for action generation is shown in Table 1. For grounding via textual choices, image annotation, and element attributes, the prompts are shown in Tabs. 2, 3 and 4, along with specific task and examples in Figs. 1, 2, 3, 4 and 5.
Appendix B Online Experiment Details
We develop an online evaluation tool using Playwright to load webpages and conduct operations generated by web agents. We manually monitor each step of the model and assess whether it finishes the tasks. We explicitly prohibit attempts to log in or perform final submissions to prevent potentially harmful effects.
MindAct. We strictly follow the original settings in MindAct-FLAN-T5 and MindAct-GPT4. The action space only contains Click, Type, and Select.
SeeActOracle We manually implemented the model’s intended actions. The action history was automatically generated by the model, with an added requirement of summarizing the actions into the "Element", "Operation", "Value" format.
SeeActChoice We still adopt the top- and batch them into three option groups as described in offline experiments. We allow PRESS ENTER and TERMINATE for the model to make confirmation, and stop the process.
During our tests, pop-up ads on webpages were manually closed. The MindAct model was not trained on pop-up ads and hence lacks the feature to automatically manage ads, which could result in stalling. Conversely, SeeActcould proactively suggest ad closure through visual analysis and reasoning.
Appendix C Error Examples for Grounding via Image Annotation
In grounding via image annotation method, we observe significant hallucination errors that can be classified into the following categories:
Making up bounding box & label. In our grounding method, if the correct element is absent from the set of candidate elements, the model is anticipated to generate "None" as the answer. However, as depicted in Figure 6 and Figure 7, the model erroneously claims the element is included within a red bounding box and makes up a wrong index label as the answer.
Failure to link bounding boxes with the correct labels. Another challenge arises in accurately linking bounding boxes to their corresponding index labels. This challenge can be attributed to both LMMs’ limitations in understanding relative spatial positions and the complex, dense layout of webpage elements. The model often mistakenly associates the labels of adjacent elements (as illustrated in Figure 8 and Figure 9) or makes up a label entirely (as demonstrated in Figure 10), rather than accurately predicting the intended index label for the targeted element.
Appendix D Strong Capability of Planning
GPT-4V shows remarkable understanding and planning capabilities during our experiments. As depicted in Figure 11, the model is capable of understanding the website and generate a full plans for the given task involving multiple low-level tasks. Specifically, GPT-4V could understand reasonably well about the process and the remaining work of the task by its careful examination of the webpage, as shown in Figure 12.
Appendix E Challenges in Grounding via Textual Choices
Although textual choices achieved the best results among the three grounding approaches, it still suffers from challenges of similar or identical elements which are common in webpages. The model tends to choose the first text choice that seemingly corresponds to its intention. Moreover, this is inevitable, as web pages indeed contain many elements that may even have exact identical HTML information, as the "Schedule" button shown in Figure 13.
Appendix F Knowledge and Reasoning Requirement
Some tasks require a certain degree of reasoning and knowledge, which may be challenging for fine-tuned models like MindAct. For instance, the task in Figure 14 necessitates the model to know the specific district of Dublin in Virginia. In the task of Figure 15, the model correctly provided the IATA airport code of airports in Indira Gandhi and Los Cabos.
Appendix G Path Variation and Awareness of Error Correction
On webpages, multiple paths often exist to accomplish a given task. For instance, varying the execution order of actions within an interchangeable sequence can result in diverse routes to task completion. Additionally, the agent can navigate to different webpages but still accomplish the give tasks. Figure 16 presents a straightforward example where the model chose a more direct route that differs from the ground truth annotated in the dataset.
When running on live website, the agent’s previous action histories is likely to be filled with redundant, unnecessary, erroneous, failed operations generated, or merely exploratory attempts by the model, resulting in a final path that deviates significantly from the ground truth. Despite these circumstances, the model can still accomplish the task amidst numerous incorrect explorations. The process of exploration and correction requires the model to possess a sense of self-correction. As shown in Figure 17, GPT-4V demonstrates this awareness of correcting errors caused by previous steps.