GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation
An Yan, Zhengyuan Yang, Wanrong Zhu, Kevin Lin, Linjie Li, Jianfeng Wang, Jianwei Yang, Yiwu Zhong, Julian McAuley, Jianfeng Gao, Zicheng Liu, Lijuan Wang
Introduction
Building autonomous agents capable of interacting with computing devices and following human commands has been a long-standing topic in the machine learning community (Bolt, 1980; Lieberman et al., 1995). Since the advent of smartphones, there has been a practical demand for creating virtual assistants, like Siri, Cortana, and Google Assistant, which have the potential to significantly enhance user experience and assist individuals who are physically or situationally impaired. Ideally, these assistants would competently carry out everyday tasks based on natural language instructions, ranging from simple actions like setting a timer to more complex tasks such as locating the ideal hotel for a family vacation.
Recent studies have started to explore mobile device control and smartphone task automation following human instructions (Rawles et al., 2023; Wen et al., 2023; Zhan and Zhang, 2023; Wang et al., 2023). Representative approaches include describing screen images with text and processing converted text with large language models (LLMs) (Rawles et al., 2023; Wen et al., 2023), or training a vision-language model to generate actions in a supervised manner (Rawles et al., 2023; Zhan and Zhang, 2023). However, these supervised models, when trained on specific types of screens and instructions (Rawles et al., 2023), exhibit limited effectiveness in generalizing to real-world scenarios. On the other hand, the LLM-based approaches generalize better, but the intermediate step of converting screen images to text results in information loss and consequently hurts performance. Inspired by the efficacy and broad applicability of recent large multimodal models (LMMs), we explore utilizing an LMM, GPT-4V (OpenAI, 2023a, b, c; gpt, 2023; Yang et al., 2023c), for zero-shot smartphone GUI navigation, aiming to set a new strong baseline for this intriguing task.
We identify two primary challenges for GUI navigation with LMMs, namely intended action description and localized action execution. First, the model should understand the screen image and text instruction input, and reason over the query to determine the appropriate action to take, such as providing a natural language description “clicking the Amazon icon in the third row and fourth column.” Second, the model should convert such high-level understanding into a formatted action that can be easily executed based on rules, such as “{Action: Click, Location: (0.31, 0.57)}.” In our approach, we prompt GPT-4V with an image and text for action planning, and place set-of-mark tags (Yang et al., 2023b) to anchor the generated outputs. Specifically, we associate these marks with spatial locations with the help of segmentation or OCR models. To this end, our proposed GPT-4V-based system, namely MM-Navigator, can generate executable actions conditioned on the screen image, the text instruction and its interaction history.
We benchmark MM-Navigator on two datasets. We start with an iOS GUI navigation dataset with screenshots and user instructions that we manually collected. This clean analytic dataset is designed to probe insights for the two challenges in GUI navigation: intended action description and localized action execution. Human evaluations are used to assess GPT-4V on these two tasks, with accuracy rates of and , respectively. Additionally, we assess the model on a random subset from the recently released Android navigation benchmark (Rawles et al., 2023). We follow the proposed evaluation protocol in the benchmark, together with extra human evaluations. The strong performance demonstrates that MM-Navigator is an effective GUI navigator for smartphones, significantly outperforming previous LLM-based approaches. We provide in-depth analyses of the representative success and failure cases. We find that the current state of GPT-4V may already be effective in aiding humans in various real-world GUI navigation scenarios, as evidenced by the multi-screen results in Figure 4. However, continued enhancements are still essential to further increase the system’s reliability, as revealed in our analyses.
Our contributions are summarized as follows.
We present MM-Navigator, an agent system built on GPT-4V for smartphone GUI navigation. MM-Navigator effectively incorporates action histories and set-of-mark tags to produce precise executable actions.
We collect a new analytic dataset with diverse iOS screens and user instructions, which evaluates two main challenges in GUI navigation with LMMs: intended action description and localized action execution.
We perform extensive evaluations, both automatic and human, on two datasets and provide detailed analyses. The impressive results demonstrate the effectiveness of MM-Navigator for GUI navigation.
Related Work
Autonomous GUI navigation involves a model following instructions to maneuver through different graphical user interfaces, such as websites or applications, to perform the user-queried task. Current benchmarks collected either synthetic or real-world user-generated instructions to evaluate models’ abilities in identifying specific UI elements (Shi et al., 2017; Li et al., 2020; Bai et al., 2021), or achieving overarching task objectives by interacting with a series of GUI views (Li et al., 2020; Burns et al., 2021; Venkatesh et al., 2022; Deng et al., 2023; Rawles et al., 2023). To understand the visual information from these GUI views, one line of work adopts a model structure that can process multimodal inputs (Sun et al., 2022; Redmon et al., 2016). Other methods focus on converting the UI scene text and icons into the text-only HTML format, such as single-module LLMs can process these text inputs for GUI navigation (Zhang et al., 2021; Rawles et al., 2023; Wen et al., 2023).
Multimodal agents.
Recent advancements in LLMs (Brown et al., 2020; OpenAI, 2023a; Chowdhery et al., 2022; Anil et al., 2023; Touvron et al., 2023; Hoffmann et al., 2022) have catalyzed the exploration of LLM-based agent systems (Madaan et al., 2023; Shinn et al., 2023; Pan et al., 2023; Yao et al., 2022; Schick et al., 2023; Paranjape et al., 2023; Pryzant et al., 2023; Guo et al., 2023; Zhao et al., 2023; Yang et al., 2023a), which integrate reasoning logic and external tools for a variety of complex language tasks. Inspired by the success in the NLP domain, multimodal researchers delve into multimodal agents. The line of research begins with LLM-based multimodal agents (Gupta and Kembhavi, 2023; Surís et al., 2023; Wu et al., 2023; Yang* et al., 2023; Shen et al., 2023; Lu et al., 2023; Yu et al., 2023; Li et al., 2023), such as MM-ReAct (Yang* et al., 2023) for advanced visual reasoning and Visual ChatGPT (Wu et al., 2023) for iterative visual generation and editing. Propelled by the rapid advancements of LMMs (Alayrac et al., 2022; Driess et al., 2023; OpenAI, 2023a, b, c; gpt, 2023; Yang et al., 2023c; Google, 2023), the latest studies have begun to investigate the LMM-powered multimodal agents (Yang et al., 2023; Liu et al., 2023), thereby surpassing the need for basic visual description tools like caption models (Wang et al., 2022a; Wu et al., 2022). Our proposed methodology represents a specialized LMM-based agent for GUI navigation. We aim to provide a comprehensive analysis and a strong baseline for this task.
MM-Navigator
When presented with a user instruction in natural language, the agent is asked to complete a series of actions on the smartphone to complete this instruction. The entire process of agent-environment interactions from initial to final states is called an episode. At each time step of an episode, the agent will be given a screenshot , and decide the next step action to take in order to complete the task.
2 Screen Grounding and Navigation via Set of Mark
GPT-4V serves as a multimodal model that takes visual images and text as inputs and produces text output. One challenge is how do we communicate with GPT-4V to perform actions on screen. A possible solution is to ask the model to reason about coordinates to click given a screen. However, based on our preliminary exploration, though GPT-4V have a good understanding of the screen and approximately where to click to perform an instruction by describing the corresponding icon or text, it appears to be bad at estimating accurate numerical coordinates.
Therefore, in this paper, we seek a new approach, to communicate with GPT-4V via Set-of-Mark prompting (Yang et al., 2023b) on the screen. Specifically, given a screen, we will detect UI elements via the OCR tool and IconNet (Sunkara et al., 2022). Each element has a bounding box and either OCR-detected text or an icon class label (one of the possible 96 icon types detected by (Sunkara et al., 2022)) are contained. At each step time , we add numeric tags to those elements, and present GPT-4V with the original screen and the screen with tags . The output text of GPT-4V will be conditioned on the two images. If GPT-4V decides to click somewhere on the screen, it will choose from the available numeric tags. In practice, we found this simple method works well, setting up a strong baseline for screen navigation with large multimodal models.
3 History Generation via Multimodal Self Summarization
Set-of-Mark prompting bridges the gap between text outputs from GPT-4V and executable localized actions. However, the agent’s ability to maintain a historical context is equally important in successfully completing tasks on smartphones. The key difficulty lies in devising a strategy that allows the agent to effectively determine the subsequent action at each stage of an episode, taking into account both its prior interactions with the environment and the present state of the screen. The naive approach of feeding all historical screens or actions into the agent is computationally expensive and may decrease the performance due to information overload. For example, screens at each step can change rapidly, and most of the historical screen information is not useful for reasoning about future actions. Humans, on the other hand, can keep track of a short memory of the key information after performing a sequence of actions. We aim to find a more concise representation than a sequence of screens or actions. Specifically, at each time step, we ask GPT-4V to perform multimodal self summarization, which converts the historical actions and current step information into a concise history in the form of natural language, which is formulated as follows:
where is the action to take at current step , is the summarized history based on and , is the parameterized GPT-4V model. In this way, the trace of history will be generated auto-regressively when an episode is rolled out.
iOS Screen Navigation Experiment
We begin by conducting analytical experiments on iOS screens to understand GPT-4V’s capability in GUI navigation. Successfully operating smartphones in a human-like manner involves different types of screen understanding abilities. Firstly, there is the semantic reasoning ability, which involves comprehending screen inputs and articulating the necessary actions to fulfill given instructions. Secondly, there is the need to translate these action descriptions into specific localized actions, such as determining the precise location for a screen click. Correspondingly, we develop two sets of test screens to disentangle these two aspects, which are referred to as “intended action description” and “localized action execution,” respectively.
In this study, we gather instructions from human annotators, evenly divided into two distinct sets, containing iOS screens with and without added marks. The first set, “intended action description,” involves GPT-4V taking an iOS screenshot image and an instruction as inputs, and generating an open-ended text description of the desired action to perform. This set aims to assess GPT-4V’s ability to reason the correct action to perform. Moving beyond having someone click the screen for GPT-4V (Yang et al., 2023c; Lin et al., 2023), we investigate directly generating formatted executable actions. In the second set, “localized action execution,” we add marks (Yang et al., 2023b) to ground screen locations with interactive SAM (Kirillov et al., 2023), and let GPT-4V use the mark indexes to perform localized actions. Other approaches, such as textualized box coordinates (Chen et al., 2022; Yang et al., 2022; Wang et al., 2022b), screen visual grounding (Yu et al., 2016; Mao et al., 2016; Plummer et al., 2015; Yang et al., 2019; Deng et al., 2021), object detectors (Ren et al., 2015; Carion et al., 2020) could also translate action descriptions into executable actions.
Human evaluation metrics.
We use human evaluation for the analytical experiments on iOS screens, with a binary score for each sample indicating if the output is correct. For “intended action description”, human annotators determine if the output text description could lead to the correct output. For “localized action execution,” human annotators assess if clicking the location (i.e., location of the selected mark) fulfills the given instruction. Each sample is assigned a binary score, either 0 or 1, to reflect its correctness.
2 Intended Action Description
Table 1 reports an accuracy of on generating the correct intended action description, quantitatively supporting GPT-4V’s capability in understanding the screen actions to perform (Yang et al., 2023c; Lin et al., 2023). Figure 1 showcases representative screen understanding examples. Given a screen and a text instruction, GPT-4V gives a text description of its intended next move. For example, in Figure 1(a), GPT-4V understands the Safari browser limits of “the limit of 500 open tabs,” and suggests “Try closing a few tabs and then see if the "+" button becomes clickable.” Another example is telling the procedure for iOS update: “You should click on "General" and then look for an option labeled "Software Update” in (b). GPT-4V also effectively understands complicated screens with multiple images and icons. For example, in (c), GPT-4V mentions, “For information on road closures and other alerts at Mt. Rainier, you should click on "6 Alerts" at the top of the screen.” Figure 1(d) gives an example in online shopping, where GPT-4V suggests the correct product to check based on the user input of the desired “wet cat food.”
3 Localized Action Execution
A natural question is how reliable GPT-4V can convert its understanding of the screen into executable actions. Table 1 shows an accuracy of on selecting the location that could lead to the desired outcome. Figure 2 shows the added marks with interactive SAM (Yang et al., 2023b; Kirillov et al., 2023), and the corresponding GPT-4V outputs. As shown in Figure 2(a), GPT-4V can select the “X” symbol (ID: 9) to close the tabs, echoing its previous description in Figure 1(a). GPT-4V is also capable of selecting the correct location to click from the large portion of clickable icons, such as the screen shown in (b). Figure 1(c) represents a complicated screen with various images and icons, where GPT-4V can select the correct mark for the reading the “6 Alerts.” Within a screen with various texts, such as (d), GPT-4V can identify the clickable web links, and locate the queried one with the correct position .
4 The Current State with GPT-4V
From the analytical experiments on iOS screens, we find GPT-4V is capable of performing GUI navigation. Although several types of failure cases still occur, as outlined below, MM-Navigator shows promise for executing multi-screen navigation to fulfill real-world smartphone use cases. We conclude the section with qualitative results on such episode-level navigation queries.
Despite the promising results, GPT-4V does make errors in the zero-shot screen navigation task, as shown in Table 1. These errors are illustrated through representative failure cases as follows. (a) GPT-4V might not generate the correct answer in a single step when the query involves knowledge the model lacks. For example, GPT-4V is not aware that only “GPT-4” can support image uploads, hence it fails to click the “GPT-4” icon before attempting to find the image uploading function. (b) Although usually reliable, GPT-4V might still select the incorrect location. An example of this is selecting the mark for the “ChatGPT” app instead of the correct mark . (c) In complex scenarios, GPT-4V’s initial guess might not be correct, such as clicking the “numeric ID 21 for the 12-Hour Forecast” instead of the correct answer of mark 19. (d) When the correct clickable area is not marked, like a “+” icon without any marks, GPT-4V cannot identify the correct location and may reference an incorrect mark instead. Finally, we note that many of those single-step failures may be corrected with iterative explorations, leading to the correct episode-level outcome.
From single screens to complete episodes.
MM-Navigator shows an impressive capability in performing GUI navigation in a zero-shot manner. We further extend MM-Navigator from processing a single cellphone screen to recursively processing an episode of screen inputs. Figure 4 shows the qualitative result. In each step, we include the objective, “You are asked to shop for a milk frother, your budget is between 100.” and its previous action in the prompt to GPT-4V. We show that the model can effectively perform multi-step reasoning to accomplish the given shopping instruction.
Android Screen Navigation Experiment
We use the AITW dataset (Rawles et al., 2023) for our evaluation on Android screen navigation. AITW is a large-scale benchmark dataset for UI control, which contains natural language instructions, screenshots on different Android systems with different resolutions, and user-annotated actions. It covers diverse multi-step tasks such as various web and application operations, app installation, and tasks with Google apps, with 715K episodes and 30K unique instructions in total. Table 2 shows the basic statistics of the dataset. We follow the split from previous work (Zhan and Zhang, 2023). Following the previous experiment setting (Rawles et al., 2023) that evaluates PaLM 2 on a randomly sampled 288 episodes, we sample 300 episodes from the test split as our test set.
Metrics.
Following previous work (Rawles et al., 2023; Zhan and Zhang, 2023), we compute the screen-wise partial action matching score as the main evaluation metric, defined as the number of correct actions divided by the episode length, then this score is averaged over all tested episodes. A predicted action from GPT-4V is considered correct if both the action type and gesture match the gold ones, i.e., user actions. For click actions, it is considered correct if the selected element falls within a 14% screen distance from the gold gestures or occurs within the same detected bounding box with user gestures. For scroll actions, it is considered correct if the selected direction has the same scroll direction (up, down, left, and right) as user gestures. The partial score has been shown to correlate with the task complete score estimated by human evaluations (Rawles et al., 2023) to measure the action success rate of this task.
Baselines.
We compare with the following baselines (Rawles et al., 2023; Zhan and Zhang, 2023):
PaLM-2 ZS (Rawles et al., 2023): Zero-shot performance with PaLM-2 (Anil et al., 2023), by feeding a textual description of the screen and ask it to predict an action among the supported actions in AITW. We adopt a previously proposed LLM-based design for device control (Wang et al., 2023), where the input screen description is converted to HTML syntax.
PaLM-2 5-shot (Rawles et al., 2023): Five examples of navigation are designed as Chain-of-thought prompts. The history of prior actions taken by the agent is also fed into the model input.
ChatGPT 5-shot (Zhan and Zhang, 2023). The input prompts are of the same format as PaLM-2 5-shot. Experiments are conducted via the ChatGPT API.
Fine-tuned Llama-2 (Zhan and Zhang, 2023): Fine-tuning Llama-2 model (Touvron et al., 2023) with LoRA (Hu et al., 2021), by feeding the model with the user instruction and screen descriptions in HTML syntax (the same that are used for in-context learning LLMs) and predict user actions. The model is fine-tuned with randomly sampled training data to help adapt to this task.
2 Performance Comparison
Our main results are shown in Table 3. First, GPT-4V outperforms previous LLMs that take ground-truth descriptions of the screens as inputs. Compared with previous text-only LLMs, taking screen images as visual inputs provides an easier way for human-model interactions. It also better preserves the screen information and avoids the information loss when converting screens to text descriptions. Additionally, adding screen descriptions still improves the performance of GPT-4V. Giving the agent access to its historical interactions is helpful for better conditioned and grounded generation, and our in-context self-summarization module provides an efficient way to achieve this. Overall, we find GPT-4V presents a strong level of screen understanding of icons and text, showing the potential of visual-based device control with LMMs.
3 Ablation Studies
For the ablation studies, we randomly sampled 50 episodes in total from 5 categories, which is a different subset used by the main results.
We first perform an ablation study to compare the performance with different methods to add tags on screen, shown in Table 4. We consider three methods: (1) By side which adds tags with black squares (same style as (Rawles et al., 2023) by the left side of each detected icon; (2) Red which uses red circles for each tag; (3) Center which adds tags with black squares at the center of each detected box. First, adding tags by the left side of boxes may cause problems, for example, some icons may be too close to each other, hence leading to slightly worse results. For tagging styles, we didn’t find a significant difference between red cycles and black rectangles, though empirically black rectangles (Yang et al., 2023b) perform slightly better.
Different prompts.
We then perform robustness check with different prompting variants: (1) Baseline: Simply ask GPT-4V to take actions; (2) Think: Prompt GPT-4V to think step by step (Kojima et al., 2022); (3) Detail: Provide more context for this task. Overall, we did not observe improvements by “thinking step by step”, but adding more task descriptions helps GPT-4V to better execute actions.
4 Error Analysis
We look into GPT-4V prediction traces and attempt to categorize common types of errors that cause mismatching between GPT-4V predictions and human annotations.
We notice false negative cases where the mismatches are rooted in inaccurate Set-of-Mark (Yang et al., 2023b) annotation parsing or imperfect dataset annotation. In these cases, the predictions made by GPT-4V are correct after manual justification, but are classified as wrong predictions in automatic evaluation because the target regions are over-segmented (e.g., Figure 5(a)(b)), or because the ground-truth annotation only covers one of the many valid actions (e.g., Figure 6(a) has two Google Play logo; Figure 6(b) has multiple ways of accessing Google Search; and users may lookup “Setting” by direct search as GPT-4V, or by scrolling down as the human annotation in Figure 6(c)).
Figure 7 shows a few true negative examples of GPT-4V failing the designated tasks. In our zero-shot testing setup, GPT-4V is not provided with demonstrative examples to learn user action patterns. In this case, while users may scroll down or up to explore the GUI, we notice GPT-4V is more likely to perform the action of “click” on each screen, leading it to occasionally make short-sighted decisions. In Figure 7(a), GPT-4V attempts to look for “improve location accuracy” in “Network&Internet” among the listed visible tabs, while the user decides to scroll down and look for more aligned setting tabs. In Figure 7(b), GPT-4V clicks on “Accept All”, which is not a button. In Figure 7(c), GPT-4V also shows a more literal understanding of the instruction and the current observation as in (b), clicking the “News” tab in the Google Search platform instead of actually visiting the news website.
Discussion
For future benchmarks, more dynamic interaction environments are needed. Even humans can make mistakes sometimes, and in this case, it is important that the evaluation benchmark would allow the model to explore and return to previous status when a mistake is made and realized by the model. It is also interesting to explore how to automatically evaluate success rates for this task, e.g., by using LMMs (Zhang et al., 2023). Another direction is to build GUI navigation datasets with different devices and diverse contents, e.g., personal computers and iPads.
Error correction.
A pretrained LMM may make mistakes due to data or algorithm bias. For example, if the agent fails to complete tasks in certain novel settings, how do we correct its errors to avoid mistakes in the future? Moreover, it would be interesting to study this in a continual learning setting, where the agent keeps interacting with new environments and receives new feedback continually.
Model distillation.
Using a large-scale model such as GPT-4V for GUI navigation is costly. In the future, it would be interesting to explore model distillation (Polino et al., 2018) for this task, to obtain a much smaller model with competitive navigation performance, which may achieve lower latency and higher efficiency.
Conclusion
We have presented MM-Navigator, a GPT-4V-based multimodal agent system designed for the GUI navigation task. The system is benchmarked on both our collected iOS dataset and a public Android navigation dataset, revealing GPT-4V’s exceptional capabilities in understanding, reasoning, and planning over the screen environments. For future works, one promising direction is to establish a simulator-based benchmark that incorporates multi-step and episode-level automatic evaluations.