You Only Look at Screens: Multimodal Chain-of-Action Agents

Zhuosheng Zhang, Aston Zhang

Introduction

Building intelligent autonomous agents that are capable of task planning, decision making, and action execution in a particular environment is a long-standing goal of artificial intelligence (AI) (Searle, 1969; Wooldridge & Jennings, 1995; Maes, 1995; Hendler, 1999). The advent of large language models (LLMs) (Brown et al., 2020; Chowdhery et al., 2022; OpenAI, 2023) has flourished promising opportunities for developing autonomous agents to assist users in completing tasks in distinct environments such as operation systems, specific applications, and web browsers (Adept, 2022; Rawles et al., 2023; Liu et al., 2023; Zhou et al., 2023; Wang et al., 2023c).

Recent studies have explored prompt engineering (Richards, 2023; Nakajima, 2023; Reworkd, 2023; Sumers et al., 2023; Liu et al., 2023) and fine-tuning techniques (Rawles et al., 2023; Wen et al., 2023; Sun et al., 2022) to elicit the capability of language models to execute actions in interactive environments. However, there are at least two major challenges that have limited real-world applications of autonomous agents.

First, existing approaches commonly rely on external tools such as optical character recognition (OCR) and icon detectors (Zhang et al., 2021; Sunkara et al., 2022) to parse the environment into textual elements (e.g., HTML layouts) as inputs to a language model (Figure 1(a)) (Rawles et al., 2023; Wen et al., 2023). On the one hand, the parsed elements generate lengthy inputs, thus leading to inference efficiency. Since computational latency is a key measure in deployment, using lengthy inputs would increase inference cost and may even exceed the input length limit of the language model. On the other hand, parsing the visual environment into textual elements may also be prone to error propagation or information loss because parsing mistakes are inevitable using external tools.

Second, most existing approaches are under the sand-box setting that requires accessing internal APIs to interact with the environment (Zhou et al., 2023; Gur et al., 2023), e.g., using a JavaScript element selection on a webpage or a Python interpreter to execute actions. However in practice, the API interface is often inaccessible in third-party applications (Apps).

These challenges have motivated more advanced techniques that are capable of first principles thinking (Aristotle, ; Irwin, 1989)—allowing direct interactions on the screen without needing access to intermediate environment parsing or interval application-dependent APIs (Figure 1(b)). To address the challenges, we introduce Auto-UI, a multimodal approach that directly interacts with the interface. To improve the agent’s action prediction capability, we propose a novel chain-of-action technique, where a chain of action is a series of intermediate previous action histories and future action plans that lead to action prediction.

We evaluate Auto-UI on a new device-control benchmark AITW (Rawles et al., 2023) with 30KK unique instructions, spanning multi-step tasks of application operation, web searching, and web shopping. Experimental results show that Auto-UI achieves state-of-the-art performance with an action type prediction accuracy of 90% and an action success rate of 74%.

In summary, our work makes the following technical contributions:

We introduce Auto-UI, a multimodal agent for autonomous UI control that can directly interact with the screens, thus circumventing the constraints of environment parsing and application-specific API access.

We propose a chain-of-action technique that leverages the previously executed actions and future action plans to help the agent decide what action to execute at each step.

Auto-UI achieves state-of-the-art performance with an action type prediction accuracy of 90% and an action success rate of 74%. Notably, Auto-UI can infer an action as fast as within less than one second.

Related Work

Our work falls into the field of language agents. This section will first review the recent progress in building language agents and then discuss the approaches to conduct user interface control with language agents.

Language agents refer to those agents that can follow user instructions and interact with environments to complete tasks. Such agents expand the landscape of language models to compete in specific fields, including application operation, web searching, and web shopping. There are two popular types of language agents, autonomous agents and communicative agents. Autonomous agents aim to assist humans to achieve specific goals in the real world. Typical examples of autonomous agents are AutoGPT (Richards, 2023), BabyAGI (Nakajima, 2023), and AgentGPT (Reworkd, 2023). In contrast, communicative agents are personalized and socialized agents (Park et al., 2023; Wang et al., 2023b; Zhu et al., 2023; Hong et al., 2023) with human behaviors that can communicate and collaborate with each other. They are often deployed in immersive environments. Inspired by the potential in real-world applications, this work focuses on autonomous agents, especially those working in mobile devices. We aim to assist users by completing multi-step tasks (e.g., manipulating Apps, web shopping, and question answering) without any manual intervention. Given a user instruction in natural language, the agent is required to interpret the instruction and execute actions by directly controlling its user interface. Due to the requirement in real-world applications, the agent is expected to be both effective and efficient.

2 UI Control with Natural Language

Recently, LLMs have shown promise in building autonomous UI agents with abilities of instruction following (Sanh et al., 2021; Taori et al., 2023b; Chiang et al., 2023) and chain-of-thought (CoT) prompting (Nye et al., 2022; Wei et al., 2022). Especially, CoT prompting (Wei et al., 2022; Kojima et al., 2022; Zhang et al., 2023a) elicit LLMs’ capacities of step-by-step planning, decision making, and action execution. Those capacities have been shown to be effective in UI control tasks (Rawles et al., 2023). However, the task environments are graphical user interfaces (GUIs), instead of natural language that LLMs can directly process. Therefore, the GUI states and actions are required to be converted to textual formats to conform to the input and output formats of LLMs. For example, it is feasible to parse the UI screens by icon recognition and OCR (Zhang et al., 2021; Sunkara et al., 2022) and organize the parsed elements into HTML layouts. As a compromise, existing approaches are restricted in a sandbox setting where they rely on external tools (Rawles et al., 2023; Wen et al., 2023) and application-specific APIs (Zhou et al., 2023; Gur et al., 2023) for environment parsing and action interpretation; thus, commonly suffer from inference efficiency and error propagation. Although there are studies that have considered multimodal architecture to process inputs in different modalities (Sun et al., 2022), however, those studies still rely on fine-grained environment parsing to ensure competitive performance. In contrast, this work is established upon first principles thinking, which directly reads the UI without additional environment parsing and provides the action (e.g., action type, gesture coordinate, and typed text) that can be executed without needing any extra APIs.

Methodology

In this section, we will first introduce the basic concepts for the UI control task and then describe the design of our proposed Auto-UI framework.

Given a user instruction (also known as a goal), the agent needs to complete the task with multiple steps of interactions. The entire process is called an episode, which is composed of a series of screens. For each step in the episode, the agent will be provided with a screenshot, and the agent is required to predict the action until the task is complete. Detailed examples can be found in Appendix A.1.

2 Framework Overview

Auto-UI is a multimodal agent that decides what action to take given the input screenshot and a user instruction. To empower the agent’s decision making capability, we introduce a chain-of-action approach by leveraging a series of intermediate previous action histories and future action plans to predict actions.

The model architecture of Auto-UI is illustrated in Figure 2. On a high level, Auto-UI consists of three stages. First, we acquire encoded features from both vision and language inputs. Specifically, the vision input, i.e., a screenshot, is encoded by a frozen vision encoder. Meanwhile, the language input, consisting of the goal and a chain of previous action histories—each history contains a tuple {action type, touch point, lift point, and typed text}, is encoded by a language encoder. Second, the encoded vision and language representations are integrated by a self-attention module. Third, the fused representation is fed to the decoder to generate a chain of future action plans (i.e., action types to execute in future steps) followed by action prediction. A chain of action consists of two parts in the procedure above: a chain of previous action histories on the input side and a chain of future action plans on the output side. In the following, we describe the entire procedure in detail.

If t>1t>1, there will be a chain-of-action history that has already been executed before step tt. We denote the chain of action histories as Xhistory=[m1,…,mt]X_{\textrm{history}}=[m_{1},\dots,m_{t}] where mim_{i} contains a tuple of action type, touch point, lift point, and typed text. Otherwise, if t=1t=1, XhistoryX_{\textrm{history}} will be set empty:

We concatenate XgoalX_{\textrm{goal}} and XhistoryX_{\textrm{history}} as the input to the language encoder:

Then, we obtain the encoded representations of the vision and language inputs as follows:

where WW is a trainable projection matrix to convert HscreenH_{\textrm{screen}} into the same dimensionality as HlanguageH_{\textrm{language}}.

Interaction

where dkd_{k} is the same as the dimension of HlanguageH_{\textrm{language}} because a single head is used.

where WlW_{l} and WvW_{v} are learnable parameters.

Decoding

The fused representation HfuseH_{\textrm{fuse}} is fed to a Transformer decoder to generate the target predictions in a string format. The target predictions consist of a chain of future action plans YplanY_{\textrm{plan}} and the current action prediction YactionY_{\textrm{action}} separated by specific prompts: {Action Plan: YplanY_{\textrm{plan}}, Action Decision: Yaction}Y_{\textrm{action}}\}. Concretely, YplanY_{\textrm{plan}} is a chain of action types to execute in future steps: YplanY_{\textrm{plan}} = [action_typet, …\dots, action_typek]. YactionY_{\textrm{action}} contains four components: YactionY_{\textrm{action}} = {“action_type”: , “touch_point”: , “lift_point”: , “typed_text”: }. These four components will be explained in the following subsection.

3 Coordinate Normalization

Recall that a target action consists of four components: action type, touch point, lift point, and typed text. We consider six action types: dual-point gesture, type, go_back, go_home, enter, and status_complete. A dual-point gesture comprises a touch point and a lift point with [y,x][y,x] coordinates. The gesture actions ensure a flexible action space and can represent clicks and scrolls at arbitrary locations. For example, a gesture action {“touch_point”: [0.7761, 0.7089], “lift_point”: [0.7761, 0.7089]} means clicking at the coordinate [0.7761, 0.7089], while a gesture action {“touch_point”: [0.1898, 0.4477], “lift_point”: [0.8242, 0.4077]} means scrolling down. A type action means typing a text and the text is placed in the field. The other action types, i.e., go_back, go_home, enter, and status_complete are system actions, whose corresponding , fields are filled with -1, and the is empty.

We observe that high-precision coordinates are not necessary for representing a click or scroll action. Therefore, we apply normalized values of the coordinates, which helps accelerate convergence and mitigate the ambiguity of coordinates. The normalization is applied to click and scroll actions. For click actions, we keep four decimal places. For scroll actions, we first determine the scroll direction with the touch point and lift point. Then, we transform the touch and lift points into fixed directional coordinates as follows: “up”: {[0.8, 0.5], [0.2, 0.5]}, “down”: {[0.2, 0.5], [0.8, 0.5]}, “left”: {[0.5, 0.8], [0.5, 0.2]}, “right”: {[0.5, 0.2], [0.5, 0.8]}, where {[⋅\cdot], [⋅\cdot]} consists of the touch point and lift point in the first [⋅\cdot] and second [⋅\cdot]. We provide examples of target actions in Appendix A.2.

Experiments

We use the AITW benchmark dataset (Rawles et al., 2023) for our evaluation. AITW is a large-scale benchmark dataset for UI control, which contains natural language instructions, screenshots, and actions. There are 715KK episodes spanning 30KK unique instructions, covering diverse multi-step tasks such as application operation, web searching, and web shopping, on more than 350 Apps and websites. To ensure generality, this dataset also covers various device types and operation systems in varying screen resolutions.

There are five subsets in the benchmark dataset, namely, General, Install, GoogleApps, Single, and WebShopping. Table 1 presents the data statistics. Each subset is split episode-wise into a training, validation, and test set (80/10/10%).

(i) General contains miscellaneous tasks that need interaction with third-party Apps and websites, as well as question answering.

(ii) Install contains tasks related to installing and uninstalling Apps, App login, and App login support.

(iii) GoogleApps contains tasks about manipulating various Google applications such as Gmail, Calendar, Photos, and Settings.

(iv) Single contains atomic tasks (e.g., “upvote the post”) whose preceding actions have been already completed (e.g., opening Instagram, going to home feed, looking at a post).

(v) WebShopping contains tasks related to online shopping on E-commerce websites, e.g., searching for an item, adding an item to the cart, and viewing the shopping cart.

2 Baselines

We adopt three types of baselines for comparisons.

(i) Specialized UI Agent. We adopted the Behavioural Cloning (BC) agent, which reported the state-of-the-art performance in Rawles et al. (2023). BC is a Transformer-based architecture that takes a task instruction, the current screen, and a stacked history of screen observations and actions as input. The task instruction and OCR-detected texts are encoded by a pre-trained BERT. The icons are represented by the embeddings for each of the bounding box points. The screen history is modeled by the {x,y}\{x,y\} positions of the touch and lift actions. All the embedded representations are fused to predict the action by a decoder. There are two BC variants, BC-single and BC-history, depending on whether the model takes as input the screen-action history.

(ii) In-context Learning LLMs. PaLM 2 and ChatGPT (turbo-3.5) are adopted. Following previous studies (Rawles et al., 2023; Wang et al., 2023a), we feed the LLM a textual description of the screen and a user instruction. The textual description of the screen is formatted as an HTML syntax, providing the information of UI elements derived from OCR detection and icon detection from external tools (Rawles et al., 2023). The model is required to predict an action among pre-defined actions. If the action is clicking, the model will be required to provide the index of the clicked UI element. Alternatively, the model needs to provide the scroll direction if the action is scrolling. In addition, 5-shot CoT prompting is leveraged to improve the performance (Appendix A.3).

(iii) Fine-tuned LLMs. We adopt Llama 2 (Touvron et al., 2023) as the baseline and fine-tune it with LoRA. We feed the model with the user instruction and the screen descriptions in HTML syntax (the same as adopted for in-context learning LLMs). The model is expected to predict the action in the same output format as in-context learning LLMs. As fine-tuning an LLM is expensive, we randomly sample 1% training data to help the LLM adapt to our tasks.

3 Evaluation Measures

We compute the screen-wise action matching score as the main evaluation measure, defined as the number of correct actions divided by the episode length. A predicted action is considered correct if the action type and dual-point gesture match the gold ones. As we described in Section 3.3, the gesture actions can represent the click actions and scroll actions at arbitrary locations. A click action is considered correct if its touch point and lift point fall within a 14% screen distance from the gold gestures or occur within the same detected bounding box with the gold gestures. A scroll action is considered correct if it has the same scroll axis as the gold gesture.

The screen-wise action matching score has been shown to correlate with the task complete score estimated by human evaluations (Rawles et al., 2023) and is appropriate to measure the action success rate for user instructions. Besides the overall matching score, we will also compare the click region accuracy, scroll direction accuracy, action type accuracy, and typed text accuracy for a more comprehensive reference (Section 5.1).

The evaluation criteria apply to the BC baselines and our Auto-UI. For the LLMs, they can only click on detected UI elements, rather than clicking at arbitrary locations. Therefore, we consider if the clicked UI element is matched for click actions instead of comparing dual-point gestures for LLMs.

4 Implementation Details

We adopt the encoder-decoder architecture (Raffel et al., 2020) under small (60M), base (200M) and large (700M) settings in our framework. We apply FLAN-Alpaca to initialize our model weights.https://github.com/declare-lab/flan-alpaca. The vision features are obtained by the frozen BLIP-2 encoder (Li et al., 2023) (version: blip2_t5_instruct). We fine-tune the models up to 10 epochs, with a learning rate of 1e-4. The maximum input sequence length is 512. The batch size is 4. Our experiments are run on 8 NVIDIA Tesla V100 32G GPUs. Training the large and base models takes 75 and 25 hours, respectively.

We develop two kinds of approaches to analyze their generalization abilities, namely Auto-UIseparate{}_{\text{separate}}, and Auto-UIunified{}_{\text{unified}}. Specifically, Auto-UIseparate{}_{\text{separate}} is trained and evaluated independently on each subset. Auto-UIunified{}_{\text{unified}} is a unified model trained on the training sets of each subset and evaluated on each test set. As the GoogleApps subset is 10-100 times larger than the other subsets, using all the training data to train a unified model would suffer from the data imbalance issue (Zhang et al., 2022). Therefore, we only use 10% training data of GoogleApps. At the same time, the overall computation cost can also be saved by 80%. We use Auto-UIunified{}_{\text{unified}} as the default model for analysis unless otherwise stated.

5 Main Results

Table 2 shows the main results. Auto-UIunified{}_{\text{unified}} achieves the best overall performance compared with all the baselines. When compared with separate (not unified) models, Auto-UIunified{}_{\text{unified}} shows general effectiveness across various task scenarios. The results show that a unified multimodal model out of first principles thinking can serve as a strong autonomous agent. Compared with previous BC models, Auto-UIunified{}_{\text{unified}} has two major advantages. First, Auto-UIunified{}_{\text{unified}} is a unified model that can be adapted to different scenarios without the need to train specific models for each task. Second, Auto-UIunified{}_{\text{unified}} does not need additional annotations (screen parsing) and is easy to use. We will provide a more detailed analysis of the generality of computation efficiency in Section 5.2 and 5.4.

The ablation study in Table 3 verifies that both the chain of actions and coordinate normalization contribute to the overall performance (+5.74% and 4.04%, respectively). We set the maximum numbers of the previous actions and future actions to 8 and 4, respectively. The choice is made according to our analysis on the General subset with Auto-UIseparate{}_{\text{separate}} (Figure 3). The model under those setups achieves the optimal performance and both the input and output sequence lengths would not exceed the model limit.

For the LLMs, using either prompting or fine-tuning techniques does not achieve competitive performance compared with the other approaches. The most plausible reason is that they learn from the parsed HTML elements of the screen so that they may suffer from information loss compared with more informative vision features of the screens.

It is reasonable that Auto-UIunified{}_{\text{unified}} performs relatively inferior to BC-history on the two App-centered subsets, Install and GoogleApps, because we only use 10% training data of GoogleApps considering the data balance and computation overhead. We observe that the performance does not improve when we use all the training data of GoogleApps, possibly due to the data imbalance issue (Zhang et al., 2022). In contrast, our separate model Auto-UIseparate{}_{\text{separate}} can achieve better performance than BC-history, showing that our approach is better than BC-history under the same training setting. As we aim to study a simple and unified approach that achieves generally strong performance, we leave the treatment of the data imbalance issue in future work.

Analysis

To dive into the capability of Auto-UI, we calculate the click region accuracy, scroll direction accuracy, action type accuracy, and typed text accuracy. Figure 4 presents the results. We see that Auto-UI achieves over 90% action type accuracy on average. In contrast, the major challenges lie within the click region and scroll direction predictions. Although the model is able to predict the right action most of the time, it tends to click a wrong place or scroll in a wrong direction. The result reveals a future direction of improving the model’s ability to understand the screen layouts, e.g., using more advanced vision features.

2 Generalization Ability

As our approach is designed under first principles thinking and does not rely on pre-defined internal APIs, it could be easily generalized to new task domains. To verify the generality, we evaluate the performance of Auto-UIseparate{}_{\text{separate}} on each subset in Figure 5. For example, we train an Auto-UIseparate{}_{\text{separate}} model on the training set of General and then test its performance on the tests of each subset. We see that our approach is able to achieve a decent performance though the domains vary. This result reveals that the model could capture general knowledge for the UI control task; thus is applicable to different domains. In addition, the unified model Auto-UIunified{}_{\text{unified}} can serve as a potential choice in real-world applications owing to more coverage of training data.

3 Comprehensive Analysis

Here we present a comprehensive analysis of the choice of pre-trained features and model scale. The results are summarized in Table 4.

∙\bullet Pre-trained Features. There are two kinds of pre-trained features used in this work, the vision features and language model weights. For vision features, we compare two popular types, CLIP (Radford et al., 2021) and BLIP-2 (Li et al., 2023). We observe that BLIP-2 achieves relatively better performance. Therefore, we use BLIP-2 by default in Auto-UI. For pre-trained language model weights, we compare initializing the model with the vanilla T5 (Raffel et al., 2020), FLAN-T5 (Chung et al., 2022), and FLAN-Alpaca (Taori et al., 2023a) weights under the large size. We see that FLAN-Alpaca achieves the best performance as it has been optimized with Stanford Alpaca synthetic instruction tuning data.

∙\bullet Model Scale. Compared with the performance gains from our technique components (chain of actions and coordinate normalization) in Table 3, the benefit of scaling parameter size becomes relatively marginal. As we observe that a larger model size does not lead to significant improvement in performance, we do not scale the model scale but focus on the base (220M) and large (770M) models in this work. In addition, our choice is also based on other considerations, including the constriction of GPU memory and computation budget.

4 Computation Cost

Table 5 compares the inference speed and GPU memory cost for Auto-UI and Llama 2. Auto-UI is able to achieve nearly real-time inference (within less than one second for an action prediction) with less than 10GB GPU memory. The inference speed is over 10 times faster than Llama 2. Our work shows the strength of the medium-sized language model in building autonomous agents, which is able to achieve competitive performance with fast inference.

Conclusion

This work presents an autonomous UI agent called Auto-UI that can interact in a multimodal UI environment without environment parsing or application-dependent API access. In addition, we propose a chain-of-action technique that leverages the previously executed actions and future action plans to help the agent decide what action to execute. Experimental results show that Auto-UI achieves superior performance to previous prompting-based and fine-tuning baselines. Besides the strong performance and generality across domains, Auto-UI can infer an action as fast as within less than one second.

Acknowledgements

We thank Christopher Rawles for providing dataset and baseline details for the AITW benchmark.

References

Appendix A Appendix

We show the task examples from the AITW benchmark dataset (Rawles et al., 2023). Figures 6-10 show the examples in each subset, i.e., General, Install, GoogleApps, Single, and WebShopping. The gold actions for each screen are depicted in the illustrations for reference.

A.2 Coordinate Normalization

A.3 LLM Prompt

We use the following prompt for PaLM 2-CoT and ChatGPT-CoT due to its optimal performance reported in Rawles et al. (2023).

A.4 Using Screen Descriptions

We are interested in whether Auto-UI can be further improved when screen annotations are available. Therefore, we incorporate screen descriptions containing icon and text information, organized in HTML syntax, into our language input XlanguageX_{\textrm{language}}. Detailed examples of screen descriptions can be found in the “Screen” section in A.3.

In Table 7, we see that Auto-UI can perform better when the annotated screen descriptions are available. The results show that there is still room for performance gains for Auto-UI. However, as the annotations are not always available in real-world applications, we do not include them by default in our framework.