Multimodal Web Navigation with Instruction-Finetuned Foundation Models
Hiroki Furuta, Kuang-Huei Lee, Ofir Nachum, Yutaka Matsuo, Aleksandra Faust, Shixiang Shane Gu, Izzeddin Gur
Introduction
Web navigation is a class of sequential decision making problems where agents interact with web interfaces following user instructions (Shi et al., 2017; Liu et al., 2018; Gur et al., 2019). Common web navigation tasks include, for example, form filling (Diaz et al., 2013), information retrieval (Nogueira & Cho, 2016; Adolphs et al., 2022), or sending emails via a sequence of interactions with computer interface such as click or type (Figure 1). Recently, there has been a growing interest in developing agents to automate these actions and free humans from repetitive interactions (Mazumder & Riva, 2020; Li et al., 2020; Shvo et al., 2021).
Most prior works studied web navigation problems as online RL to learn the optimal action distribution with task-specific models from scratch (Liu et al., 2018; Gur et al., 2019; Jia et al., 2019; Humphreys et al., 2022). However, online RL requires massive trials-and-errors and is often infeasible in practice since the failure in web navigation would result in undesirable consequences; for instance, wrong password may lead to account freeze, and sending email to the wrong person could be problematic in a business scene. In contrast, offline training from the static dataset enables safe development of web agents, but the performance has been sub-optimal compared to those online RL counterparts (Humphreys et al., 2022; Gur et al., 2022). Furthermore, many of the prior works was unable to leverage rich out-of-domain data for generalization, as they usually use specialized models to explicitly handle the hierarchical structures of document object model (DOM) and their dependencies, for example, with LSTM (Gur et al., 2019; 2021), self-attention (Liu et al., 2018), or GNN (Jia et al., 2019). And many of them only output a fixed set of categorical actions (Humphreys et al., 2022), which is unfavorable for truly open-ended web navigation in the real world.
Recently, foundation models (Bommasani et al., 2021), especially large language models (LLM) (Brown et al., 2020; Chowdhery et al., 2022), have demonstrated superior performance in commonsense, symbolic, arithmetic, and multi-step logical reasoning (Wei et al., 2022b; c; Kojima et al., 2022). These models enable transformative generalization and are capable of solving wide ranges of interactive decision making problems in the wild, including but not limited to task planning in robotics (Huang et al., 2022a; b; Shah et al., 2022; Ahn et al., 2022), board game (Meta Fundamental AI Research Diplomacy Team et al., 2022), web-based retrieval and browser crawling (Nakano et al., 2021; Yao et al., 2022b; Zaheer et al., 2022).
In this work, we leverage pre-trained vision and language foundation models and introduce a competitive offline learning recipe for autonomous web agents: First, we hypothesize that grounded spatial understanding is important for web navigation (Humphreys et al., 2022; Toyama et al., 2021) and thus enables our agent to observe both HTML and screenshots by combining a language model and a ViT (Dosovitskiy et al., 2020), from semantically rich multimodal tokens that perceive local and temporal information. Second, we observe that web navigation tasks are by nature instruction-following and thus base the language model on an instruction-tuned LLM (Wei et al., 2022a; Chung et al., 2022; Ouyang et al., 2022; Iyer et al., 2022) instead of self-supervisedly pre-trained LLMs (Raffel et al., 2020; Brown et al., 2020) as in Gur et al. (2022). Third, we collect a large multimodal corpus, with both HTML and screenshots, to finetune the language model and ViT jointly. Fourth, our model outputs action in free-form text. These four key pieces together give us a multimodal web agent, which we call Web navigation via Grounded Understanding Models or WebGUM in short. As shown in Figure 1, our model takes in a command for a web-based task via a natural language instruction (e.g., in an email client, Find Gisele’s email and forward it to Siana, please.) and uses multimodal observations of the computer interface to complete the task via a sequence of computer actions.
On MiniWoB++ (Shi et al., 2017; Liu et al., 2018), a simulated web navigation environment benchmark, WebGUM outperforms previous best offline approaches trained with HTML inputs (Gur et al., 2022) by 45.8%, and even the best existing online RL approaches (Humphreys et al., 2022), despite being trained fully offline with much fewer experiences. WebGUM also shows better performance than humans and private-LLM-based agents (Kim et al., 2023; Sun et al., 2023). We perform extensive ablations and analysis in Section 5 to demonstrate WebGUM’s advantages in (1) temporal and local multimodal perception, (2) dataset and model size scaling, (3) better HTML understanding, and (4) ability of multi-step reasoning. WebGUM grounds vision and HTML understanding on the computer interface, which is critical for solving multi-step tasks with dynamic page transitions or tasks that require visual contexts, such as booking flights (+50%), shape recognition (+22%), or crawling social media (+21%). Using instruction-finetuned language models (Chung et al., 2022), compared to using vanilla models (Raffel et al., 2020), improves the success rate on MiniWoB++ by 25%, and is especially adept at handling the unknown composition of the tasks or out-of-distribution HTML inputs synthesized with realistic perturbations. On the WebShop benchmark (Yao et al., 2022a), we demonstrate that the capability of multi-step reasoning (Wei et al., 2022c) in language models enables better performance than existing state-of-the-art few-shot PaLM-540B (Yao et al., 2022b; Chowdhery et al., 2022), while our model only has 3 billion parameters. WebGUM exhibits strong positive transfer to the real-world action prediction tasks on the Mind2Web while surpassing GPT-4. Finally, we collect 347K multimodal expert demonstrations on MiniWoB++, 38 times larger than the existing unimodal dataset (Liu et al., 2018), and make these publicly available for future research https://console.cloud.google.com/storage/browser/gresearch/webllm. We believe that incorporating foundation models for efficient offline training is a scalable approach towards real-world web automation where online interactions are prohibitively costly.
Related Work
Web Navigation Among many proposed benchmarks for autonomous web navigation (Toyama et al., 2021; Burns et al., 2022; Yao et al., 2022a), one of the most inclusive and representative benchmark to test the capability of autonomous agents is MiniWoB++ (Shi et al., 2017; Liu et al., 2018), which consists of a set of simulated websites with various user instructions from primitive tasks to complex multi-step decision making tasks, such as sending emails or booking flights. Prior works have tried to solve this benchmark using a variety of techniques; Liu et al. (2018) and Gur et al. (2019; 2021) leverage the guidance during online RL from high-level workflow (Liu et al., 2018) or curriculum learning (Gur et al., 2019; 2021), which should be, however, designed per task, and then would not be scalable methods. Other approaches have employed supervised learning (SL) with a large million-scale dataset and following RL-finetuning (Humphreys et al., 2022), or SL with LLM-based agents (Gur et al., 2022). Offline SL agents often suffer from sub-optimal behavior, and online RL with tremendous exploratory experiences has been critical for proficient navigation on the web (Humphreys et al., 2022), which is, however, difficult to conduct in real websites as there is typically no reward signal and interactions are prohibitively costly. As shown in Appendix I, many of these approaches depend on task-specific hierarchical structures of DOM (Jia et al., 2019; He et al., 2020), tailored architectures to encode their dependencies such as LSTM (Gur et al., 2019; 2021), self-attention (Liu et al., 2018), or GNN (Jia et al., 2019), and task-dependent categorical output space (Humphreys et al., 2022), which could not handle open-ended multi-task settings similar to real world, or incorporate pre-trained models. In contrast, we remove such web-specific architectures and convert web navigation into visual question-answering format (text, image text), which allows us to leverage pre-trained foundation models (Chung et al., 2022; Dosovitskiy et al., 2020) as rich prior knowledge on the web, and then to learn the capable agents even with offline training.
Large Language Models for Web Navigation Concurrently, private-LLM-based agents, such as InstructGPT (text-davinci-003) (Ouyang et al., 2022) and GPT-3.5-turbo, have achieved competitive performance to RL-fintuned models and humans by leveraging a handful of few-shot demonstrations with self-improvement (Kim et al., 2023), code generation (Sun et al., 2023), and structured prompts (Zheng et al., 2023). In contrast, WebGUM focuses on multimodality and finetuning with domain-specific data. With those, we show very competitive performance compared to PaLM-540B with only 3 billion parameters. WebGUM can also handle long HTML observation tasks, such as book-flight or choose-date-hard, where agents that rely on in-context few-shot learning tend to run out of input tokens. In addition, our models do not requires ad-hoc prompt engineering.
In Appendix B, We discuss additional related works on multimodal large-scale models and foundation models for decision making.
Preliminaries
We formulate autonomous web navigation as a deterministic sequential decision making problem; composed of a state space , action space , deterministic transition function , instruction space , reward function (or episodic success criteria) . At each time step , the agent follows a parameterized policy conditioned on previous states and actions , and transits to the next state: . This process continues until the agent reaches the terminal state (e.g. Submit button is clicked) or the max time step is exceeded. An episode is treated as a success if given instruction is satisfied (i.e. ), and as a failure if the agent takes a invalid action or reaches a wrong terminal state.
In autonomous web navigation, the state is a web page consisting of the raw HTML as a text sequence and a screenshot as an image. Following prior works (Shi et al., 2017; Liu et al., 2018; Gur et al., 2019; 2021), we assume the constraint action space: function(selector, text). function is either click or type, selector is an integer index that can uniquely specify the element, and text is a text input for type function.
Figure 1 presents an example episode of MiniWoB (Shi et al., 2017), which involves multi-step decision making. To meet the given instruction, the agent clicks an email from the proper sender and types the correct receiver to forward that email. MiniWoB also has primitive behavioral tasks such as clicking buttons or entering texts. For the examples of WebShop (Yao et al., 2022a), see Appendix L.
WebGUM
In this work, we follow Gur et al. (2022) to use T5 (Raffel et al., 2020), an encoder-decoder architecture, for HTML-based web navigation, as its bi-directional nature could be a good fit for the tree structure of HTML and the architecture has been shown to scale well. We combine T5 with a vision transformer (ViT) (Dosovitskiy et al., 2020) for multimodality as illustrated in Figure 2. Specifically, we use the ViT to map image observations (screenshots) into image tokens. The ViT is pre-trained on ImageNet-21K classification (Deng et al., 2009). The T5 encoder then consumes both visual and HTML tokens in a unified manner, and the decoder predicts actions in text. See Appendix C for more implementation details.
Encoding Temporal and Local Visual Tokens For language models to be aware of task temporal information and local scene recognition, the encoder considers multimodal tokens extracted from a history of patched screenshots ( steps). Temporal visual tokens contribute to predict the consistent actions in a multi-step tasks. To better extract spatial and semantic information across the local parts of websites, our ViT encodes one local token per patch rather than global one per image (i.e. CLS-token). We divide an input image into patches – giving a total of visual tokens. We crop the screenshots of MiniWoB++ to remove the yellow instruction part, and the image size becomes 160 160. We pad cropped images with white pixels to fit them into 224 224; the default input size for ViT.
2 Instruction-Finetuned Large Language Models
We base our language model on Flan-T5 (Chung et al., 2022), an instruction-finetuned T5, as opposed to using a vanilla pre-trained T5 as in Gur et al. (2022). Flan-T5 is finetuned with large-scale instruction-following format problems and chain-of-thought examples across a variety of domains, including reasoning or programming. Considering that web navigation is inherently an instruction-following task, we hypothesize that carefully trained instruction-finetuned models could generalize well to enhance the alignment with user instruction and zero-shot reasoning in the web-navigation, interactive decision making context. For the same reason, we also hypothesize that these high-performing instruction-finetuned models enable better sample efficiency and downstream performance, and thus are well-suited for offline learning. We further finetune the Flan-T5 language model and the ViT vision encoder jointly (Figure 2) on a large corpus of instruction-following multimodal web navigation data, which we describe in Section 4.3. In Section 5, we empirically demonstrate that this instruction-finetuned recipe improves HTML comprehension, multi-step reasoning and decision making significantly.
3 Large-scale Data Collection with Language Model Agents
Recent successes of foundation models are largely powered by internet-scale data (Brown et al., 2020; Radford et al., 2021; Chen et al., 2022; Wang et al., 2023). While large amount of data is critical, for web navigation domain, there is only a small public dataset for MiniWoB++, consisting of 12K episodes of human demonstration (Liu et al., 2018). Moreover, the dataset only consists of DOM observations and lacks any visual features, which might limit the fine spatial perception of the elements on the page. A large-scale multimodal dataset, including screenshots of websites, is required to build a better navigation policy at scale.
To collect a huge amount of multimodal behavioral dataset on MiniWoB++, we leverage the finetuned-LLM policy from Gur et al. (2022), instead of human demonstrators (Liu et al., 2018; Humphreys et al., 2022). This significantly reduces the cost to construct a new dataset by leveraging the prior success of autonomous agents. We first rollout a LLM policy with 100 episodes per task, which results in a 2.8K successful episodes. Then, we finetune Flan-T5-XL models with this small dataset and run those with 10,000 episodes per task. Lastly, we collect additional 54K demonstrations with Synapse (Zheng et al., 2023), a private-LLM-based agents with prompting, for the tasks where the finetuned-LLM may not complete well. Such efforts result in a multi-task dataset with 401K (347+54K) episodes including HTML and screenshots at each step. See Appendix F for more details.
Results
We test our method on MiniWoB++ (Shi et al., 2017; Liu et al., 2018) with 100 evaluation episodes per task, taking the average success rate over 56 tasks taken from Gur et al. (2022). Table 1 shows that WebGUM, with a small 2.8K dataset and Base-size model (310M parameters), significantly outperforms previous offline methods for web navigation (Humphreys et al., 2022; Gur et al., 2022). While they used 2.4 million episodes or 3 billion parameters, WebGUM could improve the data and parameter efficiency to achieve superior performance in offline regime, which is realized by the problem simplification of web navigation in order to leverage temporal-local visual perception and instruction-finetuned LLMs as strong inductive bias on web environments. In addition, scaling dataset and model size, WebGUM achieves 94.2% success rateVideos are available at https://sites.google.com/view/mm-webnav/, exceeding the previous best offline model, WebN-T5 (Gur et al., 2022), by over 45.8% and even surpassing the online RL-finetuned SoTA, CC-Net (Humphreys et al., 2022) (+0.7%), despite our fully offline training and much fewer data. Moreover, WebGUM surpasses humans and recent LLM-based agents, such as RCI (Kim et al., 2023) and AdaPlanner (Sun et al., 2023), even with GPT-4 (OpenAI, 2023). The per-task comparison and error analysis (Appendix G, L) imply that there is room for improvement in complex reasoning tasks requiring memory such as guess-number.
In the following sections, we perform extensive and precise ablations of WebGUM to clearly identify the source of improvement. Especially, we will demonstrate the contribution of (1) temporal and local multimodal perception (Section 5.1), architectures and pre-trained models, and (2) dataset and model size scaling (Section 5.2). We will also point out (3) better HTML comprehension (Section 5.3) and (4) capability of multi-step reasoning (Section 5.4) from instruction-finetuned LLMs. Furthermore, we prove that WebGUM can be transferable to the real-world tasks (Section 5.5).
To verify the importance of image modality, we design three ablations: (i) input replacement, (ii) removing visual perception tokens, and (iii) employing different pre-trained ViT. We first replace image observations with completely white images, and with randomly sampled MiniWoB++ screenshots taken in the initial states at test time. For visual token and pre-trained ViT ablations, we prepare various pre-trained weights with ImageNet-21K (IN) + AugReg (Steiner et al., 2022), JFT-300M (Sun et al., 2017), or JFT-3B (Zhai et al., 2022), and with self-supervised objectives such as CLIP (Radford et al., 2021), MAE (He et al., 2021), or DINO (Caron et al., 2021), and then finetune Base-size models as a proxy of larger-size models (Hoffmann et al., 2022) to reduce the computational costs.
In Figure 3 (left), the performance of the model with white images is comparable to the unimodal model. Presumably because the model with randomly-taken images may accidentally contain the images from the target task, WebGUM (random) slightly surpasses WebGUM (white). These results prove WebGUM successfully obtains grounded vision and HTML understanding by leveraging temporal and local fine perception. In the visual token ablation, Figure 4 (left) shows that combining both temporal and local visual tokens (66.1%) improves the performance than temporal (64.2%) or local tokens only (64.0%). Interestingly, the effects of different pre-trained ViT are marginal, compared to visual tokens, which highlights our contribution on designing suitable architecture for multimodal web navigation.
We also compare per-task performance gaps caused by adding vision modality to language models. Figure 3 (right) presents top-10 absolute performance improvement, suggesting WebGUM leverages visual inputs for multi-step tasks with dynamic page transitions (e.g. book-flight; +50%) or tasks requiring visual context understanding (e.g. click-shape; +22%) (see Appendix G and L).
2 Scaling Effect in Dataset and Model Size
In this section, we show the importance of scaling up the dataset and model size in WebGUM, similar to the observations in the language and vision domain (Shoeybi et al., 2019; Kaplan et al., 2020; Rae et al., 2021; Wei et al., 2022b; Chowdhery et al., 2022). To investigate data scaling, we prepare three dataset: minimal 2.8K demonstrations, 347K demonstrations, and its 20%-size demonstrations (68K), and then finetune Flan-T5-Base with them. Figure 4 (middle) proves that increasing dataset size leads to the improvement of success rate. Because multimodal models benefit from the scaling more, the larger dataset size might be more crucial in multimodal models, which also supports our attempts to construct large-scale multimodal dataset for web navigation. Notably, Base-size WebGUM with 2.8K episodes already achieves 55.7%/66.1%, surpassing previous best SL models (49.8%/55.6% we trained with 347K episodes). This surprising data efficiency comes from the sufficient inductive bias and alignment with the user intentions in instruction-finetuned LLMs.
In addition to dataset size, Figure 4 (right) shows that the performance of WebGUM improves as the number of parameters in T5 model increases from Base (220M) to XXL (11B). These results also reveal that scaling the models might be more important than the dataset; the low-capacity model may cap the performance at a lower level. In contrast, decoder-only Flan-PaLM-8B only achieves 72.8% success, comparable to WebGUM-Large (770M), which emphasizes the advantage of encoder-decoder models in web navigation. See Appendix D for further details.
3 Better HTML Comprehension from Instruction-Finetuned LLMs
We have demonstrated that instruction-finetuned LLMs outperforms vanilla LLMs in web navigation. To analyze the effect of instruction-finetuning more precisely, we here focus on the capability of HTML understanding. Since instruction-finetuned LLMs perform well on many NLP tasks with content comprehension (Chung et al., 2022; Iyer et al., 2022), web navigation should also benefit from them. As a test bed for HTML comprehension, we investigate (1) generalization to unseen compositions of known tasks, and (2) robustness to the realistic input perturbations, which are also important challenges for the web agents to be deployed on the real-world internet. We also provide the base language model comparison on a standard HTML comprehension benchmark, WebSRC (Chen et al., 2021d) in Appendix E, where Flan-T5 achieves better EM/F1 scores than T5 after finetuning.
For the compositional tasks, we pick up 4 click-“something” (link, button, checkboxes, dialog) tasks and make 6 combinations of these by naively stitching with 2 or 3 tasks (e.g. Figure 5). See Appendix H for further details. The results show that WebGUM with HTML and image inputs outperforms prior finetuned-LLM (Gur et al., 2022) and Synapse (Zheng et al., 2023), a SoTA LLM agent in MiniWoB++, which implies WebGUM has obtained better reading skills for web navigation and could transfer them to handle unseen HTML in compositional tasks robustly.
To test the robustness against input corruptions, we test three different realistic perturbations; adding extra HTML at the top or bottom of the original HTML, and adding attributes of coordinates (left, right, top, bottom; they are unrelated to solving the tasks) in each element of HTML at test time. These perturbations often happen in the real world due to the renewal or API changes, not to mention unknown websites, but rule-based pre-processing may not fully cover them. The results show that while all the methods are affected by the input corruptions to some extent, WebGUM, with both HTML and HTML plus image modalities, achieves significantly better performances than Gur et al. (2022). Notably, WebGUM outperforms prior finetuned LLM (+ 56.2% in multimodal and +33.4% in unimodal models) even when extra distracted attributes are added to HTML. They support our hypothesis: instruction-finetuning imporves HTML comprehension in LLMs, which enables the downstream agents to deal with out-of-distribution inputs or tasks robustly.
4 Ability of Multi-Step Reasoning as a Prior for Interactive Decision Making
Another notable feature in instruction-finetuned LLMs is an ability of multi-step reasoning (Chung et al., 2022). We hypothesize this reasoning capability would play an important role as a prior for interactive decision making. To decouple the evaluation of reasoning capability from visual page perception, HTML understanding, and the benchmark simulator (MiniWoB++), we extensively evaluate our WebGUM on WebShop (Yao et al., 2022a), another online-shopping website simulator with a large amount of real-world product data. Because it requires complex multi-step decisions considering previous contexts for item comparison, WebShop is suitable for investigating the capability of multi-step reasoning from instruction-finetuned LLM in depth (Yao et al., 2022a; b). WebShop provides a user instruction that describes the features of item (e.g. I need a long clip-in hair extension which is natural looking, and price lower than 20.00 dollars). The agents should search, compare and choose a proper product that matches the given instruction. The performance score is evaluated by the percentage of required attributes covered by the chosen product, and if the product meets all the requirements, that episode is labeled a success. See Appendix K for further details.
Table 2 shows that WebGUM achieves 45.0% success, significantly outperforming not only simple baselines, such as supervised imitation learning (IL), IL plus RL-finetuing and WebN-T5 (by more than 15%), but also recent prompt-based LLM agents, including ReAct (Yao et al., 2022b) (i.e. PaLM-540B (Chowdhery et al., 2022) with one-shot prompt and reasoning annotations), while our model only has 3 billion parameters. Due to the consistent reasoning and enhanced alignment with user intentions, WebGUM could compare the products with backtracking, and choose proper options (see Appendix L). Our results imply that ability of multi-step reasoning in Flan-T5 works as strong and transferable prior knowledge for downstream decision making.
5 Strong Transfer to Real-World Action Prediction
Lastly, we demonstrate the applicability of WebGUM to real-world problems. We test WebGUM on Mind2Web (Deng et al., 2023), a real-world demonstration dataset with about 2K instructions on 137 websites. In the action prediction tasks, we transfer WebGUM finetuned for MiniWoB++ with 401K dataset into real-world Mind2Web by further finetuning with the training set. WebGUM takes top-50 relevant HTML snippet candidates, instructions, and action history as inputs and outputs next actions by predicting the element id, operations (e.g. click, type), and values. Table 3 reveals that WebGUM, transferred from MiniWoB, achieves superior performance to MindAct-Large/XL and even GPT-4 in all the categories (cross-task/website/domain). Because both MindAct and WebGUM are based on Flan-T5, these results support that WebGUM exhibits strong positive transfer to real-world tasks.
Discussion and Limitation
Throughout the paper, we present an effective and practical methodology to simplify web navigation into offline training in order to leverage the inductive bias of web environments in instruction-finetuned LLMs. While WebGUM exhibits positive transferability to real-world problems in Mind2Web, we leave it as future work to scale multimodal foundation models into the deployment for real-world web navigation (Gur et al., 2023).
We collect and release a multimodal expert dataset with 347K episodes on MiniWoB++. However, this is still far from internet-scale dataset that is necessary for generalist models. Collecting behavioral data at scale by iterative data-collection and deployment (Ghosh et al., 2021; Matsushima et al., 2021; Li et al., 2022a) might be a key for practical interactive agents. Since our approach – taking raw HTML and screenshots as inputs and predicting parsable actions in text – only has minimal assumptions which constraint model architectures, it might be applicable to any advanced LLMs or open-ended situations. While WebGUM could deal with out-of-distribution compositional and perturbed tasks in a robust manner, human-level broader generalization to the diverse real websites or instructions is still a hard problem to be resolved.
Conclusion
We develop Web navigation via Grounded Understanding Models (WebGUM), learning an instruction-following visual language foundation model for web navigation. WebGUM significantly improves the success rate on MiniWoB, compared to previous offline-trained SoTA from 48.4% to 94.2%. Our detailed ablations show that temporal and local visual tokens capture dynamic transition and visual context of the page, and that instruction-finetuned language models significantly improves web navigation performance due to the better HTML comprehension and capability of multi-step reasoning. Multi-step reasoning enables more robust generalization to out-of-distribution tasks, and outperforms PaLM-540B in WebShop. WebGUM also demonstrates strong positive transfer to real-world action prediction tasks in Mind2Web. Furthermore, we scale the existing MiniWoB dataset into multimodal 347K expert demonstrations, about 38 times larger than before. We believe that our work is an significant step towards building more capable and scalable models for autonomous web navigation.
Acknowledgements
HF was supported by JSPS KAKENHI Grant Number JP22J21582. We thank Yusuke Iwasawa, Mustafa Safdari, Austin Huang, Heiga Zen for helpful feedback on this work, and Shunyu Yao for setting up WebShop experiments.
References
Appendix
Appendix A Broader Impacts
While WebGUM is evaluated only in realistic web simulators (Shi et al., 2017; Liu et al., 2018; Yao et al., 2022a), we should carefully conduct it if we deploy the autonomous web agent on the real-world Internet because of security and safety reasons. For instance, the wrong password may cause an account freeze, and emailing the wrong person is problematic in a business scene. Training with online RL may often be infeasible for this reason, while we demonstrate an alternative approach; data-driven, fully offline training by leveraging inductive bias in foundation models. Autonomous agents, well-grounded with the user’s intention, should be helpful in our daily lives by reducing our burden on computer tasks. Because a part of our training corpus (54K) includes the demonstrations taken from the output of LLMs (Anil et al., 2023), we will exclude those from the dataset release and it will result in 347K episodes.
Appendix B Extended Related Works
Foundation Models for Decision Making Recently, the ability of multi-step reasoning and inductive bias in foundation models have been leveraged to solve text-based interactive tasks via sequential decisions considering few-shot in-context examples (Ahn et al., 2022; Huang et al., 2022a; b; Zeng et al., 2022; Yao et al., 2022b; Meta Fundamental AI Research Diplomacy Team et al., 2022). Even in continuous control (Chen et al., 2021a; Janner et al., 2021; Furuta et al., 2022b; Brohan et al., 2022) or computer games (Reed et al., 2022; Lee et al., 2022b; Fan et al., 2022), high-capacity transformer models are trained with a large amount of diverse dataset via multi-task behavioral distillation (Chen et al., 2021c; Gu et al., 2021a; DeepMind Interactive Agents Team et al., 2021; Furuta et al., 2022a; Shridhar et al., 2022; Jiang et al., 2022). To build autonomous web navigation agents, we also leverage pre-trained LLM (Raffel et al., 2020; Chung et al., 2022), by finetuning with massively-curated multimodal demonstrations, and we point out that the better content comprehension and multi-step reasoning abilities, obtained through instruction-finetuning of LLM (Chung et al., 2022), are essential for the notable performance on downstream decision making aligned with human instructions.
Multimodal Large-scale Models Large language models have demonstrated extraordinary emergent abilities on a variety of NLP tasks, such as commonsense question answering, arithmetic, logical reasoning, open-ended text generation (Radford et al., 2019; Brown et al., 2020; Chowdhery et al., 2022; Wei et al., 2022b; Tay et al., 2022), or code completion (Chen et al., 2021b; Austin et al., 2021; Li et al., 2022b). In addition, some works have investigated vision-and-language understanding to improve the accuracy of common vision-based tasks such as open-ended image/object classification (Radford et al., 2021; Gu et al., 2021b; Kamath et al., 2021), image captioning, or visual question-answering (Lu et al., 2022; Alayrac et al., 2022; Chen et al., 2022; Reed et al., 2022; Liu et al., 2023; Dai et al., 2023; Li et al., 2023). Several works also have tackled document understanding with (multimodal) transformer models (Xu et al., 2019; Li et al., 2021a; c; Appalaraju et al., 2021; Tang et al., 2022; Wang et al., 2022a; b), including markup languages such as HTML (Aghajanyan et al., 2021; 2022; Li et al., 2021b; Lee et al., 2022a) for summarization of the documents or question answering on the contents. Despite the great efforts on document understanding, these works are less connected to interactive decision making problems. Our model obtains not only a grounded understanding of websites in a multimodal manner but also the ability to decide the optimal actions to achieve given instructions in web navigation, helping multi-step decisions and visual context understanding.
Appendix C Implementation Details
We adopt the encoder-decoder models proposed by Raffel et al. (2020) as multimodal transformers, and vision transformer (Dosovitskiy et al., 2020) pre-trained with ImageNet-21K (Deng et al., 2009) as an image encoder for the visual tokenshttps://github.com/google-research/scenic. We especially use ViT-B16, a small-size transformer with 86 million parameters, which divides an input image into -size patches. We use publicly available checkpoints of T5 (Raffel et al., 2020)https://github.com/google-research/t5x/blob/main/docs/models.md#t5-11-checkpoints, Flan-T5 (Chung et al., 2022)https://github.com/google-research/t5x/blob/main/docs/models.md#flan-t5-checkpoints, and T5-XL finetuned with MiniWoB++ demonstrations (Gur et al., 2022)https://console.cloud.google.com/storage/browser/gresearch/webllm/webn_t5_3b for the experiments. To construct the training pipeline, we leverage SeqIO (Roberts et al., 2022) library, and use SentencePiece (Kudo & Richardson, 2018) vocabulary with 32K tokens from C4 dataset (Raffel et al., 2020) for text tokenization. The batch size for training is 128, and input sequence length is set to 4096 tokens. Due to the huge computational requirements, we run one seed to train each model throughout the paper (Humphreys et al., 2022; Gur et al., 2022). We use cloud TPU-v4, which has a 32 GiB HBM memory space for the experiments. Base-size models require 256 cores and XL-size models do 512 cores, which takes 1-2 days for finetuning.
Appendix D Details on Dataset and Model Size Scaling
We here present how critical it is to scale up the dataset and model size in WebGUM. For the dataset size ablation, we use Flan-T5-Base and ViT-B16. As for both HTML and multimodal models, we could observe the scaling effects in web navigation: the larger the dataset (Table 4) and model (Table 5) size are, the higher the success rates are. Surprisingly, our approach even with only 2.8K HTML episodes (about 25% of the previous one curated by Liu et al. (2018)) and Base-size model (about 7.3% parameters) already achieves 55.7%, surpassing previous SL state-of-the-art (48.4% by Gur et al. (2022)). This surprising efficiency might come from the sufficient inductive bias and alignment with the user intentions in instruction-finetuned LLMs, and WebGUM could fully leverage them for web automation problems. The margin of improvement might be smaller than expected due to the limited capacity of transformer to obtain the grounded understanding of natural language instructions, HTML, and screenshots. In fact, the results also reveal that scaling the models might be more important than the dataset; the low-capacity model may cap the performance at a lower level.
Appendix E WebSRC
We extensively evaluate the capability of HTML comprehension in instruction-finetuned LLMs with WebSRC (Chen et al., 2021d) where the models are asked to solve contextual QA problems understanding a given HTML and its structure. Those problems are curated from real websites to include key-value extraction, entity comparison, and table understanding problems. The answer formats are either text span in HTML or binary (yes/no). Because the context length is insufficient for raw HTML, we preprocess context HTML by extracting a snippet that includes the answers in advance. We finetune both T5-XL and Flan-T5-XL with the training dataset. Table 6 shows that Flan-T5 records better HTML comprehension performance than T5, which may accelerates the web navigation performance on MiniWoB++ and Mind2Web.
Appendix F Dataset Details
To construct a large-scale multimodal behavioral dataset on MiniWoB++, we leverage a public finetuned-LLM policy (Gur et al., 2022) trained with multi-task human demonstration dataset (Liu et al., 2018)https://github.com/stanfordnlp/miniwob-plusplus-demos as a demonstrator. We run such LLM policies with 10,000 episodes per task and only keep successful trajectories to maintain the quality of dataset, following Humphreys et al. (2022). Lastly, we collect additional 54K demonstrations with Synapse (Zheng et al., 2023)https://github.com/ltzheng/synapse, a private-LLM-based agents with prompting, for the tasks where finetuned-LLM may not complete well such as click-scroll-list and enter-time, and also write a scripted policy for book-flight. We use PaLM 2 (Anil et al., 2023) as a base LLM for Synapse. Such efforts result in a multi-task dataset with 401K (347K+54K) episodes including HTML and screenshots at each time step. Table 7 shows the details of our multimodal dataset (347K), consisting of HTML, screenshots, actions, and instructions at each time step.
Appendix G Per-Task Performance of MiniWoB++
In this section, we present per-task success rate on MiniWoB++ (LABEL:tab:per_task_miniwob_results) and absolute performance improvement by adding image modality to HTML input for WebGUM (Figure 7).
As for LABEL:tab:per_task_miniwob_results, we refer to Gur et al. (2022) and Zheng et al. (2023) for the baseline performances. We use 56 tasks as benchmark, while removing some duplicated tasks (e.g. “-nodelay” tasks) from 62 tasks adopted in Gur et al. (2022). During the evaluation on MiniWoB++, we ignore the time limit due to the computational constraints.
Figure 7 presents full results of the absolute performance improvement, subtracting the success rates: (Success Rate of WebGUM(HTML+Image)) - (Success Rate of WebGUM(HTML)). The results suggest WebGUM leverages visual inputs for multi-step tasks with dynamic page transitions (e.g. book-flight or search-engine) or the tasks that require global contexts of the page (e.g. tic-tac-toe or click-shape). See Appendix L for the visualization.
Appendix H Compositional Evaluation on MiniWoB++
For the compositional evaluation, we pick up 4 click-“something” (link, button, checkboxes, dialog) tasks and make some combinations of those by naively stitching with 2 or 3 tasks. Then, we prepare the following 6 combinational tasks,
These tasks should be resolved in order of the name: for instance, in click-link_click-button_click-dialog task, the agent should click the proper link, click the proper button, click the proper dialog, and then the task results in success. In click-button_click-link task, the agent should click the proper button, and then click the proper link. The instructions for compositional tasks are also simply combined among original task instructions in order of the name. This evaluation could test the ability to transfer primitive skills to control computers to solve unseen tasks. Table 9 shows the per-task average success rate among 6 combinations above. WebGUM can solve the compositional tasks much better than baselines (Gur et al., 2022; Zheng et al., 2023) .
Appendix I Comparison against Prior Web Navigation Agents
Appendix J Input Perturbation Evaluation on MiniWoB++
Appendix K Evaluation on WebShop
In addition to MiniWoB++, we extensively evaluate our WebGUM on WebShop (Yao et al., 2022a) benchmark, another online-shopping websites simulator with a large amount of real-world product data. WebShop provides user instruction that describes the feature of items (e.g. I need a long clip-in hair extension which is natural looking, and price lower than 20.00 dollars). The agents should search, compare and choose a proper product that matches the given instruction. Since WebShop requires complex multi-step reasoning considering previous contexts for comparison (Yao et al., 2022a; b), we can test the capability of instruction-finetuned LLM in decision making tasks in depth. The performance score is evaluated by the percentage of required attributes covered by the chosen product (from 0 to 100), and if the product meets all the requirements, that episode is labeled a success.
Because WebShop does not have API to get the screenshot of rendered websites, we focus on WebGUM with text inputs, parsed from noisy HTML in the real world.WebShop just provides visual features of item pictures when the agents reach the product page. These features are extracted by ResNet-50 (He et al., 2016), rather than raw images or screenshots of the website. Some baseline agents (IL and IL+RL) incorporate such embeddings. We convert the actions from raw texts (e.g. search[a long clip-in hair extension] or click[
Table 11 shows that WebGUM achieves 45.0% success, significantly outperforming not only simple baselines, such as supervised imitation learning (IL) and IL plus RL-finetuing (by more than 15%), but also recent prompt-based LLM agents, including ReAct (Yao et al., 2022b) (i.e. PaLM-540B (Chowdhery et al., 2022) with one-shot prompt and reasoning annotations), while our model only has 3 billion parameters. IL and IL plus RL-finetuning baselines use BART (Lewis et al., 2019) model for the search policy, and BERT (Devlin et al., 2019) model for the click policy. The better performance of WebGUM proves the hypothesis that the ability of multi-step reasoning in instruction-finetuned language models works as a prior for decision making problems.