A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis
Izzeddin Gur, Hiroki Furuta, Austin Huang, Mustafa Safdari, Yutaka Matsuo, Douglas Eck, Aleksandra Faust
Introduction
Large language models (LLM) (Brown et al., 2020; Chowdhery et al., 2022; OpenAI, 2023) can solve variety of natural language tasks, such as arithmetic, commonsense, logical reasoning, question answering, text generation (Brown et al., 2020; Kojima et al., 2022; Wei et al., 2022), and even interactive decision making tasks (Ahn et al., 2022; Yao et al., 2022b). Recently, LLMs have also demonstrated success in autonomous web navigation, where the agents control computers or browse the internet to satisfy the given natural language instructions through the sequence of computer actions, by leveraging the capability of HTML comprehension and multi-step reasoning (Furuta et al., 2023; Gur et al., 2022; Kim et al., 2023).
However, web automation on real-world websites has still suffered from (1) the lack of pre-defined action space, (2) much longer HTML observations than simulators, and (3) the absence of domain knowledge for HTML in LLMs (Figure 1). Considering the open-ended real-world websites and the complexity of instructions, defining appropriate action space in advance is challenging. In addition, although several works have argued that recent LLMs with instruction-finetuning or reinforcement learning from human feedback improve HTML understanding and web automation accuracy (Furuta et al., 2023; Kim et al., 2023), their architectures are not always suitable to process real-world HTML documents; as presented in Figure 2, HTML tokens of real websites are much longer than those of simulators, and most LLMs have shorter context lengths than the average HTML tokens in real websites. It is prohibitively costly to treat such long documents as inputs directly, and even to adopt prior techniques for structured documents, such as text-XPath alignment (Li et al., 2021b) or text-HTML token separation (Wang et al., 2022a). To prioritize broad task generalization and model-size scaling, such domain knowledge for HTML codes is not applied in recent LLMs.
In this work, we introduce WebAgent, an LLM-driven autonomous agent that learns from self-experience to complete user instructions on real websites by combining canonical web actions in a program space (Figure 3). WebAgent (i) plans sub-instructions per step by decomposing natural language instructions, (ii) summarizes long HTML pages into task-relevant snippets based on sub-instructions, and (iii) acts via programming on real websites by grounding sub-instruction and HTML snippet into executable Python codes. We combine two LLMs to form WebAgent: Flan-U-PaLM (Chowdhery et al., 2022; Chung et al., 2022) for grounded code generation, and newly introduced HTML-T5, a domain-expert pre-trained language model, for task planning and conditional HTML summarization. HTML-T5 has an encoder-decoder architecture and is specialized to capture the structure – syntax and semantics – of long HTML pages better by adopting local and global attention encoder (Guo et al., 2022). It is self-supervisedly pre-trained with a mixture of long-span denoising objectives (Tay et al., 2022) on a large-scale HTML corpus from CommonCrawl. To ground language model agents into real websites, we introduce self-experience supervision, where the domain-expert language models are finetuned with self-generated demonstrations.
Existing LLM-driven agents often solve decision making tasks with a single LLM conditioned on different prompts per role (Kim et al., 2023; Sun et al., 2023; Zheng et al., 2023), which is, however, not enough for real-world tasks whose complexity is higher than that of simulators. The empirical evaluations reveal that our method incorporating self-bootstrapped specialist language models improves HTML understanding and grounding, and achieves better generalization than single LLM agent. In real-world web automation, WebAgent significantly increases the success rate by 50%, and error analysis emphasizes that coupling task planning with HTML summarization in specialized language models is essential for task success. Moreover, HTML-T5 not only works as a core module for WebAgent but also achieves strong results by itself on the web-based tasks. On MiniWoB++ (Liu et al., 2018; Shi et al., 2017), HTML-T5 achieves 18.7% higher success than previous language model agent (Gur et al., 2022) while also outperforming competitive baselines, such as naive local-global attention models (Guo et al., 2022) and its instruction-finetuned ones (Chung et al., 2022). On Mind2Web (Deng et al., 2023), an offline task planning dataset, HTML-T5 achieves SoTA performances among MindAct with FLan-T5-XL and GPT-4 (OpenAI, 2023). In summary, our key contributions are:
We propose WebAgent, integration of two modular LLMs under self-supervision for real-world web automation. The domain-expert language model handles planning and HTML summarization, and a generalist language model generates executable programs.
We newly introduce HTML-T5, pre-trained language models with local-global attentions and a mixture of long-span denoising on large-scale HTML corpus, which capture the syntax and semantics of HTML better.
WebAgent notably improves the success rate by over 50% in real websites. HTML-T5 itself outperforms prior language model agent by 18.7% in MiniWoB++, and realizes SoTA performance in Mind2Web while surpassing GPT-4.
Related Works
Web Automation Web automation is a sequential decision making task where agents manipulate browsers following given instructions (Shi et al., 2017), such as form filling (Diaz et al., 2013) or information retrieval (Adolphs et al., 2022) through the sequence of computer actions (Li et al., 2020; Mazumder & Riva, 2020; Shvo et al., 2021). Prior works have realized the web automation via reinforcement learning (Gur et al., 2019; Humphreys et al., 2022; Jia et al., 2019; Shaw et al., 2023), finetuned (Furuta et al., 2023; Gur et al., 2022) or prompted LLMs (Kim et al., 2023; Sun et al., 2023; Yao et al., 2022b; Zheng et al., 2023) on the simulated websites (Shi et al., 2017; Toyama et al., 2021; Yao et al., 2022a). However, there are still huge gaps between simplified simulators and real web environments; for instance, the average tokens for HTML pages are about 15 times larger (Figure 2), and pre-defined action space for specific websites is a strong assumption that may harm the generalization to out-of-distribution web pages or instructions.
MindAct (Deng et al., 2023) could be the most relevant work, where finetuned language model summarizes the raw HTML document into task-relevant snippets, and another model predicts the web actions in a multi-choice QA format. While MindAct also combines several language models, it has just adopted DeBERTa (He et al., 2021) and Flan-T5 (Chung et al., 2022) for summarization and actor modules, and evaluated it on the offline dataset. In contrast, we design HTML-T5, specialized for web-based tasks, to handle long HTML documents. WebAgent leverages HTML-T5 finetuned with self-experience for summarization and planning, and Flan-U-PaLM as a capable programmer, which enables it to generate open-ended web actions and to act on online real-world websites.
Program Synthesis In addition to common LLMs (Brown et al., 2020; Chowdhery et al., 2022; Touvron et al., 2023), several works have proposed programming-focused language models (Chen et al., 2021a; Feng et al., 2020; Li et al., 2022; Wang et al., 2021) and their benchmarks (Austin et al., 2021; Hendrycks et al., 2021a; Lu et al., 2021). Another line of work has investigated the tool augmentation of LLMs (Parisi et al., 2022) by decoding API calls (Schick et al., 2023) or Python snippets to be parsed with the interpreter (Gao et al., 2023). Most works deal with the program synthesis on the static dataset, except for the attempts in robotics (Liang et al., 2023) and game (Trivedi et al., 2022; Wang et al., 2023a), where LLMs output Python or JavaScript snippets to command the agents. Similarly, we leverage the ability of code generation as an open-ended action space for web-based agents to manipulate the real website, and demonstrate LLMs can sequentially decode Python selenium codes considering the given sub-instructions and HTML in the prompts.
See extended related works on document understanding and LLM for task planning in Appendix B.
WebAgent
WebAgent is composed of interactions between HTML-T5, a domain-expert language model, which predicts the sub-instruction for the next-step program and conditionally summarizes long HTML documents, and Flan-U-PaLM (Chowdhery et al., 2022; Chung et al., 2022), an instruction-finetuned LLM for grounded program synthesis (Figure 3). In contrast to a single LLM conditioned on different prompts per role, such a modular approach can deal with real-world tasks better. Moreover, to align WebAgent with real websites, we introduce self-experience supervision to ground the agent into real-world tasks. We describe the details of each component in the following sections, and provide the example workflow in Appendix D.
Previous works demonstrate that generalist LLMs, such as T5 (Raffel et al., 2020), Flan-T5 (Chung et al., 2022), and InstructGPT (Ouyang et al., 2022), have a capability of manipulating the web environments (Shi et al., 2017) with great HTML comprehension (Furuta et al., 2023; Gur et al., 2022; Kim et al., 2023). However, they have not fully leveraged the HTML-specific inductive bias on syntax and semantics considered in the prior specialist transformer models (Li et al., 2021b; Wang et al., 2022a; Zhao et al., 2022). We here introduce HTML-T5, a pre-trained encoder-decoder language model, by interpolating the generalist and specialist nature of language models to solve downstream HTML-based web automation tasks efficiently. HTML-T5 processes HTML documents in a text-to-text manner, and leverages local and global attentions (Ainslie et al., 2020; Guo et al., 2022) in the encoder to handle the hierarchical structure of long HTML inputs. We pre-train it with large-scale HTML corpus curated from CommonCrawl on a mixture of long-span denoising objectives (Raffel et al., 2020; Tay et al., 2022), and finetune it for each downstream task. Especially, for WebAgent, we employ self-experience supervision to align the model with real websites.
Model Architecture In contrast to natural language texts, HTML documents have an explicit hierarchy from the tree structure; the relation of each element (e.g. ,