WebLINX: Real-World Website Navigation with Multi-Turn Dialogue

Xing Han Lù, Zdeněk Kasner, Siva Reddy

Introduction

Proprietary conversational assistants like ChatGPT (OpenAI_2022_ChatGPT) are capable of more than just conversing; they can also browse websites through plugins (chatgpt-plugins; bartextensions), allowing them to perform actions and provide more useful responses. However, this capability is limited: the plugins must be developed separately for each website and may not cover all of a website’s functionality. This limitation raises an important research question: can we leverage the models behind those assistants to navigate websites directly in the user’s browser, while retaining their conversational capabilities?

Motivated by this question, we define the problem of conversational web navigation: given the initial user instruction, an agent must complete a real-world task inside a web browser while communicating with the user via multi-turn dialogue. This problem is relevant in many real-world scenarios: helping visually impaired users efficiently navigate websites through a chat interface, enhancing smart speakers and digital assistants with voice-controlled web navigation, and improving the productivity of knowledge workers by reducing highly repetitive steps while staying in control. From a research perspective, this problem can be used to assess the ability of LLM agents to not only follow self-contained instructions, but also engage with their environment through dialogue and generalize to unforeseen situations.

To address this problem, we introduce WebLINXWeb Language Interface for Navigation through eXemplars (§3), a benchmark containing 2337 demonstrations of conversational web navigation produced by human experts across 155 real-world websites. Figure 1 shows a demonstration. Each demonstration captures the full sequence of actions performed by a human navigator when interacting with the user (known as instructor) through a conversational interface. We record over 100K occurrences of actions and utterances, where each action is associated with a Document Object Model (DOM)Tree representation of HTML page as rendered in the browser. tree, browser screenshots, and frames from demonstration-level video recordings. WebLINX: Real-World Website Navigation with Multi-Turn Dialogue highlights the unique aspects of WebLINX. Unlike previous works focused on mobile apps or specialized applications, ours is the first large-scale benchmark that can be used to train dialogue-enabled navigation agents and evaluate their generalization capabilities to realistic scenarios, such as adapting to new websites, categories, and geographies; we also reserve a split to assess the ability of agents to interact with instructors without visual access to the browser.

A naive way to use this benchmark would be to give the full DOM tree directly to an agent and instruct it to predict the correct action. As some HTML pages contain thousands of elements, fitting them completely within the context of a LLM poses a significant challenge; even if it was possible, existing LLMs would be unable to process them in real-time. Consequently, we design a method called Dense Markup Ranking (§5.1), which compares each element in an HTML page with the full action history. By using a similarity-based approach to both learn and rank elements, we can leverage compact architectures used in text retrieval. This lets us find the most relevant elements and prune irrelevant ones to obtain a compact representation of the DOM. We combine it with the action history, detailed instruction and screenshot (in a multimodal context) to construct an input representation for LLMs, which can now meaningfully predict which actions to take. However, even if a predicted action is correct, it may be identified as incorrect by existing metrics, which can happen when there are minor differences in an agent’s response or when an overlapping element is selected. Thus, we design a suite of evaluation metrics (§4) tailored for specific types of action (for instance, clicking should be evaluated differently from what the navigator says).

We examine 19 models based on 8 architectures (§LABEL:sec:results), including smaller image-to-text, larger text-only decoders, LLMs, and multimodal models (capable of accessing both image and text). Among them, 5 are in the zero-shot setting, and the remaining are finetuned using the training split of WebLINX. We find that even the best zero-shot model, GPT-4V (GPT4V_System_card), is surpassed by finetuned models (§LABEL:sec:overview_results). Notably, a smaller model like Sheared-LLaMA (xia2023sheared) outperforms the much larger Fuyu (fuyu_blog_post), which was pretrained with browser screenshots. However, all models face challenges in generalizing to new settings, such as unseen websites from a different geographic location or when the instructor gives instructions without seeing the screen. Those findings prompted us to qualitatively look at the behavior of the models (§LABEL:sec:qualitative_assessment), where we find that GPT-4V lacks situational awareness and can make obvious blunders. However, the best finetuned models still fail in simple cases, such as clicking on non-existing links or failing to change the language of a translation app. Thus, we believe that significant effort will be needed to make progress on the problem of conversational web navigation, as we discuss in LABEL:sec:discussion.

Our contributions are summarized as follows:

We introduce the task of conversational web navigation and a large-scale expert-annotated benchmark for it, named WebLINX (§3).

We propose a suite of action-specific metrics, which we combine to assess overall model performance (§4).

We design a method to simplify HTML pages (§5.1), allowing us to evaluate a wide range of models (§5.2).

We find that smaller text-only decoders outperform multimodal LLMs, but all finetuned models struggle to generalize to novel scenarios (§LABEL:sec:results).

Previous work predominantly focused on building web agents for a single task. A prominent work for task-driven web navigation is MiniWoB++ (Shi_Karpathy_Fan_Hernandez_Liang_2017; Liu_Guu_Pasupat_Shi_Liang_2018), a simulated web environment with an extensive list of task primitives (e.g., select value from a dropdown or date from a calendar). Its well-defined input space and the flexibility of its simulated environments lead to reinforcement learning approaches reaching human-level performance (Liu_Guu_Pasupat_Shi_Liang_2018; Humphreys_Raposo_Pohlen_Thornton_Chhaparia_Muldal_Abramson_Georgiev_Goldin_Santoro_2022). However, the ability of those methods to transfer to realistic settings have been limited, even after introducing environment extensions (Gur_Jaques_Miao_Choi_Tiwari_Lee_Faust_2022) and sample-efficient methods (kim2023language). Other works also explored grounding language commands to web elements and mobile UIs (Pasupat_Jiang_Liu_Guu_Liang_2018; li2020mapping; Burns_Arsan_Agrawal_Kumar_Saenko_Plummer_2022), or question answering (QA) by navigating Wikipedia (Nogueira_Cho_2016).

In an effort to build more realistic environments, Yao_Chen_Yang_Narasimhan_2023 introduced WebShop, an e-commerce environment with over 12K human-written task instructions. Models trained on WebShop achieved strong performance, but still relied on clean HTML and simple visual representations (Furuta_Nachum_Lee_Matsuo_Gu_Gur_2023). Instead, we aim to build agents that can act on any real-world website, often existing in noisy and dynamic environments.

The prospect of using LLMs to act on real websites (Nakano_Hilton_Balaji_Wu_Ouyang_Kim_Hesse_Jain_Kosaraju_Saunders_2022) has lead to the development of LLM-based navigation services (adept-act1; multi-on; hyperwrite-ai), which has set the stage for academic counterparts. Mind2Web (deng2023mind2web) and WebArena (zhou2023webarena) are large-scale resources for building autonomous navigation agents like SeeAct (zheng2024gpt4vision) and WebVoyager (he2024webvoyager). On the other hand, WebLINX is a benchmark for building agents that can interact with users in a multi-turn dialogue fashion, allowing them to be steered towards precise goals.

In this section, we introduce WebLINX, a large-scale benchmark for conversational web navigation consisting of 2337 demonstrations with an average of 43 turns. It contains interactions between a human user (referred to as instructor) and human assistant (navigator) aiming to complete tasks across 155 real-world websites selected from 15 geographic areas. We classify the websites into 8 categories and 50 subcategories based on their domains.

The data statistics are summarized in WebLINX: Real-World Website Navigation with Multi-Turn Dialogue and a breakdown by category and split is illustrated by Figure 2. Additional statistics about the dataset, including the number of demonstrations in split, can be found in LABEL:appendix:supplementary_statistics, along with the list of categories in LABEL:appendix:categories_and_subcategories.

Demonstration Framework

The demonstrations capture real-time interactions, which are recorded by the navigator controlling the web browser. Each demonstration D={s1,a1,…,sn,an}\mathcal{D}=\{s_{1},a_{1},\ldots,s_{n},a_{n}\} is a sequence of nn states s∈Ss\in\mathcal{S} and actions a∈Aa\in\mathcal{A} . At each turn t∈{1,…,n}t\in\{1,\ldots,n\}, the state sts_{t} contains the representation of the website. Each action follows one of the 5 core intents described in Table 3. The full list of intents is provided in LABEL:appendix:intents_descriptions.

Data Collection

To collect the demonstrations, we worked with a professional data labeling company,EsyCommerce: esycommerce.com who enlisted 8 expert annotators that received detailed instructions and extensive training to complete our tasks. The annotators worked in pairs: an instructor interacts with a navigator who completes the tasks in a web browser (see Figure 3). Both use the chat interface to communicate, but only the navigator controls the browser. We designed an app, browser extension, and processing pipeline to record the demonstrations, which are subsequently validated by a different annotator under the supervision of the original navigator (details in LABEL:appendix:data_collection_details).

Evaluation Splits

In addition to a Train split, we create Valid and \textscTest\textsciid\textsc{Test}_{\textsc{iid}} to assess in-domain generalization, and 4 out-of-domain splits for various scenarios (see Section 2.1).

1 Representing actions and states for modeling

At each turn tt, we have access to the state sts_{t} to predict an action ata_{t}. The state consists of the following (if available):

ctc_{t}: Candidate elements that can be targeted by ata_{t},

iti_{t}: Screenshot of the navigator’s browser,

vtv_{t}: Viewport size (height and width),

Note that a state need not contain all of the above. For example, at the start of a demonstration, the instructor and navigator may need multiple rounds of dialogue to properly define the objective, in which case the initial states do not have DOM trees or screenshots. A model mm predicts an action ata_{t} for a given state sts_{t} based on a prompt template pmp_{m} which indicates how to make use of the contents in a state.

Since a model mm has a limited input length in practice, we represent history hh as the set of past five actions (denoted as ara_{r}) and five utterances (uru_{r}). We could not include the representation of past states such as elements or screenshots.

Parsing Action Output

An action consists of an intent and argument and can be generated by an agent in a textual format. It must follow a pre-defined structure (see Table 3) that allows it to be parsed into a structured form, which can be executed in a browser using tools like Selenium.https://www.selenium.dev/ We discuss additional details in LABEL:appendix:output_processing.

In this section, we describe the evaluation metrics (§4.1) and their applicability to specific groups of intents (§4.2).

A commonly used metric in prior work on web navigation is task success rate, which measures the proportion of demonstrations where the model reached the desired final state (Shi_Karpathy_Fan_Hernandez_Liang_2017; Yao_Chen_Yang_Narasimhan_2023; deng2023mind2web). However, this metric is inappropriate for our benchmark because the objective is not fully defined in the first turn or later turns; instead, it evolves as the conversation proceeds. We instead leverage turn-level automatic evaluation metrics, following established approaches in dialogue systems (rastogi2020towards; zhang2020dialogpt). The aim of the metrics is to provide a heuristic estimate of the similarity between the predicted action and the reference action.

Given prediction a′a^{\prime} and reference aa, the intent match is \textscIM(a′,a)=1\textsc{IM}(a^{\prime},a)=1 if the intents are equal, otherwise \textscIM(a′,a)=0\textsc{IM}(a^{\prime},a)=0. This tells us if a model can correctly identify which action to perform, but does not indicate if the model can predict the correct arguments.

Element Similarity using IoU

For actions with elements as arguments (click, textinput, submit), we compute the intersection over union (IoU; jaccard1912distribution). Given the area of a bounding box B\mathcal{B}, we have:

To compute the area, we use (x,y) coordinates of the reference and predicted bounding boxes. This formulation (1) favors elements with high visual overlap, (2) penalizes predicting elements much smaller or larger than reference elements even if one is completely contained by the other, and (3) assigns 0 if the elements do not overlap.

Text Similarity using F1

To measure lexical similarity of text arguments in say and textinput, we calculate chrF (popovic-2015-chrf), an F1-score for character n-gram matches (we use the default setting of n=6n=6). Similar to IoU, we scale by the IM, resulting in \textscIM(a′,a)×\textscchrF(a′,a)\textsc{IM}(a^{\prime},a)\times\textsc{chrF}(a^{\prime},a). In the case of load intent, URLs follow a structure that can be consistently segmented, which leads us to apply the F1-score on segments instead of n-grams; we call this measure URLF. We use F1 to refer to either chrF and URLF, depending on whether an action contains a text or URL argument.

2 Turn-level score and overall score

To allow better comparisons between models, we divide the intents into groups: The element group (EG) contains click, textinput, and submit, and is evaluated with IoU. The text group (TG) encompasses load, say, and textinput, and is evaluated with F1.

We assign a turn level score based on the following: If the turn involves an action in EG, the score is the same as IoU, i.e. score is 0 when the intent is incorrect or the element doesn’t overlap, it is 1 when intent is correct and the element perfectly overlaps, and it is somewhere in between for the rest. For TG actions load and say, the score is same as F1, i.e., score is 0 when either intent is incorrect or there is no text overlap, it is 1 when intent is correct and the text matches exactly, and it is somewhere in between for the rest. For textinput, the turn score is IoU×F1\text{IoU}\times\text{F1} since it contains both text and element arguments. Finally, we compute the overall score using the micro-average of turn-level scores.

In this section, we describe a method for selecting candidate elements (§5.1) and how to use them in textual input. We use these methods to build models that can accurately predict actions (§5.2). We report results in LABEL:sec:results and provide implementation details in LABEL:appendix:additional_modeling_details.

To choose a set of suitable candidates for the model input (§3.1), we need a candidate selection stage that filters the full set of elements in the DOM tree. deng2023mind2web proposed to pair each DOM element with the task query and input them into a DeBERTa model (he2021deberta), which is finetuned using a cross-encoder loss (reimers-gurevych-2019-sentence). We found this method takes on average 916ms to select candidates for a given turn.Calculated on the training set, see LABEL:appendix:paragraph_empirical_speed_improvements. When factoring in network latency and LLM inference, this would result in poor processing time. It is thus crucial that we use efficient ranking method to build agents that can operate in real time and learn from interactions with users.

To solve this, we propose Dense Markup Ranking (DMR), which is 5 times faster than the previous approach, at the cost of slightly lower recall. The method consists of: (1) a simplified element representation to reduce computational overhead; (2) a dual encoder-based approach (reimers-gurevych-2019-sentence; karpukhin-etal-2020-dpr); (3) similarity-based learning between the text representation of sts_{t} and a1:t−1a_{1:t-1} and corresponding HTML elements. Using this method, we finetune a variant of MiniLM (Wang2020MiniLMDS). We formulate the cosine-based learning objective, examine the inference speed improvements, and evaluate alternatives in LABEL:appendix:dmr_details.

Even after our candidate selection, the input sequence length to a model can exceed its limit, so we truncate the sequence. To reduce information loss from traditional truncation (e.g., for large DOM elements and long history), we design a strategy that leverages the hierarchical nature of the input to determine which subsection should be truncated. We introduce several improvements to the representation used in prior works by including the full HTML attributes, viewport size, XML Path, and the bounding boxes of candidate elements (implementation details in LABEL:appendix:otr_details and LABEL:appendix:truncation).

2 Modeling Actions

Upon selecting the most promising candidates for a given state sts_{t}, we can combine them with the remaining information in sts_{t} to construct a representation that can be used to predict action strings, which can be parsed and executed (§3.1). To understand which factors matter for predicting actions, we examine 19 zero-shot and finetuned models (using the Train split) with different input modalities: image-only, text-only, and both. We provide implementation details in LABEL:appendix:implementation_experiments and hyperparameters in LABEL:appendix:hyperparameters_details.

We categorize action models by the input modality, since the output is always in a structured format (§3.1). We define the following types: (1) text-only, which receives instructions, pruned DOM tree, candidate element description and history; (2) image-to-text, which receives the screenshot, instructions and past actions directly embedded in the image; (3) multimodal, which receives the screenshot, instructions, pruned DOM tree, candidate description and history directly as text. Additional discussions are found in LABEL:appendix:details_on_model_categorization.

Text-only models

The recent MindAct (deng2023mind2web) model is a Flan-T5 (chung2022scaling__flant5) model that has been finetuned on Mind2Web. We further fine-tune it on WebLINX using its original configuration.

To quantify the improvements brought by DMR-based representation (§5.1), we directly finetune Flan-T5 checkpoints, allowing us to control for size and architecture with respect to MindAct. We also finetune LLaMA-2 (touvron2023llama; touvron2023llama2)We use the variants finetuned on chat. and a distilled version, Sheared-LLaMA (S-LLaMA; xia2023sheared).

Proprietary text-only LLMs

We report results for GPT-3.5 Turbo (brown2020language; Andrew_Peng_Michael_Wu_Logan_Kilpatrick_Steven_Heidel_2023), in both zero-shot (3.5T) and finetuned (3.5F) settings. We also include zero-shot results for GPT-4T (OpenAI2023GPT4TR).

Image-to-text modeling

We explore Pix2Act (shaw2023pixels) an encoder-decoder (Vaswani2017AttentionIA) purely finetuned on pixels. It uses a Pix2Struct backbone (Lee_2022_Pix2Struct), which is pretrained on screenshots using a Vision Transformer encoder (dosovitskiy2020image) and a text decoder. We follow the behavior cloning approach used by Pix2Act by finetuning the same backbone on WebLINX.

Multimodal models

We finetune Fuyu-8B (fuyu_blog_post), a base model pretrained on browser screenshots by modeling images and text using a unified architecture. We also report zero-shot results for the variant of GPT-4 with vision capabilities (GPT-4V; GPT4V_System_card).