Can Foundation Models Wrangle Your Data?

Avanika Narayan, Ines Chami, Laurel Orr, Simran Arora, Christopher Ré

Introduction

Foundation Models (FMs) (Bommasani et al., 2021) are models trained on broad data that can be adapted to a wide range of downstream tasks. These models have achieved substantial gains across many semantically challenging tasks such as question answering (Brown et al., 2020), knowledge base construction (Petroni et al., 2019), and information retrieval (Guu et al., 2020). As they have scaled to hundreds of billions of parameters (e.g. GPT-3 (Brown et al., 2020), PaLM (Chowdhery et al., 2022)), large FMs have demonstrated surprising emergent behaviors and good zero-shot generalization to new tasks (i.e. no task-specific finetuning) on domains vastly different from the data they were pretrained on (Chowdhery et al., 2022). These large FMs are often autoregressive language models (e.g. GPT-3 and PaLM) that are trained to predict the next word in large text corpora and can be adapted to new tasks given a simple natural language description of the task (see Figure 1). These breakthrough capabilities have led to a race for building bigger and better models, and innovations continue to push the boundaries of what large FMs can do on a variety of hard language tasks.

A natural question that arises is whether these advances can benefit hard classical data tasks (e.g. data cleaning and integration). While it is clear that FMs benefit text-intensive tasks, it is not clear whether these models can be applied to data tasks over structured data. The symbols commonly found in structured data (e.g. dates, numbers, alphanumeric codes) are less frequent in natural language text so it is unclear that FMs possess the ability to reason over them. Moreover, since FMs are trained to predict the next word, it is non-obvious that they can work out-of-the-box on complex data tasks. This paper explores the aforementioned question and introduces a new research vision for leveraging FMs for data management, focusing on data cleaning and integration tasks—two keys steps in data-driven enterprise pipelines.

Recently, a large body of research has applied machine learning (ML) (Konda et al., 2016) and deep learning (DL) (Li et al., 2020; Mudgal et al., 2018) methods—namely pretrained language models (PLMs) like BERT (Devlin et al., 2018)—to semantically-complex data tasks. However, these approaches still require a significant amount of engineering effort as they rely on:

Task-specific architectures: Data cleaning and integration encapsulate many different tasks such as entity matching (Papadakis et al., 2020), schema matching (Sutanta et al., 2016), and error detection (Heidari et al., 2019). Existing approaches, whether they are rule-, ML- or DL-based, vary greatly from one task to the other, often with complex, task-specific architectures. For instance, adapting BERT to data tasks requires architectural changes and finetuning the entire model for each task. This leads to siloed and hard-to-maintain systems.

Hard-coded knowledge: Data tasks often rely on domain knowledge (e.g. understanding the relationship between a city and its zip code for data cleaning constraints) and commonsense reasoning. These are usually hard-coded with human-engineered rules or external knowledge bases (Chu et al., 2015; Rekatsinas et al., 2017). Consequently, systems can be brittle and fail to generalize to a diverse set of domains.

Labeled data: ML- and DL-based solutions require copious amounts of hand-labeled data (Adadi, 2021). For instance, PLMs that have achieved state-of-the-art (SoTA) results on data tasks (e.g. Ditto (Gururangan et al., 2020b)) require a significant amount of task-specific labeled data and fine-tuning to achieve good performance. Labeling data for each task is engineering intensive and adds to the difficulty of maintaining data cleaning and integration systems.

Excitingly, FMs display several useful properties that make them an appealing choice compared to traditional approaches:

Task-agnostic architecture: As a result of their natural language interface, FMs can be applied to a wide-range of tasks. For instance, Figure 1 shows how an entity matching task—which requires identifying whether two table entries refer to the same entity—can be cast as a prompting task. This unifying interface eliminates the need for siloed architectures, in contrast to existing learned approaches where architectures need to be carefully crafted for each task (e.g. task-specific classification layer).

Encoded knowledge: Because FMs are trained on large, generic corpora of data, they contain knowledge about an extensive set of common entities, and thus do not rely on human-engineered rules to acquire knowledge (Razniewski et al., 2021).

Limited to no labeled data: FMs can be applied to a breadth of tasks with little to no labeled data (e.g. few-shot and zero-shot). When a FM needs to be fine-tuned, it typically needs dramatically less labeled data to achieve competitive results (Kaplan et al., 2020).

Our goal is to better understand if large FMs can be applied to data integration and cleaning tasks. We study the behavior of GPT-3—an early and promising FM. While GPT-3 is already a high quality model, we expect the significant investment in FMs from both academia and industry to lead to more performant and scalable FMs over time. Like many other communities, the data management community stands to benefit from these trends. As such, we aim to understand the advantages and limitations of FMs on data tasks, by focusing on three key questions.

How well do large FMs transfer to data tasks? To answer this, we cast several data tasks as natural language generation tasks (Section 3) and explore whether a single FM can generalize well to these tasks. In Section 4.2, we quantify the zero- and few-shot performance of FMs on five enterprise data tasks: entity matching, error detection, schema matching, data transformation, and data imputation. We find that the largest GPT-3 variant (175B parameters) outperforms SoTA ML-and DL-based approaches on these tasks with few examples. This is particularly surprising since prior approaches are fully-finetuned on task-specific labeled data for these tasks, while GPT-3-175B is simply pretrained to generate text.

What are the caveats in applying FMs to data tasks? In Section 4.3, we unpack the few-shot “prompt tuning” process—serializing tabular data to text, casting data tasks as text generation tasks and constructing demonstrative task examples—for applying FMs to data tasks. We quantify the effects of prompt formatting variations on performance and the differences between manually and randomly selecting task examples. We find that FMs are brittle to differences in prompt formatting and that performance improves when prompts are manually selected versus randomly selected.

What opportunities do FMs present for data tasks and what are the relevant research challenges? Finally, in Section 5, we discuss the potential challenges and related research questions with using FMs in data management pipelines. We discuss the forthcoming shift in how ML systems are built, challenges around updating FM knowledge, and opportunities and considerations pertaining to private, temporal and local data.

We hope that our preliminary exploration will encourage the data management community to explore the effectiveness of FMs for other data tasks and develop techniques to overcome the shortcomings of FMs in this setting.

Background

We first give some background on the different data tasks considered in this paper and then provide a brief review of FMs.

We focus on entity matching (EM), error detection (ED), and data imputation (DI) and describe the setup for these tasks. We denote DD, a structured dataset with nn entries, such that each entry is represented by a collection of mm attribute value pairs: for entry ei∈De_{i}\in D we have ei={ei,j}1≤j≤me_{i}=\{e_{i,j}\}_{1\leq j\leq m} where for attribute jj, ei,j={attrj,valj}e_{i,j}=\{\texttt{attr}_{j},\texttt{val}_{j}\}.

Entity Matching The goal of EM is to match entities (real-world objects like people, places and things) across different datasets. Formally, given two structured datasets (D,D′)(D,D^{\prime}) and pairs of entries e,e′∈D×D′e,e^{\prime}\in D\times D^{\prime}, the goal is to predict whether these entries represent the same entity or not. This problem is usually solved as a classification problem, and real-world EM systems are often preceded by blocking heuristics which are used to remove obvious non-matches.

EM has been extensively studied over the past decade (see (Papadakis et al., 2020) for a survey) and methods broadly fall into three categories: rule-based, crowd-based (Gokhale et al., 2014; Wang et al., 2012) and ML/DL-based (Konda et al., 2016; Mudgal et al., 2018). Recently, methods relying on PLMs (Li et al., 2020) have become SoTA for this task.

Error Detection ED is an important step in data cleaning pipelines. Given an entry ee, the goal is to detect attributes jj where valj\texttt{val}_{j} has an error. The task is framed as a classification problem where the goal is to predict if valj\texttt{val}_{j} is correct for a given ee.

ED has been studied extensively in both academic and industrial settings (Abedjan et al., 2016). There are a number of successful commercial offerings including Trifacta (Tri, 2022) and Tamr (tam, 2022). Traditionally, ED systems have been heavily reliant on rule-based algorithms which enforce data constraints through functional dependencies or knowledge bases (Chu et al., 2015, 2013; Dallachiesa et al., 2013). Additionally, there are statistical-based approaches such as pattern enforcement (Kandel et al., 2011; Chu et al., 2015), outlier detection (Dasu and Loh, 2012), and record deduplication (Stonebraker et al., 2013) algorithms. Recent efforts have developed SoTA ML models for ED (Heidari et al., 2019).

Data Imputation DI is a critical step for repairing dirty data sources. Given an entry ee with missing attribute values {attrj,NULL}\{\texttt{attr}_{j},\texttt{NULL}\}, the goal of DI is to infer the missing values. The full range of plausible values for the missing value is not known apriori.

Prior works in DI falls into three categories—clustering and/or statistical-based (Mayfield et al., 2010), generative model-based (Ghysels et al., 2007), ML/DL-based (Biessmann et al., 2019b; Mei et al., 2021), and tabular data pretraining-based approaches (Deng et al., 2022)—and struggle when needing to impute values not seen in the training set (Mei et al., 2021).

2. Background on Foundation Models

We now give an overview of language FMs, starting from very early, smaller FMs (i.e. PLMs) and moving to large-scale FMs, the latter of which is the focus of this paper.

Pretrained Language Models Pretrained Language Models (PLMs) are neural networks pretrained on large corpora of publicly available text (e.g. web pages). The first breed of PLMs—ELMo (Peters et al., 2018), BERT (Devlin et al., 2019), RoBERTA (Liu et al., 2019)—learned the semantics of natural language by predicting masked works during pretraining. These models have on the order of hundreds of millions of parameters. Traditionally, PLMs are adapted to downstream tasks through a task-specific prediction layer (e.g. classification layer) and a task-specific finetuning step wherein all model weights are updated.

Large Autoregressive Language Models In 2020, GPT-3 (Brown et al., 2020) marked a significant shift in the the ML community. It represented a new class of large-scale language models: autoregressive language models pretrained to predict the next word in a sequence. These models have billions of parameters and have been used for language generation tasks such as question answering and summarization. Since the release of GPT-3, bigger and better performing autoregressive language models have been developed (Chowdhery et al., 2022).

Emergent Behaviors Interestingly, the biggest GPT-3 variant (175B parameters) has the capacity to solve natural language tasks with only a few examples (few-shot prompting), and, in some cases, just a task description (e.g. “Translate French to English”). Unlike traditional finetuning, no model parameters are updated to fit the task. Few-shot prompting has proven to be effective on tasks widely different from the FMs pretraining objective (e.g., code generation (Xu et al., 2022), Trivia QA (Lin et al., 2021; Brown et al., 2020) and arithmetic (Brown et al., 2020)). This in-context learning behavior can be thought of as the FM “locating” an already-learned behavior (Reynolds and McDonell, 2021). Smaller models (¡10B parameters) typically require some form of task-specific finetuning. We study ways to use smaller FMs on data tasks in the full report (Narayan et al., 2022).

Foundation Models for Data Tasks

Our goal is to understand whether FMs can benefit data cleaning and integration tasks. The procedure for applying FMs to data tasks requires adapting structured data inputs to textual inputs (Section 3.1), casting data tasks as text generation tasks (Section 3.2) and, for few-shot prompting, constructing demonstrative task examples to help the FM learn new data tasks (Section 3.3).

FMs take text as input and generate text as output. In order to apply these models to data tasks, we first need to convert structured tabular data inputs to text representations.

Given a structured dataset, we first convert an entry to text. Concretely, for entry e, we follow previous work (Li et al., 2020) and serialize the attribute value and entry as follows:

If the attribute value is NULL, it is serialized as the empty string. Based on the task and dataset, serialization only happens over a subset of attributes relevant to the task. FMs can be sensitive to the prompt format and the specific serialization used (Zhao et al., 2021). In Section 4.3 we study the sensitivity of FMs to sub-selection choices.

2. Data Tasks as Natural Language Tasks

Next, we need to convert data tasks to text generation tasks. We construct natural language descriptions of each task (i.e. prompts) that use the serialized representations from Section 3.1. The prompts are passed to the FM, whose generated output is the answer to the given task. For entity matching and error detection, the model generates a “Yes” or “No” response,Interestingly, the model is not constrained to produce a Yes/No answer on the output side, but we find that this happens most of the time. In the rare examples where the model does not predict a Yes/No answer, we use No as the default answer. and for imputation it generates the missing value. We now enumerate the prompts for each task.

Entity matching: given two entries (e,e′)(e,e^{\prime}) the template is

Data imputation: given an entry ee and attribute jj to infer, we use

Error detection: given an entry ee and attribute jj to classify as erroneous, we use

These templates highlight the generality of this framework which could be extended beyond the tasks considered in this paper.

3. Task Demonstrations

Task demonstrations can be included in the prompt to help the model learn a new task (see Figure 2 for an error detection example). These demonstrations are used to show the model how the task should be completed (e.g. should it generate Yes/No or a missing value) as well as understand the finer-grained semantics of DD. We explore two approaches for selecting task demonstration examples.

Random One approach is to sample random examples from a labeled dataset. However, this approach often causes high variance in FM performance (Liu et al., 2021a). Moreover, the ordering of examples in the prompt can have a non-trivial impact on the downstream performance (Lu et al., 2021; Zhao et al., 2021; Liu et al., 2021a). Although recent work has tried to systematize the prompt tuning process, it is still an open area of research (Li and Liang, 2021a; Sun and Lai, 2020).

Manual Another approach is to manually construct examples that yield good performance on a held-out validation set that is 10% of the original labeled dataset. This approach is more costly (requires more time) in comparison to random sampling but improves performance when examples are carefully constructed.

In our manual prompt tuning experiments, we spend at most 1 hour per task analyzing errors on the validation set. We then manually construct demonstrations that help the model correct these errors. We liken this step to the canonical hyperparameter tuning process in ML where time and compute are spent finding the best model parameters for the given task. However, the engineering effort for prompt tuning is significantly lower. Concretely, it takes less time (minutes vs. hours and days), is more compute effective (inference vs. full training), and gives the user finer-grained control (natural language guidance vs. blackbox parameter updates).

Experiments

We compare FMs to SoTA methods on a variety of data cleaning and integration tasks. Our goal is to understand whether large FMs can transfer to data tasks in zero- and few-shot settings (Section 4.2), and the nuances in applying FMs to data tasks (Section 4.3).

We begin by describing our experimental protocol, including models, datasets, metrics and baselines used.

Models For few-shot prompting we use the GPT-3-175B parameter model (text-davinci-002) in the OpenAI API endpoint (OpenAI, 2021).

Datasets For entity matching, we use the standard Magellan benchmark (Konda et al., 2016). For schema matching, we choose a challenging dataset, Synthea, from the OMAP benchmark (Zhang et al., 2021). For data transformation, we choose two challenging datasets from the TDE benchmark (He et al., 2018). For imputation, we choose two challenging datasets from (Mei et al., 2021): Restaurants and Buy. Finally, for error detection, we evaluate on the benchmark Hospital and Adult datasets which are used across several data cleaning papers (Heidari et al., 2019; Rekatsinas et al., 2017; Chu et al., 2013).

We use the provided train/test/dev splits for all entity matching datasets. For cleaning tasks, the splits are not available but we follow the protocol of (Heidari et al., 2019; Mei et al., 2021) to generate the dataset splits. For the Adult dataset, we evaluate over a randomly sampled set of 1K rows due to budget constraints.

Evaluation Metrics For the error detection, and schema/entity matching, we evaluate performance using F1 score. For imputation and data transformation, we evaluate using accuracy.

Baselines We compare against the SoTA methods for each task. For entity matching, we benchmark against Ditto (Li et al., 2020), the current SoTA DL-based approach which finetunes BERT (Devlin et al., 2019). For data imputation, we benchmark against IMP (Mei et al., 2021), which finetunes RoBERTa (Liu et al., 2019), and HoloClean (Rekatsinas et al., 2017), a statistical-based SoTA data repair engine. For schema matching, we compare against the SoTA model, SMAT (Zhang et al., 2021), which finetunes an attention-based BiLSTM. For data transformations, we compare against TDE (He et al., 2018), a SoTA search-based solution. Finally, for error detection, we compare against HoloClean and HoloDetect (Heidari et al., 2019), a data-augmentation based ML approach.

2. Zero/Few-shot Performance of Large FMs

In this section we explore the zero and few-shot performance of GPT-3-175B. Our goal is to understand whether large FMs transfer to data tasks in zero- and few-shot settings.

We first review the zero- and few-shot results.

Few-shot Performance Our results show that in the few-shot setting, with manually curated task demonstrations, GPT-3-175B achieves SoTA performance on 4 entity matching, 2 imputation, 2 data transformation, 2 error detection and 1 schema matching benchmark dataset(s) (see Table 1, Table 2, Table 3). The FM outperforms fully-finetuned, SoTA PLM-based approaches for entity matching (Li et al., 2020), and data imputation (Mei et al., 2021). For error detection, the few-shot approach outperforms the ML-based SoTA method, HoloDetect, which uses data augmentation and weak supervision to perform well.

Zero-shot Performance In the zero-shot setting, we observe that the FM outperforms statistical-based approaches and standard data repair engines (Rekatsinas et al., 2017) for imputation. On entity matching, the zero-shot performance is significantly lower than the few shot performance. This performance gap suggests that demonstrations are very important for the task, and we study the impact of task demonstrations in more detail in Section 4.3.

2.2. Discussion

These results show that large FMs can transfer to data tasks. These results are particularly exciting given that FMs are trained to model English language and have no prior exposure to the semantics of data tasks nor the syntax of tabular data. Furthermore, the zero-shot performance on imputation suggests that large FMs not only have an understanding of how to complete the tasks, but also have encoded knowledge that is needed to correct and complete records (e.g. functional dependencies between address and zip code). We analyze encoded large FM knowledge in the full report (Narayan et al., 2022).

On the entity matching datasets that the FM does not achieve SoTA on, we find that the FM struggles on data domains that contain jargon not commonly found in text. In such cases, the model lacks a strong semantic understanding of the input and has difficulty reasoning over the data. For example, in the Amazon-Google dataset, the model has difficulty matching samples due to the high volume of product-specific identifiers in the descriptions. Here, for instance, the model fails to accurately match the two entries: “name: pcanywhere 11.0 host only cd-rom xp 98 nt w2k me. manufacturer: symantec. price: NULL” and “name: symantec pcanywhere 11.0 windows. manufacturer: NULL. price: 19.99.”. We discuss ways to adapt FMs to domain-specific data in more detail in Section 5.

3. Prompt Tuning Ablations

In this section, we analyze the performance impact of the three different choices made during prompt tuning: attribute selection, prompt formatting, and task demonstration curation.

We run our ablations on three entity matching datasets (see Table 4). For all datasets, we evaluate on up to 200 samples for cost purposes. We discuss our findings next.

Attribute Selection First, we find through experimentation that sub-selecting attributes during row serialization can have a non-trivial impact on entity matching performance. In particular, we observe that sub-selecting attributes that are essential in determining whether two entities match (e.g. name) and removing noisy attributes improves model accuracy. To better illustrate this point, we evaluate FM performance on three datasets when not sub-selecting attributes. We find that, on average, attribute selection results in a 13.7 F1 point performance improvement (Table 4).

Prompt Formatting Second, we observe that FMs can be brittle to subtle variations in prompt templates. We investigate this brittleness by replacing the span “Are Product A and Product B the same?” (Prompt 1) with “Are Product A and Product B equivalent?” (Prompt 2). This minor modification results in an average 9.4 F1 point performance gap in the datasets in Table 4.

Task Demonstrations Finally, we find that the choice of task demonstrations has a significant impact on downstream performance. We conduct an ablation where we replace manually curated task demonstrations with randomly selected demonstrations (see Prompt 1 (w/o Example Select.) in Table 4). We run this experiment over three different random seeds and report the results in Table 4. Across all cases, manually curated examples outperform randomly selected examples by an average of 14.7 F1 points.

3.2. Discussion

The aforementioned results demonstrate that successful prompt tuning requires (1) selecting an informative set of attributes required for the task, (2) crafting a well-formatted prompt that the FM understands, and (3) constructing a set of instructive task demonstrations that condition the model to the data at hand. For (1), we find that attribute sub-selection boosts performance by removing noisy attributes that hurt performance. For (2), our ablations show that prompt formatting (e.g. word choice, punctuation) can have significant impact on model performance. For (3), our results indicate that examples need to be carefully crafted for FMs to learn new tasks. We conjecture that on more reasoning-intensive tasks (e.g. matching), prompts are important as they help teach the model how to reason about entities and how to complete the task. Moreover, we emphasize that all three steps require some form of iteration to develop the most effective prompt (e.g. passing various inputs to the model and inspecting its outputs).

These findings are aligned with existing literature on prompt tuning that observe non-trivial variance in prompt-based learning settings (Zhao et al., 2021). This performance variance suggests that iterative prompt programming is an essential human-in-the-loop process for FM usage (Liu et al., 2021b). However, some works suggest that smarter, automatic example selection methods can help close the gap between random example selection and human-in-the-loop prompt selection (Liu et al., 2021a). These results highlight the paradigm shift induced by building systems centered around FMs: instead of spending time tuning models, we now need to invest time finding the right examples and engineering useful prompts for each task. We discuss these paradigm shifts in more detail in Section 5.

Research Agenda

Because of their natural language interface and vast internal knowledge, FMs provide an interface for unifying a wide-range of siloed, hand-engineered data integration pipelines. Consequently, we envision that the data orchestration workbenches of the future will be centered around FMs. We discuss the opportunities, practical considerations, and technical challenges associated of this vision next.

Natural Language Interactions FMs usher in a new era of human-machine collaborative workflows wherein users spend less time labeling data and finetuning models and more time writing natural language prompts that are representative of the task at hand. As the necessity of writing code decreases, we envision systems that are more accessible to non-machine learning experts (e.g., business users). In future work, we seek to better understand the human-in-the-loop prompt engineering pipeline, especially in the context of data management practices.

Model Prototyping The data integration and management pipeline can be categorized into three distinct stages: discovery and design, development, and deployment (IBM, [n. d.]). We propose that FMs will be most useful in the discovery and design phase when training data is less available. In this setting, FMs enable rapid prototyping of data models via prompting. In some cases, the FM’s out-of-the-box performance will be sufficiently high, obviating the need to train a task-specific model. In others, we can use the FM to label and generate data with human-in-the-loop feedback. When a sufficient amount of data has been collected, transitioning to the fully-supervised model development regime is the optimal choice.

Passive Learning From User Exhaust In the enterprise setting, organizations accumulate a staggering amount of data exhaust—the informational byproduct that streams from devices, products, and workforce management practices (dat, 2018). Because FMs are pretrained in an unsupervised fashion with a simple token prediction objective (see Section 2.2), they can learn over any raw and unlabeled sources of data (Dhariwal et al., 2020; Yan et al., 2021; Reed et al., 2022). As such, FMs can effectively ingest the exhaust from the entire data stack (from system logs to structured data). Learning from data analyst exhaust (e.g., clicks over GUI) is also an opportunity to improve FM performance for these data tasks.

2. Practical Considerations of FMs

Integration in Data Management Workflows FMs take text as input and generate text as output. As a result, they are limited in their ability to directly take actions over the graphical user interfaces (GUIs) of data management software (DMS) (e.g., Snowflake). Given that a majority of data analyst time is spent working in these environments, we need ways of translating natural language specifications of data tasks (e.g., “remove all nan cells”), to actionable operations on these GUIs. Excitingly, new work demonstrates how to augment FM capabilities such that they can effectively take actions on web and mobile interfaces (Wang et al., 2022; Adept, [n. d.]). These works suggest the possibility of directly integrating FMs with DMS.

Interaction with Existing Systems FMs can interact with existing DMS by either replacing them or utilizing their outputs to conduct downstream tasks. In terms of system replacement, for tasks where the required rules and logic are not encoded in the FMs knowledge, it is an open question (Desmond et al., 2022) as to how to systematically translate the rules and logic (e.g., domain-specific dependencies) to natural language inputs to the FMs. For system integration, we need ways of systematically incorporating the outputs of existing systems (e.g., dataset pattern discovery) in natural language prompts. Prior work has demonstrated how to “ground” FM-based text-to-SQL prompts with database schemas and metadata (Deng et al., [n. d.]). Similar ideas could be used when incorporating system outputs with FMs.

Debuggability FMs are non-deterministic and can make unexpected errors. In order to use these models in data management pipelines, we need mechanisms for increasing transparency and debuggability of pipelines. One possible approach is to collect and monitor model confidence scores. Prior work demonstrates that a FM can “learn to express uncertainty about its own answers in natural language” (Lin et al., 2022). Another approach is to decompose tasks into chains of “primitive operations”, enabling for more transparency and visibility of specific failure points (Wu et al., 2022).

3. Technical Challenges of FMs

Domain Specificity A key challenge in applying FMs to data management tasks is operating over highly specialized domains (e.g., medical, financial, and insurance data). The existing literature suggests that domain-specialization is best achieved by continuously pretraining on relevant data (e.g., earning reports, medical reports) (Yang et al., 2020; Gururangan et al., 2020a). However, training models with billions of parameters can be costly. This has motivated several works which attempt to more quickly and cheaply adapt models to new domains by finetuning only a few layers of the model or simply training a small neural network (e.g., adaptor) on top of a frozen FM (Labs, 2022).

Privacy Organizations are often unable to pass sensitive data to third party APIs for privacy reasons (Arora and Ré, 2022; Arora et al., 2022a). Unfortunately, this constraint is in conflict with the contemporary state of FMs where the best performing models (e.g., GPT-3), can only be accessed via API requests. This motivates the need for better open-source models which are competitive with closed-source models. Excitingly, recent work (Arora et al., 2022b) has demonstrated that techniques such as prompt ensembling and prompt reframing can enable open-source models like GPT-J-6B (Wang and Komatsuzaki, 2021) to out-perform GPT3-175B on popular natural language understanding benchmarks. Extending this work to data management tasks is an interesting research direction.

Prompt Engineering and Automation Prompting is the primary mechanism for adapting FMs to new tasks. However, prompting requires some manual effort to design performant prompts. In the DI setting, prompts can be sensitive to things such as data schemas and minor formatting variations. A few automated approaches have been developed which reduce the manual effort needed to construct prompts: soft prompt tuning (e.g., optimizing a sequence of task-specific vectors which are appended to the text prompt) (Lester et al., 2021a; Li and Liang, 2021b) and learning to retrieve better in-context examples (Rubin et al., 2021). Adapting these works to the challenges of prompting for data tasks is an open area of research.

Conclusion

In this work we investigate the applicability of FMs to classical data tasks. We find that large FMs can achieve SoTA performance on many data tasks with 0-to-few natural language task demonstrations. The ability of these models to transfer to data tasks with no task-specific finetuning is particularly interesting given that these models are simply trained to predict next words. Our work builds upon years of important work on integration and cleaning tasks in the data management community. We hope that our results gesture towards the possibilities of using language guided models for human-in-the-loop data integration practices across a broader range of data management tasks.

Acknowledgments

We are thankful to Ihab Ilyas, Theo Rekatsinas, Mike Cafarella, Ce Zhang, Sen Wu, Christopher Aberger, Neel Guha, Beidi Chen and Xiao Ling for their helpful discussions and feedback. We gratefully acknowledge the support of DARPA under Nos. FA86501827865 (SDH) and FA86501827882 (ASED); NIH under No. U54EB020405 (Mobilize), NSF under Nos. CCF1763315 (Beyond Sparsity), CCF1563078 (Volume to Velocity), and 1937301 (RTML); ONR under No. N000141712266 (Unifying Weak Supervision); the Moore Foundation, NXP, Xilinx, LETI-CEA, Intel, IBM, Microsoft, NEC, Toshiba, TSMC, ARM, Hitachi, BASF, Accenture, Ericsson, Qualcomm, Analog Devices, the Okawa Foundation, American Family Insurance, Google Cloud, Swiss Re, Brown Institute for Media Innovation, Department of Defense (DoD) through the National Defense Science and Engineering Graduate Fellowship (NDSEG) Program, Fannie and John Hertz Foundation, National Science Foundation Graduate Research Fellowship Program, Texas Instruments Stanford Graduate Fellowship in Science and Engineering, and members of the Stanford DAWN project: Teradata, Facebook, Google, Ant Financial, NEC, VMWare, and Infosys. The U.S. Government is authorized to reproduce and distribute reprints for Governmental purposes notwithstanding any copyright notation thereon. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the authors and do not necessarily reflect the views, policies, or endorsements, either expressed or implied, of DARPA, NIH, ONR, or the U.S. Government.

References

Appendix A Small FM Finetuning Experiments

In this section, we explore ways to bridge the performance gap between the smaller and larger FMs, focusing on (1) full finetuning (updating all parameters) and (2) lightweight finetuning (updating a small number of parameters). In Section A.1, we evaluate the two approaches and unpack their sample (as measured by number of labeled samples) and training (as measured by number of parameter updates) efficiency tradeoffs. We find that both lightweight and full finetuning can be used to reduce the performance gap between a 6.7B and a 175B parameter model to an average of 3.1 points. The lightweight finetuning requires more data than the fully finetuned approach, but updates only 5% of the model parameters.

Full/Lightweight Finetuning The typical practice for achieving optimal task performance in LLMs is to finetune pretrained models with task-specific data. This approach is usually training inefficient as all weights in the model are updated. Interestingly, FMs have been shown to achieve optimal task performance by simply training a lightweight, non-linear layer (i.e. adaptor) (Figure 3) on the outputs of a frozen model (Labs, 2022; Lester et al., 2021a). This is a training efficient method that has proven to be competitive with fully-finetuned approaches (Labs, 2022). However, because this approach trains a layer from scratch, it is usually less sample efficient than the fully finetuned approach. Figure 4 visualizes the differences in sample and compute efficiency between the two approaches.

For lightweight finetuning, we recursively join a single frozen FM with a small trainable network (the adapter) (Labs, 2022) (Figure 3). The prompt input is first passed to the frozen FM whose output embeddings are transformed by the adapter and then passed again as input embeddings to the frozen FM (Labs, 2022). The generated output of the second pass yields the final answer of the model. In this setup, updates are only made to the adapter, i.e. no updates are made to the FM weights alleviating the need for costly training.

For both the full and lightweight finetuning, the model is fed the natural language prompt in Section 3.2 as input, and trained to generate the task output (e.g., a Yes/No string, or missing value).

In this section we explore the fully finetuned and lightweight finetuned performance of two small FMs (namely GPT-3-6.7B and GPT-3-1.3B). Our goal is to understand whether finetuning can bridge the performance gap with larger models, and analyze the tradeoffs between sample efficiency and training efficiency.

A.2. Experimental Setup

For adapters, we use the GPT-Neo 1.3B (Black et al., 2021) and the 6.7B parameter Neo GPT-J (Wang and Komatsuzaki, 2021) models available on HuggingFace. All finetuning experiments are run on 4-8 A100 GPU machines using Deepspeed (Rasley et al., 2020) with ZeRO-2 optimizations. We use the AdamW optimizer (Loshchilov and Hutter, 2017) with a 10%10\% learning rate warmup followed by a linear decay. We train for a maximum or 30 epochs with a learning rate of 1e-4 for adapters and 2e-5 for full finetuning, and save the best model based on validation metrics.

A.3. Experimental Results

We first review the finetuning results on Walmart-Amazon (EM), Hospital (ED) and Restaurant (DI).

Full Finetuning Our results show that in the full finetuning setting, we can effectively reduce the performance gap between GPT-3-6.7B and GPT-3-175B on all datasets (Figure 5), using as little as 10% of the training set for Walmart-Amazon. GPT-3-1.3B also matches GPT-3-175B performance on Walmart-Amazon and Restaurant, and is within 8 points of the GPT-3-175B on Hospital. Compared to the 6.7B model, the 1.3B model is less sample-efficient and needs more example to bridge the performance gap.

Lightweight Finetuning In the lightweight adapter setting, we find that we can effectively bridge the performance gap between GPT-3-6.7B and GPT-3-175B on both Restaurant and Walmart-Amazon, but not Hospital. For GPT-3-1.3B, we are less effective in bridging the gap, and see an average 25 point performance difference between the GPT-3-175B across the three datasets. Furthermore, we find that the adapter model outperforms the fully-finetuned approach on two datasets with up to 4 points.

These results show the feasibility of reducing the performance gap between GPT-3-175B and the smaller models. We observe some clear sample-training efficiency tradeoffs between the fully and lightly finetuned approaches, which we discuss next.

Sample Efficiency We find that fully finetuned models are usually more sample efficient than adapters, which could be explained by the need to train a layer from scratch in adapters. We also observe that sample efficiency improves with size as the performance of the GPT-3-6.7B trained with 10% of the data is equal to, or better than, the performance of the GPT-3-1.3B trained with 50% of the data in both the full and lightweight finetuning settings.

Finally, we comment on the performance difference between 6.7B adapter model and the 6.7B finetuned model on Hospital. We postulate that adapter model performed poorly on this dataset because the training set was particularly small (100 samples) which suggests that the adapter set-up requires more samples to learn generalizable patterns.

Training Efficiency We find that the adapter approach for GPT-3-6.7B outperforms the full finetuning approach on 2 of 3 datasets with 5% of the amount of trainable parameters, demonstrating that training efficiency need not be sacrificed for quality at this scale. However, at the smaller scale (GPT-3-1.3B), there is a non-trivial tradeoff between performance and training efficiency suggesting that training efficiency decreases as model size decreases.

Appendix B Knowledge Ablations

In this section we seek to explore the encoded FM knowledge that is useful for data tasks. We begin with a qualitative analysis (Section B.1) and follow with a slice-based analysis that unpacks the impact of different training mechanisms on FM knowledge (Section B.2).

To better illustrate the notion of encoded knowledge that large FMs possess, we conduct a qualitative analysis where we inspect the model predictions for the task of inferring missing zipcodes or cities given some context (Table 6) . GPT-3-175B is able to effectively apply its understanding of functional dependencies between address and zip code and dependencies between address, phone number and city to correctly impute the desired value. We further notice that while the smaller models fail to impute the correct values, their generated outputs have the correct semantic type of the missing attribute without out any demonstrative examples.

B.2. Slice analysis

We tease apart the differences in encoded knowledge across the different training regimes through a slice-based performance analysis on the Restaurant dataset. We use the frequency counts of city names in the training set to define three frequency-based subclasses (Table 5). We find that for entities that do not occur in the training set, GPT-3-175B is the only model that is able to correctly impute the missing values. Moreover, we observe that for rare entities—a city value of West LA—these pattern can only be learned through full or lightweight finetuning. Interestingly, we find that the lightweight approach more accurately learns the rare subclasses relative to the fully-finetuned approach in both the 100% and 50% training set regime. We hypothesize that this is because the adapter model is smaller, and is thus less prone to overfitting the training set.

Appendix C Effects of FM Bias

Because FMs are pretrained on large corpuses of web text, they inherit the biases of the data they are trained on. As a result, FMs have been found to contain a number of social biases (e.g. gender, religion, race and more) (Abid et al., 2021; Lucy and Bamman, 2021). It is important to recognize the effects of these biases when utilizing FMs in downstream applications. When applying FMs to data management tasks, we need to be wary of the fact that the inherent biases of these models may be propagated through their actions. Concretely, for data repair, the encoded biases of the FM may cause it to replace or correct entries in a biased manner (e.g. adding a suffix of Ms. to a traditionally female name). As a result, we need mechanisms for mitigating model bias and monitoring for bias in model outputs in deployment settings. Recent work takes a first step in this direction by proposing methods for automatically detecting bias-sensitive tokens and correcting these tokens to mitigate biases (Liang et al., 2021).

Appendix D Evaluating More FMs on Data Wrangling Tasks

This works focuses on evaluating GPT-3—one of the many large FMs—on a collection of data integration and cleaning tasks. We contribute tasks from this paper to the HELM benchmark (Liang et al., 2022), which evaluates performance of a broader set of large FMs on these tasks.