The Art and Practice of Data Science Pipelines: A Comprehensive Study of Data Science Pipelines In Theory, In-The-Small, and In-The-Large
Sumon Biswas, Mohammad Wardat, Hridesh Rajan
Introduction
Data science processes, also called data science stages as in stages of a pipeline, for descriptive, predictive, and prescriptive analytics are becoming integral components of many software systems today. The data science stages are organized into a data science pipeline, where data might flow from one stage in the pipeline to the next. These data science stages generally perform different tasks such as data acquisition, data preparation, storage, feature engineering, modeling, training, evaluation of the machine learning model, etc. In order to design and build software systems with data science stages effectively, we must understand the structure of the data science pipelines. Previous work has shown that understanding the structure and patterns used in existing systems and literature can help build better systems (Shaw and Garlan, 1996; Gamma et al., 1993). In this work, we have taken the first step to understand the structure and patterns of DS pipelines.
Fortunately, we have a number of instances in both the state-of-the-art and practice to draw observations. In the literature, there have been a number of proposals to organize data science pipelines. We call such proposals DS Pipelines in theory. Another source of information is Kaggle, a widely known platform for data scientists to host and participate in DS competitions, share datasets, machine learning models, and code. Kaggle contains a large number of data science pipelines, but these pipelines are typically developed by a single data scientist as small standalone programs. We call such instances DS Pipelines in-the-small. The third source of DS pipelines are mature data science projects on GitHub developed by teams, suitable for reuse. We call such instances DS Pipelines in-the-large.
This work presents a study of DS pipelines in theory, in-the-small, and in-the-large. We studied 71 different proposals for DS pipelines and related concepts from the literature. We also studied 105 instances of DS pipelines from Kaggle. Finally, we studied 21 matured open-source data science projects from GitHub. For both Kaggle and GitHub, we selected projects that make use of Python to ease comparative analysis. In each setting, we answer the following overarching questions.
Representative pipeline: What are the stages in DS pipeline and how frequently they appear?
Organization: How are the pipeline stages organized?
Characteristics: What are the characteristics of the pipelines in a setting and how does that compare with the others?
This work attempts to inform the terminology and practice for designing DS pipeline. We found that DS pipelines differ significantly in terms of detailed structures and patterns among theory, in-the-small, and in-the-large. Specifically, a number of stages are absent in-the-small, and the pipelines have a more linear structure with an emphasis on data exploration. Out of the eleven stages seen in theory, only six stages are present in pipeline in-the-small, namely data collection, data preparation, modeling, training, evaluation, and prediction. In addition, pipelines in-the-small do not have clear separation between stages which makes the maintenance harder. On the other hand, the DS pipelines in-the-large have a more complex structure with feedback loops and sub-pipelines. We identified different pipeline patterns followed in specific phase (development/post-development) of the large DS projects. The abstraction of stages are stricter in-the-large having both loosely- and tightly-coupled structure.
Our investigation also suggest that DS pipeline is a well used software architecture but often built in ad hoc manner. We demonstrated the importance of standardization and analysis framework for DS pipeline following the traditional software engineering research on software architecture and design patterns (Shaw and Garlan, 1996; Lientz et al., 1978; Parnas et al., 1985). We contributed three representations of DS pipelines that capture the essence of our subjects in theory, in-the-small, and in-the-large that would facilitate building new DS systems. We anticipate our results to inform design decisions made by the pipeline architects, practitioners, and software engineering teams. Our results will also help the DS researchers and developers to identify whether the pipeline is missing any important stage or feedback loops (e.g., storage and evaluation are missed in many pipelines).
The rest of this paper is organized as follows: in section §2, we present our study of DS pipelines in theory. Section §3 describes our study of DS pipelines in-the-small. In section §4, we describe our study of DS pipelines in-the-large. Section §5 discusses the implications, section §6 describes the threats to the validity, section §7 describes related work, and section §8 concludes.
DS Pipeline in Theory
Data Science. Data Science (DS) is a broad area that brings together computational understanding, inferential thinking, and the knowledge of the application area. Wing (Wing, 2019) argues that DS studies how to extract value out of data. However, the value of data and extraction process depends on the application and context. DS includes a broad set of traditional disciplines such as data management, data infrastructure building, data-intensive algorithm development, AI (machine learning and deep learning), etc., that covers both the fundamental and practical perspectives from computer science, mathematics, statistics, and domain-specific knowledge (Berman et al., 2018; Todd and Dietrich, 2017). DS also incorporates the business, organization, policy and privacy issues of data and data-related processes. Any DS project involves three main stages: data collection and preparation, analysis and modeling, and finally deployment (Wickham, 2019). DS is also more than statistics or data mining since it incorporates understanding of data and its pattern, developing important questions and answering them, and communicating results (Todd and Dietrich, 2017).
Data Science Pipeline. The term pipeline was introduced by Garlan with box-and-line diagrams and explanatory prose that assist software developers to design and describe complex systems so that the software becomes intelligible (Garlan, 2000). Shaw and Garlan have provided the pipes-and-filter design pattern that involves stages with processing units (filters) and ordered connections (pipes) (Shaw and Garlan, 1996). They also argued that pipeline gives proper semantics and vocabulary which helps to describe the concerns, constraints, relationship between the sub-systems, and overall computational paradigm (Garlan, 2000; Shaw and Garlan, 1996). By data science pipeline (DS pipeline), we are referring to a series of processing stages that interact with data, usually acquisition, management, analysis, and reasoning (Olson et al., 2016; Nguyen et al., 2019). The sequential DS stages from acquisition, to cleaning/curation, to modeling, and so on are referred to as data science pipeline. A DS pipeline may consist of several stages and connections between them. The stages are defined to perform particular tasks and connected to other stage(s) with input-output relations (Amershi et al., 2019). However, the definitions of the stages are not consistent across studies in the literature. The terminology vary depending on the application context and focus.
Different study in the literature presented DS pipeline based on their context and desiderata. No study has been conducted to unify the notions DS pipeline and collect the concepts (Sculley et al., 2015). While designing a new DS pipeline (Wirth and Hipp, 2000), dividing roles in DS teams (Kim et al., 2016), defining software process in data-intensive setting (Wan et al., 2019), identifying best practices in AI and modularizing DS components (Amershi et al., 2019), it is important to understand the current state of the DS pipeline, its variations and different stages. To understand the DS pipelines and compare them, we collected the available pipelines from the literature and conducted an empirical study to unify the stages with their subtasks. Then we created a representative DS pipeline with the definitions of the stages. Next, we present the methodology and results of our analysis of DS pipelines in theory.
We searched for the studies published in the literature and popular press that describes DS pipelines. We considered the studies that described both end-to-end DS pipeline or a partial DS pipeline specific to a context. First, we searched for peer-reviewed papers published in the last decade i.e., from 2010 to 2020. We searched the terms “data science pipeline”, “machine learning pipeline”, “big data lifecycle”, “deep learning workflow”, and the permutation of these keywords in IEEE Xplore, ACM Digital Library and Google Scholar. From a large pool, we selected 1,566 papers that fall broadly in the area of computer science, software engineering and data science. Then we analyzed each article in this pool to select the ones that propose or describe a DS pipeline. We found many papers in this collection use the terms (e.g., ML lifecycle), but do not contain a DS pipeline. We selected the ones that contain DS pipeline and extracted the pipelines (screenshot/description) as evidence from the article. The extracted raw pipelines are available in the artifact accompanied by this paper (Anonymous, 2021). Thus, we found 46 DS pipelines that were published in the last decade.
Besides peer-reviewed papers, by searching the keywords on web, we collected the DS pipelines from US patent, industry blogs (e.g., Microsoft, GoogleCloud, IBM blogs), and popular press published between 2010 and 2020. After manual inspection, we found 25 DS pipelines from this grey literature. Thus, we collected 71 subjects (46 from peer reviewed articles and 25 from grey literature) that contain DS pipeline. We used an open-coding method to analyze these DS pipelines in theory (Anonymous, 2021) .
1.2. Labeling Data Science Pipelines
In the collected references, DS pipeline is defined with a set of stages (data acquisition, data preparation, modeling, etc.) and connections among them. Each stage in the pipeline is defined for performing a specific task and connected to other stages. However, not all the studies depict DS pipelines with the same set of stages and connections. The studies use different terminologies for defining the stages depending on the context. To be able to compare the pipelines, we had to understand the definitions and transform them into a canonical form. For a given DS pipeline, identifying their stages and mapping them to a canonical form is often challenging. The sub-tasks, overall goal of the project, utilities affect the understanding of the pipeline stages. To counter these challenges, we used an open-coding method to label the stages of the pipelines.
Two authors labeled the collected DS pipelines into different criteria. Each author read the article, understood the pipeline, identified the stages, and labeled them. In each iteration, the raters labeled 10% of the subjects (7-8 pipelines). The first 8 subjects were used for training and forming the initial labels. After each iteration, we calculated the Cohen’s Kappa coefficient (Viera et al., 2005), identified the mismatches, and resolved them in the presence of a moderator, who is another author. Thus, we found the representative DS pipeline after rigorous discussions among the raters and the moderator. The methodology of this open-coding procedure is shown in Figure 1. The entire labeling process was divided into two phases: 1) training, and 2) independent labeling.
Training: The two raters were trained on the goal of this project and their roles. We randomly selected eight subjects for training. First, the raters and the moderator had discussions on three subjects and identified the stages in their DS pipeline. Thus, we formed the commonly occurred stages and their definitions, which were updated through the entire labeling and reconciliation process later. After the initial discussion and training, the raters were given the already created definitions of the stages and one pipeline from the remaining five for training. The raters labeled this pipeline independently. After labeling the pipeline, we calculated the agreement and conducted a discussion session among the raters and the moderator. In this session, we reconciled the disagreements and updated the labels with the definitions. We continued the training session until we got perfect agreement independently. The inter-rater agreement was calculated using Cohen’s Kappa coefficient (Viera et al., 2005). A higher () indicates a better agreement. The interpretation of of is shown in Figure 1(a). In the discussion meetings, the raters discussed each label (both agreed and disagreed ones) with the other rater and moderator, argued for the disagreed ones and reconciled them. In this way, we came up with most of the stages and a representative terminology for each stage including the sub-tasks.
Independent labeling: After completing the training session, the rest of the subjects were labeled independently by the raters. The raters labeled the remaining 63 labels: 7 subjects (10%) in each of the 9 iterations. The distribution of after each independent labeling iteration is shown in Figure 1(b). In each iteration, first, the raters had the labeling session, and then the raters and moderator had the reconciliation session.
Labeling. The raters labeled separately so that their labels were private, and they did not discuss while labeling. The raters identified the stages and connections between them, and finally labeled whether the DS pipeline involves processes related to cyber, physical or human component in it. In independent labeling, we found almost perfect agreement ( = 0.83) on average. Even after high agreement, there were very few disagreements in the labels, which were reconciled after each iteration.
Reconciling. Reconciliation happened for each label for the subject studies in the training session, and the disagreed labels for the studies in independent labeling session. In training session, the reconciliation was done in discussion meetings among the raters and the moderator, whereas for the independent labels, reconciliation was done by the moderator after separate meetings with the two raters. For reconciliation, the raters described their arguments for the mislabeled stages. For a few cases, we had straightforward solution to go for one label. For others, both the raters had good arguments for their labels, and we had to decide on that label by updating the stages in the definition of the pipeline. All the labeled pipelines from the subjects are shared in our paper artifact (Anonymous, 2021).
Furthermore, after finishing labeling the pipelines stages, we also classified the subject references into four classes based on the overall purpose of the article. First, after a few discussions, the raters and moderator came up with the classes. Then, the raters classified each pipeline into one class. We found disagreements in 6 out of 71 references, which the moderator reconciled with separate meetings with the two raters. Based on our labeling, the literature that we collected are divided into four classes: describe or propose DS pipeline, survey or review, DS optimization, and introduce new method or application. Next, we are going to discuss the result of analyzing the DS pipelines in theory.
2. Representative Pipeline in Theory
The labeled pipelines with their stages are visually illustrated in the artifact Table 3. We found that pipelines in theory can be both software architecture and team processes unlike pipelines in-the-small and in-the-large. Through the labeling process, we separated those team processes (25 out of 71), which are discussed in §2.4.
RQ1a: What is a representative definition of the DS pipeline in theory? From the empirical study, we created a representative rendition of DS pipeline with 3 layers, 11 stages and possible connections between stages as shown in Figure 3. Each shaded box represents a DS stage that performs certain sub-tasks (listed under the box). In the preprocessing layer, the stages are data acquisition, preparation, and storage. The preprocessing stage study design only appeared in team process pipelines that comprise requirement formulation, specification, and planning, which are often challenging in data science. The algorithmic steps and data processing are done in the model building layer. Modeling does not necessarily imply the existence of an ML component, since DS can involve custom data processing or statistical modeling. Post-processing layer includes the tasks that take place after the results have been generated. The DS pipeline stages are described in Table 1.
RQ1b: What are the frequent and rare stages of the DS pipeline in theory? The frequency of stage can depend on the focus of the pipeline or its importance in certain context (ML, big-data management). Among 46 DS pipelines (which are not team processes), Figure 4 shows the number of times each stage appears. A few pipelines present stages with broad terminology that fit multiple stage-definitions. In those cases, the pipelines were labeled with the fitted stages and counted multiple times. Modeling, data preparation, and feature engineering appear most frequently in the literature. While modeling is present in 93% of the pipelines, other model related stages (feature engineering, training, evaluation, prediction) are not used consistently. Often training is not considered as a separate stage and included inside the modeling stage. Similarly, we found that evaluation and prediction are often not depicted as separate stages. However, by separating the stages and modularizing the tasks, the DS process can be maintained better (Amershi et al., 2019; Sculley et al., 2015). The pipeline created with the most number of stages (11) is provided by Ashmore et al. (Ashmore et al., 2021). On the other hand, about 15% of the pipelines from the literature are created with a minimal number (3) of stages. Among them, 80% are ML processes and falls in the category of DS optimizations. We found that these pipelines are very specific to particular applications, which include context-specific stages like data sampling, querying, visualization, etc., but do not cover most of the representative stages. A pipeline in theory may not require all representative stages, since it can have novelty in certain stages and exclude the others. However, the representative pipeline provides common terminology and facilitate comparative analysis.
Clearly, preprocessing and model building layers are considered in almost all of the studies. In most of the cases, the pipelines do not consider the post-processing activities (interpretation, communication, deployment). These pipelines often end with the predictive process and thus do not follow up with the later stages which entails how the result is interpreted, communicated and deployed to the external environment. Miao et al. argued that overall lifecycle management tasks (e.g., model versioning, sharing) are largely ignored for deep learning systems (Miao et al., 2017b). Previous studies also showed that significant amount of cost and effort is spent in the post-development phases in traditional software lifecycle (Lientz et al., 1978; Rajlich, 2014). In data-intensive software, the maintenance cost can go even higher with the high-interest technical debt in the pipeline (Sculley et al., 2014). Therefore, post-processing stages should be incorporated for a better understanding of the impact of the proposed approach on maintenance of the DS pipeline.
3. Organization of Pipeline Stages in Theory
RQ2: How are pipeline stages connected to each other? In Figure 3, for simplicity, we depicted the DS pipeline as a mostly linear chain. However, our subject DS pipelines often have non-linear behavior. In any stage, the system might have to return to the previous stage for refinement and upgrade, e.g., if a system faces a real-world challenge in modeling, it has to update the algorithm which might affect the data pre-processing and feature engineering as well. Furthermore, the stages do not have strict boundaries in the DS lifecycle. In Figure 3, two backward arrows, from feature engineering and evaluation, indicate feedback to any of the previous stages. Although in traditional software engineering processes (e.g., waterfall model, agile development, etc.), feedback loop is not uncommon, in DS lifecycle, there are multiple stakeholders and models in a distributed environment which makes the feedback loops more frequent and complex. Sculley et al. pointed that DS tasks such as sampling, learning, hyperparameter choice, etc. are entangled together so that Changing Anything Changes Everything (CACE principle) (Sculley et al., 2015), which in turn creates implicit feedback loops that are not depicted in the pipelines (Van Der Weide et al., 2017; Chen and Zhang, 2014; Rehman et al., 2016; Gandomi and Haider, 2015). The feedback loops inside any specific layer are more frequent than the feedback loops from one layer to another. Also, a feedback loop to a distant previous stage is expensive. For example, if we do data preparation after evaluation then the intermediate stages also require updates.
4. Characteristics of the Pipelines in Theory
RQ3: What are the different types of pipelines available in theory? The context and requirements of the project can influence pipeline design and architecture (Garcia et al., 2018). Here, we present the types of pipelines with different characteristics that are available in theory. We classified each subject in our study into four classes based on the overall goal of the article. The most of the pipelines in theory (39%) are describing or proposing new pipelines to solve a new or existing problem. About 31% of the pipelines are on reviewing or comparing the existing pipelines. The third group of DS pipelines (14%) are intended to optimize a certain part of the pipeline. For example, Van Der Weide et al. proposes a pipeline for managing multiple versions of pipelines and optimize performance (Van Der Weide et al., 2017). Most of the pipelines in this category are application specific and include very few stages that are necessary for the optimization. Fourth, some research introduce new application or method and present within the pipeline. We observed that there is no standard methodology to develop comparable and inter-operable DS pipelines. Using the labeling methodology shown in Figure 1, we labeled each pipeline and found three types of DS pipelines in the literature: 1) ML process, 2) big data management process, and 3) team process.
ML process: 46% of all the pipelines we found in the literature are describing machine learning processes. The recent advent of artificial intelligence, supervised learning and deep learning has led to more DS systems that involve ML components. The pipelines in this category emphasize the algorithmic process, learning patterns, and building predictive models. However, the post-processing stages are rare in these type of pipelines. The ML pipelines are often thought of as algorithmic process in the laboratory scenario. But as mentioned in (Ashmore et al., 2021), incorporating the post-processing stages would be desired to ensure safe real-world deployment of such pipelines.
Big data management: The references in this category present DS pipelines that manage a large amount of data or describes a framework (software-hardware infrastructure) for data processing but do not contain machine learning components in the pipeline. Processing large amount data often requires specific algorithms and engineering methods for efficiency and further processing. We found that 18% of all the subject studies fall in this category.
Team process: We also found some DS pipelines that are not describing DS software architecture. These pipelines describe workflow of human activities that needs to be followed in a DS pipeline. These studies present a high-level view for building DS component in a team environment. The data science teams require specific expertise and management to build successful DS pipelines (Kim et al., 2016; Amershi et al., 2019). In this paper, in §3 and §4, we are only focusing on DS pipeline as software architecture, and therefore, we did not compare the team process pipelines in the rest of this section.
We identified whether the pipelines involve cyber, physical or human process, using our labeling process described in section §2.1.2. Cyber processes refer to activities that involve automated systems and machinery computations. Since modern DS systems involves large amount of data and requires extensive computation, all of the pipelines include cyber component in it. Physical processes include the activities which require real-world connections with the system. For example, collecting data using mobile sensors or cameras is a physical process. Although 23% of the big data pipelines include physical processes, only 9% of the ML pipelines include that in the pipeline. In many DS systems, developers or researchers participate in the pipelines actively to make decisions that need human interventions (Todd and Dietrich, 2017; Van Der Weide et al., 2017). For example, in many DS systems, analytical model validation, troubleshooting, data interpretation is necessary which requires human involvement. However, only 13% of the pipelines acknowledged human involvement in the pipeline.
DS Pipeline in-the-Small
Similar to the DS pipelines in large systems and frameworks, for a very specific data science task (e.g., object recognition, weather forecasting, etc.), programmers build pipeline. Different stages of the program perform a specific sub-task and connect with the other stages using data-flow or control-flow relations. In this section, we described such DS pipelines in-the-small.
We collected 105 DS programs from Kaggle competition notebooks (Kaggle, 2021a). Kaggle is one of the most popular crowd-sourced platforms for DS competitions, owned by Google. Besides participating in competitions, data scientists, researchers, developers collaborate to learn and share DS knowledge in variety of domains. The users and organizations can host a DS competition in Kaggle to solve real-world problems. A competition is accompanied by a dataset and prize money. Many Kaggle solutions have resulted in impactful DS algorithms and research such as neural networks used by Hinton and Dahl (Dahl et al., 2014), improving the search for the Higgs Boson at CERN (Jepsen, 2014), etc. We chose Kaggle solutions to analyze DS pipeline for three reasons: 1) all programs perform a DS task and provide solution to a well specified problem associated with a dataset, 2) solutions with the highest number of votes are well accepted solutions for a specific problem, and 3) the problems cover a wide range of domains.
There are 331 completed competitions in Kaggle to date. They categorized the competitions into Featured, Research, Recruitment, Masters, Analytics, Playground and Getting started. We collected solutions of all the competitions from each category except Getting started and Playground (these two categories are intended to serve as DS tutorials and toy projects). First, we filtered the competitions for which there are solutions available (many old competitions do not contain any public solution). We found 138 such competitions. For a given competition problem, we selected the most voted solution which has at least 10 votes. Thus, we got 105 top-rated DS solutions for analyzing pipelines in-the-small. This selection and pipeline creation process is shown in Figure 5.
All of the DS programs are written in Python using ML libraries like Keras, Scikit-learn, Tensorflow, etc. These packages provide high-level Application Programming Interfaces (APIs) for performing a specific task on data or model. We parsed the programs into Abstract Syntax Tree (AST) and collected all the API calls from the programs. Then the functionality of an API is used to identify the stage of the pipeline. We extracted the temporal order of API calls to identify the stages. Standard static analysis of the Python programs facilitate the extraction process. Our analysis suggests that the DS programs follow a linear structure with less than 4% AST nodes being conditional or loops. Wang et al. proposed a similar approach for extracting external dependencies in Jupyter Notebooks by creating an API database and analyzing AST (Wang et al., 2021).
We created a dictionary by mapping each API collected from the programs, to one of the 11 stages of the DS pipeline described in section §2. During the mapping, we excluded the generic APIs from the dictionary. For example, model.summary() is used to print the model parameters and does not represent any stage of the pipeline. For creating the dictionary, we taken a two-fold approach. First, we understand the context of the program and API usage. Second, we look at the API documentation to confirm the corresponding pipeline stage. We found that DS APIs are definitive in their operations and well-categorized by the library. For example, the APIs in Keras (Keras, 2021a) and Scikit-learn (Keras, 2021b) are grouped into preprocessing, models, etc. Our API-dictionary was manually validated by a second-rater and moderator who labeled DS pipelines in section §2. Then, we built a tool which takes the API dictionary and DS program, and automatically creates the DS pipeline. For a sequence of APIs with the same stage, we abstracted them into a single stage. As an example, Figure 5 shows a DS pipeline created from a Kaggle solution (Kaggle, 2021b). Each stage in the pipeline (e.g., ACQ, PRP) represents one or more API usages. The arrows in the pipeline denote the temporal sequence of stages. Note that, one stage can appear multiple times in a pipeline. The API dictionary, Kaggle programs, and tool to generate the pipelines is shared in the paper artifact (Anonymous, 2021).
2. Representative Pipeline in-the-Small
RQ4: What are the stages of DS pipeline in-the-small? Among the 11 pipeline stages described in Figure 3, we found only 6 stages in the DS programs that are depicted in Figure 6. Other stages (e.g., storage, feature engineering, interpretation, communication, deployment) are not found in these programs because these stages occur while building a production-scale large DS system and often not present in the DS notebooks. Therefore, the pipeline in DS programs consists of the subset of pipeline stages in theory.
We summarized the frequency of each stage of the DS programs in Figure 7. Among 105 programs, data acquisition and data preparation are present in almost all of them. Surprisingly, modeling is present in only 70% of the programs. We found that, in many programs, no modeling APIs had been used because developers did not use any built-in ML algorithm from libraries, e.g., LogisticRegression, LSTM, etc. In these cases, the developers use data-processing APIs on the training data to build custom model, e.g., this notebook (Kaggle, 2021c) uses data preparation APIs to produce results. To enable more abstraction of the stages in these pipelines, further modularization is necessary, which has been investigated in RQ8.
Evaluation is a tricky stage of the DS pipeline. Developers have to choose the appropriate metric and methodology to evaluate their model. Based on the evaluation result, the model is updated over multiple iterations. We found that, besides using metrics, in many cases, evaluation requires human understanding and comparison of the result produced by the model. The reason for having less number of evaluation stage in the pipeline is that often the developers evaluate the performance by plotting and visualizing the result. Since the visualization APIs are not considered as evaluation stage, we found this stage less frequently in pipelines. Also, many programs directly go to the prediction stage without going to evaluation stage at all. Furthermore, notebooks are often used for experimentation purposes so that many computations are performed during development but eliminated when the notebooks are shared (Kery et al., 2018). For example, one developer might try a number of classifiers and evaluate their accuracy. After finding the best performing classifier, it can be the only one shared in the notebook. Therefore, we experienced many missing stages in the pipeline in-the-small. The complex DS tasks require several computations which might not be used in producing the final prediction, but definitely should be considered as part of the pipeline.
3. Pipeline Organization in the Small
RQ5: How are the stages connected with each other in pipeline in-the-small? To answer RQ5, we considered each occurrence of the stages in a DS program and looked at its previous and next stage. In Figure 8, we showed which stages are followed or preceded by each stage. We found that data preparation can occur before or after all other stages. Apart from that, data acquisition is followed by data preparation most of the time, which in turn is followed by modeling. Modeling is followed mostly by training, which in turn is followed by prediction. Evaluation is mostly surrounded by prediction and data preparation. From Figure 8, we can also find some most occurring feedback loop: evaluation to preparation, evaluation to modeling and prediction to modeling.
Data preparation tasks (e.g., formatting, reshaping, sorting) are not limited to just before the modeling stage, rather it is done on a whenever-needed basis. For example, in the following code snippet from a Kaggle competition (Kaggle, 2021d), while creating model-layers, data preprocessing API has been called in line 2.
The modeling stage is always surrounded by other stages of the pipeline. However, there is often a loop around modeling, training, evaluation, and prediction. Modeling often repeats many times to improve the model over multiple iterations. For example, in the following Kaggle code snippet (Kaggle, 2021e), the model is created and trained multiple times to find the best one.
All of the DS programs fail to maintain a good separation of concerns (Dijkstra, 1982) between stages. Strong abstraction boundaries help to make the program modular and easy-to-maintain (Parnas et al., 1985; Pan and Rajan, 2020, 2022). In addition, a good DS solution should not only compute better predictive result, but also facilitate software engineering activities e.g., debugging, testing, monitoring (Hill et al., 2016). However, we found that stages are often tangled with other stages (Kiczales et al., 1997; Prehofer, 1997; Calder et al., 2003) across the pipelines. The code for one stage is interspersed with the code for other stages. For example, while building the deep learning network (modeling), the developers often switch to different data preparation tasks, e.g., reshaping, resizing (Islam et al., 2019, 2020), which tangles data preparation concern with the modeling concern. We observed some early attempts to adopt modular design practices. For instance, this notebook (Kaggle, 2021f) separated code into different high-level stages, namely, preparation, feature extraction, exploratory data analysis (EDA), topic model, etc. These high-level pipelines can improve the abstraction, which further enable the maintainability, and reusability (Rule et al., 2018). In some scenarios, reuse or maintenance might not be desired for pipelines in-the-small. However, to enhance readability (Kery et al., 2018) and repeatability (Hill et al., 2016) and ease of testing, debugging or repairing (Wardat et al., 2021, 2022), more attention on modular design practices is needed for DS pipelines.
We found that new data sources are added, new features are identified, and new values are calculated incrementally in the pipeline which evolves organically. This results in a large number of data preprocessing tasks like sampling, joining, resizing along with random file input-output. This is called pipeline jungles (Sculley et al., 2015), which causes technical debt for DS systems in the long run. Pipeline jungles are hard to test and any small change in the pipeline will take a lot of effort to integrate. The situation gets worse in case of larger DS pipelines, where several data management activities (e.g., clean, serve, validate) are necessary through the pipeline in different stages (Polyzotis et al., 2017, 2018). The recommended way is to think about the pipeline holistically and scrape the pipeline jungle by redesigning it, which in turn takes further engineering effort (Sculley et al., 2015). We found that the large DS projects, which are discussed in §4, isolate the data preparation tasks into separate files and modules (Britz, 2018; Trieu, 2018; Sandberg, 2018; Yu et al., 2018), which alleviates the pipeline jungles problem. So, DS pipeline in-the–small needs further IDE (e.g., Jupyter Notebook, etc.) support and methodologies for code isolation and modularization.
4. Characteristics of Pipelines in-the-Small
RQ6: What are the patterns in pipeline in-the-small and how it compares to pipeline in theory? We have not found many stages from Figure 3, e.g., feature engineering, interpretation, communication, in pipeline in-the-small. One reason is that the low-level pipeline extracted from the API usages cannot capture some stages. For example, even if a developer is conducting feature engineering, the used APIs might be from the data preparation stage. Fortunately, we found many Kaggle notebooks that are organized by the pipeline stages. We visited all the 105 Kaggle notebooks in our collection and extracted these high-level pipelines manually. Unlike the low-level pipelines (extracted using API usages), a high-level pipeline consists of the stages abstracted by the developers.
The Kaggle notebooks follow literate programming paradigm (Wagner, 2020; Rule et al., 2018), which allows the developers to describe code using rich text and separate them into sections. We found that 34 out of 105 notebooks divided the code into stages. We collected those stages from the Kaggle notebooks. Furthermore, we labeled these notebooks into the 11 stages from DS pipeline in theory by two raters, and extracted the stages that are not present in theory. The extracted high-level pipelines and labels are available in the paper artifact (Anonymous, 2021).
We observed that no notebooks specify these stages: storage, interpretation, communication, and deployment. These DS programs are not production-scale projects. Therefore, they do not include the post-processing stages in the pipeline. The most common stages are modeling (79%), data preparation (62%), data acquisition (53%), and feature engineering(35%), which is aligned with the finding of DS pipeline in theory. In addition, we found these stages which are not present in theory: library loading, exploratory data analysis (EDA), visualization. Among them EDA has been used most of the times (43%) and covered the most part of those pipeline. Before going to the modeling and successive stages, a lot of effort is given on understanding the data, compute feature importance, and visualize the patterns, which help to build models quickly in later stages (Ashmore et al., 2021).
Furthermore, some notebooks present library loading as separate stage. We observed that choosing appropriate library/framework and setting up the environment is an important step while developing pipeline in-the-small. We also found that data visualization is an recurring stage mentioned by the developers. Visualization can be done for EDA or feature engineering (before modeling), or for evaluation (after modeling). Based on these observations we updated the representative pipeline in-the-small in Figure 9. The high-level pipeline provides an overall representation of the system, which can be leveraged to design software process. It would be beneficial for the developers to close the gap between the low-level and the high-level pipeline by identifying the tangled stages.
DS Pipeline in-the-Large
The DS solutions described in the previous section are specific to a given dataset and a well-defined problem. However, there are many DS projects which are large, not limited to a single source file, and contains multiple modules. These solutions are intended to solve more general problems which might not be specific to a dataset. For example, the objective of the Face Classification project in GitHub (Arriaga, 2018) is to detect face from images or videos and classify them based on gender and emotion. This problem is not specific to a particular dataset and the scope is broader compared to the Kaggle solutions. We collected such top-rated DS projects from GitHub to analyze DS pipeline in-the-large.
Biswas et al. published a dataset containing top rated DS projects from GitHub (Biswas et al., 2019). From the list of projects in this dataset, we filtered mature DS projects having more than 1000 stars. Thus, we found 269 mature GitHub projects. However, there are many projects in this list which are DS libraries, frameworks or utilities. Since we want to analyze the pipeline of data science software, we removed those projects. Finally, we also removed the repositories which serve educational purposes. Thus, we found a list of 21 mature open-source DS projects. The list of projects, and their purpose are shown in Table 2.
For each project, we created two pipelines: high-level pipeline and low-level pipeline. For creating the high-level pipeline, we manually checked the project architecture, module structure and execution process. This gave us a good understanding of the source file organization and linkage between modules. After identifying the high-level pipeline and execution sequences of the source files, we used the same API based method used to analyze Kaggle programs in the previous section, to create low-level pipeline of these GitHub projects. The methodology of selecting and extracting pipelines from the GitHub projects is shown in Figure 10.
For example, the project QANet (Yu et al., 2018) is intended to do machine reading comprehension. Here, Python has been used as the primary language, and shell script has been used for data downloading and project setup. The high-level pipeline for QANet includes the stages: data acquisition, data preparation, modeling, training, evaluation and prediction. In the beginning, config.py file integrates the modules (preparation, modeling, and training) and provides an interface to configure a model by specifying dataset and other parameters. Then, the file evaluate.py is executed to perform the evaluation and prediction. For the low-level pipeline, for a specific file, we used the API based analysis to generate the pipeline, which was used to analyze pipeline in-the-small. For instance, in the project QANet, although model.py serves modeling at a high level, it also does data preparation, training, and evaluation, when APIs are considered. In addition to the pipeline stages, we also identified a few other properties of each project: 1) number of contributors, 2) AST count, 2) technology/language used, 3) entry points and 4) execution sequence. We leveraged the Boa infrastructure (Dyer et al., 2013, 2015) to analyze the different properties of the projects. These properties helped us to categorize and analyze the pipeline in-the-large. The details of the projects are available in the paper artifact (Anonymous, 2021).
The projects are from various domains: object detection, face classification, automated driving, speech synthesis, number plate recognition, predict time series sequence, etc. The number of developers in each project ranges between 1 and 40 with an average of 8. Among 21 projects, 16 of them are developed by teams and 5 of them are developed by individuals. The primary language used to develop these projects is Python.
2. Representative Pipeline in-the-Large
Compared to the Kaggle programs, we found a significant difference in the pipeline of large DS projects. Because of the larger size of the projects, the pipeline architecture is different. All the projects contain multiple source files for handling different tasks (e.g., modeling, training) and about 50% of the projects organize the source files into modules (e.g., utils, preprocessing, model, etc.).
RQ7: What is the representative DS pipeline in-the-large? Each of the projects contains six stages described in Figure 6: acquisition, preparation, modeling, training, evaluation, and prediction. However, since the projects are not coupled to a specific dataset and they solve a more general problem, the projects are not limited to one single pipeline. We found that the pipeline of each project is divided into two phases: 1) development phase and 2) post-development phase, which is depicted in Figure 11.
In development phase, the main goal is to build a model that solves the problem in general. A base dataset is used to build the model that would be used for other future datasets. After completing a modeling, training, evaluation loop, the final model is created and saved as an artifact. Afterwards, the projects also create model interfaces, which lets the user modify and exploit the model in the post-development phase. Finally, the model artifact is saved as a source file or some model archiving formats. For example, the project Person-Blocker (Woolf, 2018) and Speech-to-Text-WaveNet (Kim, 2018) saved the model in the source file (model.py) and lets the users train the model in the next phase. On the other hand, the project KittiSeg (Teichmann, 2018) and Autopilot (Alexis Chan, 2017) saved the built model artifact in JSON format (.json) and checkpoint format (.ckpt) respectively. We observed that the evaluation and prediction is often not the main goal in this phase; rather, building an appropriate model and making it available for further usage is the central activity.
In post-development phase, the users access the pre-built model and use that for prediction. After acquiring data, a few preprocessing steps are needed to feed the model. In all of the projects under this study, we found that the development phase is similar. However, we identified three different patterns in the post-development phase which are shown in Figure 11. First, the users can modify the model by setting its hyperparameters and use that to make prediction on a new dataset. Second, the users can use the model as-it-is and train the model on the new dataset to make prediction. Third, the users can also download the pre-trained model and directly leverage that for prediction. Finally, at the end of this phase, the prediction result is obtained.
The post-development phase in the pipeline enabled software reusability of the models. All of these projects have instructions in their readme or documentation explaining the usage and customization. For example, the project Deep ANPR (Earl, 2016) provides instructions for obtaining large training data, retraining the models, and build it for prediction. However, not all the projects enable reusability in the development pipelines. Only a few of them provides access to the modules by importing in new development scenario. For instance, Darkflow (Trieu, 2018) let users access the darkflow.net.build module and use it in new application development. To increase the reusability of DS programs, it would be desired to consider similar access to the development pipeline of these large projects.
3. Organization of DS Pipeline in-the-Large
RQ8: How are the stages connected in pipeline in-the-large? The abstraction in DS projects is stricter than the DS programs described in §3. The projects are built in a modular fashion, i.e., one source file for a broad task (e.g., train.py, model.py). However, inside one specific file, there are many other possible stages, especially data preprocessing appears inside all the source files. In addition, the module connectivity is not linear. All of the modules use external libraries for performing different tasks. As a result, there are a lot of interdependencies (both internal and external) in the DS pipeline. One immediate difference of these pipelines with traditional software is DS pipelines are heavily dependant on the data. For example, the project Speech-to-Text-WaveNet (Kim, 2018) requires a certain format of data. When we want to use that in a new situation, the data properties might be different. So, the usage pipelines would have a few additional stages. In some cases, the original pipeline is modified. Here, there are many sub-pipelines work together to build a large pipeline. However, we have not found any framework or common methodology these software are using. The different patterns of DS pipelines seek more advanced methodology or framework to build DS pipeline and release for production.
4. Characteristics of Pipelines in-the-Large
RQ9: What are the patterns found in the pipelines? The pipelines found in this setting can be categorized into 1) loosely coupled and 2) tightly coupled, based on their modularity. A high number of contributors in the project resulted in loosely coupled pipelines. We found the loosely coupled ones are designed in a modular fashion and one module (e.g., data cleaning, modeling) is designed to be used by other modules. Usually, there are multiple entry-points in a loosely coupled pipeline and user has more flexibility. On the other hand, in a tightly coupled pipeline, the modules are stricter and integrated tightly with other modules. There is only one or two entry-points to the pipeline, which automatically calls the other modules. We found that the projects with 6 or more contributors (75%) followed a loosely coupled architecture and projects with 1 to 5 contributors followed a tightly coupled architecture.
Although all the project under this study are written using Python, no project is using any common tool that integrates the DS modules and provides interface to the pipeline. Today, continuous integration and deployment (CI/CD) tools are widely used in traditional software lifecycle to automate compilation, building, and testing (Hilton et al., 2017; Karlaš et al., 2020). Additionally, from our subject studies of pipelines in theory, we found some CI/CD tools designed for ML pipelines available (MLOps, 2020; Microsoft Blog, 2019; Hong and Hunter, 2017). Surprisingly, here we found no projects in pipeline in-the-large are using any CI/CD tools. However, the projects demonstrate the need of CI/CD in the repositories. In most of the projects, the environment setup and access to functionalities are configured through command lines scripts (Alexis Chan, 2017; Ruan, 2019). Some projects used docker container (Sandberg, 2018; Arriaga, 2018; Woolf, 2018; Kim, 2018) to set up the environment and run the pipeline. A few others used Python notebooks that call different modules to integrate the pipeline stages (Dat Tran, 2018; Abdulla, 2017; Paino, 2017). 7 out of 21 projects used shell script for integration (e.g., sending HTTP request to download data, model reuse, etc.) (Yu et al., 2018; Qi, 2019). Although CI/CD frameworks e.g., TravisCI, GitHub Actions, Microsoft Azure DevOps are well established for traditional software such as web applications, several challenges remain for DS pipelines. Karlaš et al. outlined the probabilistic nature of ML testing as a major CI/CD challenge and pointed out the gap between recent theoretical development of CI/CD in DS and their usage in practice (Karlaš et al., 2020). Hence, further research is needed to investigate the practical challenges of using CI/CD in data science projects.
Discussion
Through our survey, empirical study, and analysis, we presented the state of data science pipeline that describes its semantics, design concerns, and the overall computational paradigm. Furthermore, the findings show the importance of studying the pipeline structure reminiscing the traditional software engineering works on design patterns and architecture.
In Theory: We presented all the representative stages and subtasks that inform the terminology of DS pipelines to be used in future works. By comparing with the available pipeline categories e.g., ML process, big data, and team processes, similarities and divergences can be directly identified. The presence of implicit feedback loops and lack of post-processing stages suggest ad hoc pipeline construction at the present time. This paper takes the first step towards comparable and reusable pipeline construction.
In-The-Small: The novel API-based analysis can be utilized for mining, extracting, and statically analyzing pipelines. We also elicited the notion of high-level and low-level pipelines, where the high-level abstraction has more similarity with that in theory. However, low-level pipelines exhibit many differences such as missing some stages, sparse data preparation, lack of modularization. The gap between low-level pipeline and its presentation in high-level can be reduced by making pipeline specific features available in development environment e.g., pipeline template in Jupyter Notebook. Additionally, the low-level pipelines often had an important stage exploratory data analysis missing which incurs much time and effort. Pipeline versioning techniques that consider data, model, and source code will facilitate storing such intermediate stages.
In-The-Large: Different pipeline patterns emerged in development and post-development phase of the large projects, which suggest creating separate developer-centric and user-centric pipeline structure. In tightly-coupled projects, the abstraction of stages are contingent upon the project-specific requirements and internal/external dependencies, whereas, in loosely-coupled projects, opportunities remain to build reusable sub-pipelines that span over project boundaries. Finally, there is a need for building automated CI/CD tools for data science specific testing, deployment, and maintenance.
To researchers and tool builders. (1) Modularization of DS pipeline into stages is challenging over all three representations. Further works are needed for standardization of pipeline architecture e.g., defining the interfaces of stages, enumerating externally visible properties, identifying domain-specific constraints, to develop reusable and interoperable DS pipelines. (2) We showed potentials for automatic pipeline analysis framework based on static analysis and API specifications. A few future directions would be mining (sub-)pipelines patterns, build AutoML pipelines (Nguyen et al., 2022), and analyzing evolution. (3) We confirmed several antipatterns of pipelines that call for actions e.g., CACE principle, pipeline jungles, scarce post-processing, implicit feedback loop, CI/CD challenges. (4) Pipeline specific tool support is needed such as version control for data and models, storing intermediate results between stages.
To data scientists and engineers. (1) Pipelines are often built for a prototype in-the-small, which might not scale to a production level system. A well-designed pipeline in the early stage will help to identify key components, estimate cost, optimize, and manage risks better in the lifecycle. (2) The representative views of pipelines will serve as a checklist of stages and their connections. (3) Data and algorithms being the focus of DS pipeline, preprocessing and modeling activities are well understood and practiced by data scientists. However, they should emphasize more on including rigorous evaluation beyond accuracy such as robustness and fairness (Biswas and Rajan, 2020, 2021). (4) Many people with diverse backgrounds are involved in a DS pipeline. A pipeline with human-in-the-loop approach will benefit identifying collaboration points, decomposing tasks, and manage transdisciplinary teams. For example, a pipeline can encourage data scientists to choose a modeling technique that is maintainable. (5) Future work is necessary to identify the interactions of DS pipeline with the real world i.e., which stages receive inputs, when a checkpoint is saved, how results are disseminated, etc.
Threat to Validity
For building the pipelines from DS programs, we relied on the APIs. One threat might be, what happens if the developer does not use any API for completing a stage in the program. We examined this possibility and found that DS programs are heavily dependent on libraries and external APIs and ML tasks are always performed using library APIs. Additionally, we validated the API-to-stage dictionary with the API documentation and manual verification.
Another possible threat is that the Kaggle solutions might not be representative. We adopted a two-fold strategy to mitigate that threat. First, we selected the solutions with the most number of votes and at least 10 votes. Second, we manually verified each program whether it is an end-to-end DS solution. Since some most voted solutions are only for introduction and exploratory analysis of the dataset, by manual verification, we excluded those programs. The GitHub projects are also taken from a previously published dataset containing DS repositories. We further filtered them based on the number of stars and whether they perform a DS task.
Moreover, since the chosen DS programs from Kaggle and GitHub are using Python as the primary language, another question might be on the generalization of them as DS programs. According to GitHub and Stack Overflow, Python has become the most growing language in recent times (Inc., 2019; Robinson, 2017). In data science, Python is the most used language because of the availability of numerous ML, DL and data analysis packages such as Pandas, NumPy, TensorFlow, Keras, Caffe, Theano, Scikit-Learn and many more.
Related Work
Many studies presented ML pipeline in their own context, which can not be generalized for all DS systems. Garcia et al. focused on building an iterative process with three main phases: development, training and inference. They described the interpretation of data and code while integrating the whole lifecycle (Garcia et al., 2018). Polyzotis et al. presented the challenges of data management in building production-level ML pipeline in Google around three broad themes: data understanding, data validation and cleaning, and data preparation (Polyzotis et al., 2017, 2018). They also provided an overview of an end-to-end large-scale ML pipeline with a data point of view. Carlton E. Sapp defined ML concepts, business challenges, stages in the lifecycle, roles of DS teams with comprehensive end-to-end ML architecture (Sapp, 2017). This gives us a holistic understanding of the business processes (e.g., acquire, organize, analyze, deliver) of a DS project.
A few other studies try to capture the DS process by surveying and interviewing developers. Roh et al. surveyed the data collection techniques in the field of big data. They presented the workflow of data collection answering how to improve data or models in an ML system (Roh et al., 2019). Another study identified the software engineering practices and challenges in building AI applications inside Microsoft development teams (Amershi et al., 2019). They found some key differences in AI software process compared to other domains. They considered a 9-stage workflow for DS software development. Hill et al. interviewed experienced AI developers and identified problems they face in each stage (Hill et al., 2016). They also tried to compare the traditional software process and the AI process. Zhou presented her own view to build a better ML pipeline (Zhou, 2019). They presented three challenges in building ML pipelines: data quality, reliability and accessibility.
Some articles described ML applications and frameworks which present DS pipelines from industry. For example, Databricks provides high-level APIs for programming languages (Hong and Hunter, 2017). Team Data Science Process (TDSP) is an agile and iterative process to build intelligent applications inside Microsoft corporation (Severtson, 2017). In a US patent, the authors compared two data analytic lifecycles (Todd and Dietrich, 2017), and presented the difference in the set of parameters with respect to time and cost. CRoss Industry Standard Process for Data Mining (CRISP-DM) is a 6-stage comprehensive process model for data mining projects across any industry (Wirth and Hipp, 2000). Google Cloud Blog described the workflow of an AI platform (Google Cloud Blog, 2019). They explained tasks completed in each stage with respect to Google Cloud and TensorFlow(Abadi et al., 2016). Although there are many papers in the literature presenting DS pipeline, there is no comprehensive study that tries to understand and compare DS pipelines in theory and practice.
Conclusion
Many software systems today are incorporating a data science pipeline as their integral part. In this work, we argued that to facilitate research and practice on data science pipelines, it is essential to understand their nature. To that end, we presented a three-pronged comprehensive study of data science pipelines in theory, data science pipelines in-the-small, and data science pipelines in-the-large. Our study analyzed three datasets: a collection of 71 proposals for data science pipelines and related concepts in theory, a collection of 105 implementations of data science pipelines from Kaggle competitions to understand data science in-the-small, and a collection of 21 mature data science projects from GitHub to understand data science in-the-large. We have found that DS pipelines differ significantly between these settings. Specifically, a number of stages are absent in-the-small, and the DS pipelines have a more linear structure. The DS pipelines in-the-large have a more complex structure and feedback loops compared to the theoretical representations. We also contribute three representations of DS pipelines that capture the essence of our subjects in theory, in-the-small, and in-the-large.