DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines

Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Sri Vardhamanan, Saiful Haq, Ashutosh Sharma, Thomas T. Joshi, Hanna Moazam, Heather Miller, Matei Zaharia, Christopher Potts

Introduction

Language models (LMs) are enabling researchers to build NLP systems at higher levels of abstraction and with lower data requirements than ever before (Bommasani et al. 2021). This is fueling an exploding space of “prompting” techniques—and lightweight finetuning techniques—for adapting LMs to new tasks (Kojima et al. 2022), eliciting systematic reasoning from them (Wei et al. 2022; Wang et al. 2022b), and augmenting them with retrieved sources (Guu et al. 2020; Lazaridou et al. 2022; Khattab et al. 2022) or with tools (Yao et al. 2022; Schick et al. 2023). Most of these techniques are explored in isolation, but interest has been growing in building multi-stage pipelines and agents that decompose complex tasks into more manageable calls to LMs in an effort to improve performance (Qi et al. 2019; Khattab et al. 2021a; Karpas et al. 2022; Dohan et al. 2022; Khot et al. 2022; Khattab et al. 2022; Chen et al. 2022; Pourreza & Rafiei 2023; Shinn et al. 2023).

Unfortunately, LMs are known to be sensitive to how they are prompted for each task, and this is exacerbated in pipelines where multiple LM calls have to interact effectively. As a result, the LM calls in existing LM pipelines and in popular developer frameworks are generally implemented using hard-coded ‘prompt templates’, that is, long strings of instructions and demonstrations that are hand crafted through manual trial and error. We argue that this approach, while pervasive, can be brittle and unscalable—conceptually akin to hand-tuning the weights for a classifier. A given string prompt might not generalize to different pipelines or across different LMs, data domains, or even inputs.

Toward a more systematic approach to designing AI pipelines, we introduce the DSPy programming model. DSPy is pronounced dee-ess-pie. It’s the second iteration of our earlier Demonstrate–Search–Predict framework (DSP; Khattab et al. 2022). This paper introduces the key concepts in DSPy. For more extensive and up-to-date documentation of the framework, we refer readers to https://github.com/stanfordnlp/dspy. DSPy pushes building new LM pipelines away from manipulating free-form strings and closer to programming (composing modular operators to build text transformation graphs) where a compiler automatically generates optimized LM invocation strategies and prompts from a program. We draw inspiration from the consensus that emerged around neural network abstractions (Bergstra et al. 2013), where (1) many general-purpose layers can be modularly composed in any complex architecture and (2) the model weights can be trained using optimizers instead of being hand-tuned.

To this end, we propose the DSPy programming model (Sec 3). We first translate string-based prompting techniques, including complex and task-dependent ones like Chain of Thought (Wei et al. 2022) and ReAct (Yao et al. 2022), into declarative modules that carry natural-language typed signatures. DSPy modules are task-adaptive components—akin to neural network layers—that abstract any particular text transformation, like answering a question or summarizing a paper. We then parameterize each module so that it can learn its desired behavior by iteratively bootstrapping useful demonstrations within the pipeline. Inspired directly by PyTorch abstractions (Paszke et al. 2019), DSPy modules are used via expressive define-by-run computational graphs. Pipelines are expressed by (1) declaring the modules needed and (2) using these modules in any logical control flow (e.g., if statements, for loops, exceptions, etc.) to logically connect the modules.

We then develop the DSPy compiler (Sec 4), which optimizes any DSPy program to improve quality or cost. The compiler inputs are the program, a few training inputs with optional labels, and a validation metric. The compiler simulates versions of the program on the inputs and bootstraps example traces of each module for self-improvement, using them to construct effective few-shot prompts or finetuning small LMs for steps of the pipeline. Optimization in DSPy is highly modular: it is conducted by teleprompters, We derive the name tele-prompters from the notion of abstracting and automating the task of prompting, in particular, such that it happens at a distance, without manual intervention. which are general-purpose optimization strategies that determine how the modules should learn from data. In this way, the compiler automatically maps the declarative modules to high-quality compositions of prompting, finetuning, reasoning, and augmentation.

Programming models like DSPy could be assessed along many dimensions, but we focus on the role of expert-crafted prompts in shaping system performance. We are seeking to reduce or even remove their role through DSPy modules (e.g., versions of popular techniques like Chain of Thought) and teleprompters. We report on two expansive case studies: math word problems (GMS8K; Cobbe et al. 2021) and multi-hop question answering (HotPotQA; Yang et al. 2018) with explorations of chain of thought, multi-chain reflection, multi-hop retrieval, retrieval-augmented question answering, and agent loops. Our evaluations use a number of different compiling strategies effectively and show that straightforward DSPy programs outperform systems using hand-crafted prompts, while also allowing our programs to use much smaller and hence more efficient LMs effectively.

Overall, this work proposes the first programming model that translates prompting techniques into parameterized declarative modules and introduces an effective compiler with general optimization strategies (teleprompters) to optimize arbitrary pipelines of these modules. Our main contributions are empirical and algorithmic: with DSPy, we have found that we can implement very short programs that can bootstrap self-improving multi-stage NLP systems using LMs as small as llama2-13b-chat and T5-Large (770M parameters). Without hand-crafted prompts and within minutes to tens of minutes of compiling, compositions of DSPy modules can raise the quality of simple programs from 33% to 82% (Sec 6) and from 32% to 46% (Sec 7) for GPT-3.5 and, similarly, from 9% to 47% (Sec 6) and from 22% to 41% (Sec 7) for llama2-13b-chat.

Related Work

This work is inspired by the role that Torch (Collobert et al. 2002), Theano (Bergstra et al. 2010; Bergstra et al. 2011; Al-Rfou et al. 2016), Chainer (Tokui et al. 2015), and others played in the development in deep learning by providing powerful abstractions. A similar transformation is emerging with higher-level pipelines of LMs, and we are seeking to offer a solid conceptual framework and programming abstractions for what we call foundation model programming. We draw on differentiable programming (Wang et al. 2018) but applied to LM calls rather than neural networks, and borrow syntactic elements from PyTorch (Paszke et al. 2019).

In-context learning (McCann et al. 2018; Radford et al. 2018; Brown et al. 2020) is a key mechanism for foundation model programming. A growing body of work has revealed that, especially with instruction tuning (Ouyang et al. 2022), we can elicit sophisticated behavior via prompting (Wei et al. 2022; Wang et al. 2022b; Press et al. 2022; Yao et al. 2022; Khot et al. 2022; Madaan et al. 2023). Similarly, forms of weak supervision that would normally require task-specific (Khattab et al. 2021a; Khattab et al. 2021b) or hand-built (Ratner et al. 2016; Hancock et al. 2018) heuristics are now done by LMs (Wang et al. 2022b; Zelikman et al. 2022; Zhang et al. 2022; Shao et al. 2023).

In-context learning methods now routinely invoke tools, leading to LM pipelines that use retrieval models (Chen et al. 2017; Lewis et al. 2020; Guu et al. 2020; Lazaridou et al. 2022; Izacard et al. 2022), multimodal foundation models, and more traditional tools like APIs (Nakano et al. 2021) and calculators. A number of toolkits have been developed to facilitate this, including LangChain (Chase 2022), Semantic Kernel (Microsoft 2023), LlamaIndex (Liu 2022), and many other retrieval and agent libraries. These toolkits provide pre-packaged chains and agents that connect LMs with numerous accessible tools. However, they suffer from the pervasive prompt engineering challenges we address in DSPy: they express task-specific behavior through hand-written prompt templates (for detailed discussion, see Appendix B).

Researchers are starting to apply discrete optimization and RL to find effective prompts, generally for a single logical LM call (Guo et al. 2023; Pryzant et al. 2023; Huang et al. 2022; Yang et al. 2023). DSPy seeks to generalize this space: it offers a rich framework for optimizing arbitrary pipelines from high-level declarative signatures, by bootstrapping high-quality multi-stage demonstrations with constraints. In this framework, DSPy teleprompters may apply optimization using model selection techniques like cross-validation or, in principle, with sophisticated techniques involving RL and LM feedback (Hu et al. 2023; Zhao et al. 2023a; Shinn et al. 2023) or learned or Bayesian hyperparameter optimization methods (Bergstra et al. 2013; Akiba et al. 2019).

The present paper seeks to motivate DSPy as a programming model and to report new empirical findings from applying the DSPy compiler. This is inspired by formative work by Bergstra et al. 2010; Bergstra et al. 2013, Paszke et al. 2019, and Wolf et al. 2020, who support their respective programming models with a mix of benchmark numbers and some qualitative measures. For the current paper, we focus on showing that DSPy and its compiler allow us to build outstanding LM systems without hand-crafted prompt strings, but instead from truly modular units, and that this opens up doors for systematically exploring a rich design space at a very high programmatic level of abstraction.

The DSPy Programming Model

We present DSPy, which treats LMs as abstract devices for text generation, We assume access to one or more LMs, which consume a prompt string and return text completions. This may be a promptable LM capable of in-context learning (e.g., GPT-3.5 or Llama2-7b) or a smaller finetuneable LM (e.g., T5-base). An LM may be selected as the default; operations will use it unless configured otherwise. and optimizes their usage in arbitrary computational graphs. DSPy programs are expressed in Python: each program takes the task input (e.g., a question to answer or a paper to summarize) and returns the output (e.g., an answer or a summary) after a series of steps. DSPy contributes three abstractions toward automatic optimization: signatures, modules, and teleprompters. Signatures abstract the input/output behavior of a module; modules replace existing hand-prompting techniques and can be composed in arbitrary pipelines; and teleprompters optimize all modules in the pipeline to maximize a metric.

Instead of free-form string prompts, DSPy programs use natural language signatures to assign work to the LM. A DSPy signature is natural-language typed declaration of a function: a short declarative spec that tells DSPy what a text transformation needs to do (e.g., “consume questions and return answers”), rather than how a specific LM should be prompted to implement that behavior. More formally, a DSPy signature is a tuple of input fields and output fields (and an optional instruction). A field consists of field name and optional metadata. String descriptions of the task and the fields are also optional and usually omitted. Fields can carry optional field prefix and description. By default, fields are assumed to hold free-form strings; we are actively exploring optional data type as a way to specify constraints on valid values (e.g., bool or int) and more gracefully handle formatting and parsing logic, though this feature is not core to DSPy at the time of writing. In typical usage, the roles of fields are inferred by DSPy as a function of field names. For instance, the DSPy compiler will use in-context learning to interpret question differently from answer and will iteratively refine its usage of these fields.

Signatures offer two benefits over prompts: they can be compiled into self-improving and pipeline-adaptive prompts or finetunes. This is primarily done by bootstrapping (Sec 4) useful demonstrating examples for each signature. Additionally, they handle structured formatting and parsing logic to reduce (or, ideally, avoid) brittle string manipulation in user programs.

In practice, DSPy signatures can be expressed with a shorthand notation like question -> answer, so that line 1 in the following is a complete DSPy program for a basic question-answering system (with line 2 illustrating usage and line 3 the response when GPT-3.5 is the LM):

In the shorthand notation, each field’s name indicates the semantic role that the input (or output) field plays in the transformation. DSPy will parse this notation and expand the field names into meaningful instructions for the LM, so that english_document -> french_translation would prompt for English to French translation. When needed, DSPy offers more advanced programming interfaces for expressing more explicit constraints on signatures (Appendix A).

2 Parameterized & templated Modules can abstract prompting techniques

Akin to type signatures in programming languages, DSPy signatures simply define an interface and provide type-like hints on the expected behavior. To use a signature, we must declare a module with that signature, like we instantiated a Predict module above. A module declaration like this returns a function having that signature.

The core module for working with signatures in DSPy is Predict (simplified pseudocode in Appendix D.1). Internally, Predict stores the supplied signature, an optional LM to use (initially None, but otherwise overrides the default LM for this module), and a list of demonstrations for prompting (initially empty). Like layers in PyTorch, the instantiated module behaves as a callable function: it takes in keyword arguments corresponding to the signature input fields (e.g., question), formats a prompt to implement the signature and includes the appropriate demonstrations, calls the LM, and parses the output fields. When Predict detects it’s being used in compile mode, it will also internally track input/output traces to assist the teleprompter at bootstrapping the demonstrations.

DSPy modules translate prompting techniques into modular functions that support any signature, contrasting with the standard approach of prompting LMs with task-specific details (e.g., hand-written few-shot examples). To this end, DSPy includes a number of more sophisticated modules like ChainOfThought, ProgramOfThought, MultiChainComparison, and ReAct. These modules generalize prompting techniques from the literature, respectively, by Wei et al. 2022, Chen et al. 2022, Yoran et al. 2023, and Yao et al. 2022 and, in doing so, generalize the ideas on zero-shot prompting and rationale self-generation from Kojima et al. 2022, Zelikman et al. 2022, Zhang et al. 2022, and Huang et al. 2022 to parameterized modules that can bootstrap arbitrary multi-stage pipelines. These can all be used interchangeably to implement a DSPy signature. For instance, simply changing Predict to ChainOfThought in the above program leads to a system that thinks step by step before committing to its output field.

Importantly, all of these modules are implemented in a few lines of code by expanding the user-defined signature and calling Predict one or more times on new signatures as appropriate. For instance, we show a simplified implementation of the built-in ChainOfThought below.

This is a fully-fledged module capable of learning effective few-shot prompting for any LM or task. We contrast that with Appendix C, which copies long reasoning prompts hand-written by sources ranging from recent research to popular prompting libraries.

Uniquely, DSPy parameterizes these prompting techniques. To understand this parameterization, observe that any LM call seeking to implement a particular signature needs to specify parameters that include: (1) the specific LM to call (Chen et al. 2023), (2) the prompt instructions (Yang et al. 2023) and the string prefix of each signature field and, most importantly, (3) the demonstrations used as few-shot prompts (for frozen LMs) or as training data (for finetuning). We focus primarily on automatically generating and selecting useful demonstrations. In our case studies, we find that bootstrapping good demonstrations gives us a powerful way to teach sophisticated pipelines of LMs new behaviors systematically.

DSPy programs may use tools, which are modules that execute computation. We support retrieval models through a dspy.Retrieve module. At the time of writing, DSPy has built-in support for ColBERTv2, Pyserini, and Pinecone retrievers, and we have explored experimental dspy.SQL for executing SQL queries and dspy.PythonInterpreter for executing Python code in a sandbox.

DSPy modules can be composed in arbitrary pipelines in a define-by-run interface. Inspired directly by PyTorch and Chainer, one first declares the modules needed at initialization, allowing DSPy to keep track of them for optimization, and then one expresses the pipeline with arbitrary code that calls the modules in a forward method. As a simple illustration, we offer the following simple but complete retrieval-augmented generation (RAG) system.

To highlight modularity, we use ChainOfThought as a drop-in replacement of the basic Predict. One can now simply write RAG()("Where is Guaraní spoken?") to use it. Notice that, if we use a signature "context, question -> search_query", we get a system that generates search queries rather than answers.

3 Teleprompters can automate prompting for arbitrary pipelines

When compiling a DSPy program, we generally invoke a teleprompter, which is an optimizer that takes the program, a training set, and a metric—and returns a new optimized program. Different teleprompters (Sec 4) apply different strategies for optimization.

In DSPy, training sets may be small, potentially a handful of examples, though larger data enables more powerful optimization. Training examples may be incomplete, i.e., only input values are necessary. Labels for the pipeline steps are not required, unless they need to be used in the metric. In practice, we typically assume labels only for (at most) the program’s final output, not the intermediate steps. This label-efficiency is critical for modularity: building a new pipeline in DSPy requires simply recompiling the new pipeline’s code, not annotating data specific to the new pipeline.

Metrics can be simple notions like exact match (EM) or F1, but they can be entire DSPy programs that balance multiple concerns. For example, we may compile the RAG module above against a dataset of question–answer pairs qa_trainset and the metric EM. The goal of optimization here is to effectively bootstrap few-shot demonstrations. The following code achieves this:

In this example, the BootstrapFewShot teleprompter (Sec 4, Appendix E.1) simulates RAG on the training example(s). It will collect demonstrations of each module (i.e., examples of its input–output behavior) that collectively lead to valid output (i.e., respecting the signatures and the metric).

If one wanted to push the compiled program to be extractive given its retrieved contexts, one could define a custom metric to use in place of dspy.evaluate.answer_exact_match:

Notice that behavior like this might be more accurately checked by another DSPy program that checks for faithful grounding of answers. Such metrics are fully supported and encouraged in DSPy.

Teleprompters can be composed by specifying a teacher program. DSPy will sample demonstrations from this program for prompt optimization. This composition can enable very rich pipelines, where expensive programs (e.g., complex expensive ensembles using large LMs) supervise cheap programs (e.g., simple pipelines using smaller LMs). One may start with compiled_rag from above (say, compiled to use a large Llama2-13b-chat LM) but now fine-tune Flan-T5-large to create an efficient program:

The DSPy Compiler

A key source of DSPy’s expressive power is its ability to compile—or automatically optimize—any program in this programming model. Compiling relies on a teleprompter, which is an optimizer for DSPy programs that improves the quality (or cost) of modules via prompting or finetuning, which are unified in DSPy. While DSPy does not enforce this when creating new teleprompters, typical teleprompters go through three stages.

The compiler first (recursively) finds all unique Predict modules (predictors) in a program, including those nested under other modules. For each unique predictor pp, the teleprompter may generate candidate values for the parameters of pp: the instructions, field descriptions, or—most importantly—demonstrations (i.e., example input–output pairs). In this iteration of DSPy, we focus on demonstrations and find that simple rejection-sampling-like approaches can help bootstrap highly effective multi-stage systems.

Consider the simplest non-trivial teleprompter in DSPy, BootstrapFewShot (simplified pseudocode in Appendix E.1). This teleprompter will simulate a teacher program (or, if unset, the zero-shot version of the program being compiled) on some training inputs, possibly one or more times with a high temperature. When running in compile mode, multi-stage traces are tracked transparently and in a thread-safe fashion throughout execution. The program’s metric is used to filter for multi-stage traces that together help the pipeline pass the metric. We thus obtain potential labels for all signatures in the program by throwing away the bad examples and using the good examples as potential demonstrations, though these design decisions are under user control.

While LMs can be highly unreliable, we find they can be rather efficient at searching the space of solutions for multi-stage designs. A well-decomposed program can typically find at least a few training examples where the LM can pass the constraints enforced by the signatures and metrics, allowing us to bootstrap iteratively if needed.

Now each parameter has a discrete set of candidates: demonstrations, instructions, etc. Many hyperparameter tuning algorithms (e.g., random search or Tree-structured Parzen Estimators as in HyperOpt (Bergstra et al. 2013) and Optuna (Akiba et al. 2019)) can be applied for selection among candidates. We report simplified implementations of DSPy’s BootstrapFewShotWithRandomSearch and BootstrapFewShotWithOptuna in Appendix E.2 and Appendix E.3.

Another type of optimization is finetuning with BootstrapFinetune, where the demonstrations are used to update the LM’s weights for each predictor. When this is applied, the LM parameter of each module is updated to the new LM weights. Typically, we are optimizing average quality using the metric with cross-validation over the training set or a validation set. This is applicable even with no labels for any stages, depending on the nature of metric.

A different type of optimization that the DSPy compiler supports is modifying the control flow of the program. One of the simplest forms of these is ensembles, which we use in the case studies in this work. An ensemble will bootstrap multiple copies of the same program, and then replace the program with a new one that runs them all in parallel and reduces their predictions into one with a custom function (e.g., majority voting). In future work, this stage can easily accommodate techniques for more dynamic (i.e., test-time) bootstrapping as well as automatic backtracking-like logic.

Goals of Evaluation

Programming frameworks can be evaluated along many dimensions: computational efficiency, developer efficiency, intuitiveness of the code and concepts, and so forth. In this paper, we focus on perhaps the most pressing issue for current LM pipelines: the role of hand-written, task-specific prompts in achieving performant systems. Our evaluations seek to test the following hypotheses:

With DSPy, we can replace hand-crafted prompt strings with concise and well-defined modules, without reducing quality or expressive power.

Parameterizing the modules and treating prompting as an optimization problem makes DSPy better at adapting to different LMs, and it may outperform expert-written prompts.

The resulting modularity makes it possible to more thoroughly explore complex pipelines that have useful performance characteristics or that fit nuanced metrics.

Our evaluation will explore these hypotheses using diverse task–program pairs. We hope this begins a shift from underspecified questions like “how do different LMs compare on GSM8K” toward “how they compare on GSM8K with program P when compiled with strategy S”, which is a well-defined and reproducible run. Ultimately, our goal is to reduce the role of artful prompt construction in modern AI in favor of the development of new modular, composable programs and optimizers.

Case Study: Math Word Problems

We evaluate on the popular GSM8K dataset with grade school math questions (Cobbe et al. 2021). We sample 200 and 300 question–answer pairs from the official training set for training and development, respectively. Our final evaluations use the 1.3k official test set examples. We report extensive comparisons on the development set to avoid overfitting on test. Following prior work on GSM8K, we evaluate the accuracy of the final numerical value that appears in the LM output.

For this task, we consider three simple DSPy programs: a one-step Predict module (vanilla), a two-step ChainOfThought module (CoT), and finally a multi-stage ComparerOfThoughts module (ThoughtReflection). These are fully defined by the following code:

In reflection, five reasoning chains are sampled from the LM (alongside their answers) and they are compared in parallel by a built-in MultiChainComparison module, which generalizes Yoran et al. 2023. This generates a new answer taking into account the patterns from the five attempts. Critically, the modules used are all generic, none is specific math problems or particular LM.

As we discussed in Section 4, DSPy programs can be compiled into new, optimized programs. In our experiments, we evaluate the programs zero-shot (no compiling) as well as a number of strategies for compiling. Our simplest compiler is LabeledFewShot:

Here, program can be any DSPy module. This simply samples k=8 random demonstrations from the trainset for the fields common to the training examples and the signature(s), in this case, question and answer, but not the reasoning for instance. We report the average of 3–5 runs (depending on the setting) when applying such random sampling.

Next, we also consider bootstrapping few-shot examples with random search:

This will generate demonstration chains for examples in the training set and optimize the selection of demonstrations (from this set) to self-improve the program’s modules. As the name indicates, this is done with random search, treating the selection of demonstrations as a parameter to optimize.

Next, if desired, this bootstrapping process can be nested in DSPy. In particular, we can use the optimized bootstrap program itself to further bootstrap another program. This is relevant, for example, whenever the original zero-shot program performs relatively poorly.

And lastly, we consider ensembling these bootstraps:

GSM8K includes human reasoning chains. Above, trainset does not include these reasoning chains. We also evaluate with trainset_human_CoT, which extends the examples in trainset with the human reasoning string. These two datasets can be used interchangeably as the value for the trainset parameter above. We note here that compiling generally runs on the order of minutes (or tens of minutes) as even the more expensive settings only require running the program a few thousand times (e.g., 10–20 trials over 150–300 validation examples) and they can occur in parallel.

Our results are summarized in Table 1, which includes dev results as well as our evaluation of promising representatives of each approach on the test set. First, the vanilla program results show that GPT-3.5 and llama2-13b-chat struggle with math word problems when they have to predict the answers directly, that is, without using a reasoning chain first. This is most pronounced in the absence of good demonstrations, which can be seen in the none compilation setting (i.e., zero-shot instruction) and the fewshot setting (i.e., sampling random question–answer pairs). Interestingly, however, vanilla is helped substantially by compiling with bootstrap and by iterating this process into bootstrap×2\times 2. On inspecting the prompts bootstrapped (Appendix F), we see that the prompt allows the LM to leverage the answer field for reasoning first, which is permitted as the metric extracts the final numerical value for evaluation.

Next, we consider the CoT program. While the expert human reasoning chains (+human_CoT) provide a large boost when available, we can match or surpass this using bootstrap, substantiating our hypothesis that DSPy can cut the need for hand-crafted prompts. Beyond this, we see that the reflection program, while only a few lines longer than the others, is a clear winner, though CoT is quite effective with ensemble. Overall, the bootstrap compilation procedure leads to large gains for every program, across both LMs. Indeed, all programs in this table are expressed by composing two to four DSPy modules and teleprompters, and they reveal overall that—in the new paradigm prescribed by DSPy—it’s composing the right generic modules, rather than manipulating string prompts, that improves different LMs from 4–20% accuracy to 49–88% accuracy.

We can informally compare with the following. Zhang et al. 2022 reports 48% for text-davinci-002, which aligns closely with our llama2-13b-chat results, and reports 59.4% with codex when employing a manual CoT approach and 62.8% with an automatic CoT method. Wang et al. 2022b report 57% for CoT prompting with PaLM 540-B, which becomes 74% upon adding self-consistency. The Llama2 authors (Touvron et al. 2023) presents 28.7% for llama2-13b, 42.2% for llama2-34b, and 56.8% for llama2-70b. Intriguingly, our program with the 13b variant of the model is competitive with their 34b-based results even though we don’t use human reasoning chains in our program. Zhao et al. 2023b reports 80.8% for CoT with gpt-3.5-turbo from April 2023. The GPT-4 authors (OpenAI 2023) reports that GPT-3.5 scores 57.1% and GPT-4 elevates this to 92% but they note that GPT-4 was in fact pre-trained on a subset of GSM8K’s training set.

Case Study: Complex Question Answering

In this case study, we explore the multi-hop question answering task with the HotPotQA (Yang et al. 2018) dataset in the open-domain “fullwiki” setting. For retrieval, we use a search index of the official Wikipedia 2017 “abstracts” dump of HotPotQA. Search is conducted by a ColBERTv2 (Santhanam et al. 2021) retriever. The HotPotQA test set is hidden, so we reserve the official validation set for our testing, and sample 1000 examples for that. We sub-divide the training set into 70%/30% train/validation splits. In the training (and thus validation) split, we keep only examples marked as “hard” in the original dataset, which matches the designation of the official validation and test sets. For training and for reporting development results, we sample 200 and 300 examples respectively.

Our simplest baseline is the vanilla program used in the previous case study on GSM8K (Sec 6); the "question -> answer" signature is universal enough that it will work for this task (and many others) when compiled appropriately.

Our baseline RAG program is the one given in Section 3.2 as a simple example of RAG with a dspy.ChainOfThought layer. We will see that this program does not excel at HotPotQA, and this motivates us to evaluate two multi-hop programs.

To that end, we first test ReAct (Yao et al. 2022), a multi-step agent for tool use, which is implemented as a built-in module in DSPy. In the simplest case, a ReAct module for a particular signature can be declared as follows in DSPy:

We also test the following custom program, which simulates the information flow in Baleen (Khattab et al. 2021a) and IRRR (Qi et al. 2020) and has similarities to IRCoT (Trivedi et al. 2022).

For compilers, we continue to use the ones that we used for GSM8K (see Sec 6). We also consider two compositions of our teleprompters. For ReAct, we consider bootstrapping with BootstrapFewShotWithRandomSearch starting from an earlier bootstrap of the ReAct program. For the simple multihop program, we also consider fine-tuning with T5-Large starting from the earlier bootstrap of that program.

Table 2 summarizes our results. Compared with the vanilla few-shot prompting, a chain-of-thought and retrieval-augmented generation (CoT_RAG) program can self-bootstrap in DSPy to increase answer EM substantially. However, this relies entirely on the ColBERTv2 retriever to find relevant passages directly from the original questions, limiting its passage recall. This is tackled in the react and multihop programs, which will generate queries for the retriever in multiple iterative “hops”. Indeed, overall, a simple multihop program performs the best, and in general bootstrap again proves to be very effective at raising its quality relative to its fewshot variant for both LMs.

In particular, we can see that bootstrap (and/or bootstrap×2\times 2) can outperform both fewshot prompting (for multihop) and expert human reasoning (for react; adapted slightly from Yao et al. 2022 to our retrieval setting). Perhaps most importantly, we can make llama2-13b-chat competitive with GPT-3.5 by simply compiling our programs.

To assess the finetuning capacity of DSPy, we also evaluated the compiler multihop_t5 defined above which produces a T5-Large (770M parameter) model. This program scores 39.3% answer EM and 46.0% passage accuracy on the dev set, using only 200 labeled inputs and 800 unlabeled questions. For compiling, we use a teacher program consisting of an ensemble (union) of two multihop with llama2-13b-chat. Considering its extremely small size and local availability, this compiled program with T5-Large would impose orders of magnitude lower costs for inference than a proprietary LM like GPT-3.5.

Our results may be pegged against the evaluation on HotPotQA in a number of recent papers, though there is significant variation in evaluation methodology and test set samples across studies in this space. Using CoT prompting, Si et al. 2022 achieve 25.2% EM. With a “recite-and-answer” technique that uses PaLM-62B (Chowdhery et al. 2022) to recite evidence passages, Sun et al. 2022 achieve 26.5% EM. Wang et al. 2022a achieve 33.8% EM and 44.6% F1 when applying self-consistency for PaLM-540B. Yao et al. 2022 achieve 27.4% EM using ReAct with PaLM-540B and 30.8 with text-davinci-002, with a tool giving it the ability for search using a Wikipedia API. They push their PaLM results to 35.1% EM by applying an additional CoT step with self-consistency, which may resemble our ensemble approach in the sense of aggregating multiple answers. Trivedi et al. 2022 reports 49% using a pipeline with code-davinci-002 LM on a sample of 500 HotPotQA questions.

Conclusion

This paper introduced DSPy, a new programming model for designing AI systems using pipelines of pretrained LMs and other tools. We presented three new concepts introduced in this abstraction (DSPy signatures, modules, and teleprompters), and showed in two very different case studies that it supports rapid development of highly effective systems that use relatively small LMs. We have maintained open-source versions of this framework for close to a year. In this period, we have seen and created a large number of programs that were compiled to high-quality systems by DSPy, spanning tasks from information extraction to low-resource synthetic data generation. In the interest of space and to maintain reasonable scope in this paper, we leave reporting on such tasks under controlled experimental conditions to future work. While in-context learning has proved transformative over the past 2–3 years of LM research, we argue that the true expressive power in this emerging paradigm is in building sophisticated text transformation graphs in which composable modules and optimizers (teleprompters) come together to leverage LMs in more systematic and reliable ways.

We thank Josh Purtell for suggesting the apt name “text transformation graph” for the computational graph abstraction of DSPy. We thank Rick Battle, Igor Kotenkov, Lisa Li, David Hall, Ashwin Paranjape, Chris Manning, Percy Liang, and many researchers, developers, and users for valuable discussions and feedback. We thank Giuseppe Attanasio for his public LaTeX GitHub-style Python code formatting gist. https://gist.github.com/g8a9/07c2be12ae02cfad4aa430d77dc940cb

This work was partially supported by IBM as a founding member of the Stanford Institute for Human-Centered Artificial Intelligence (HAI), Oracle, Virtusa, and Cigna Healthcare. It was also partially supported by an HAI Azure compute grant. This research was supported in part by affiliate members and other supporters of the Stanford DAWN project–Facebook, Google, and VMware—as well as the NSF under CAREER grant CNS-1651570. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the authors and do not necessarily reflect the views of the National Science Foundation. Omar Khattab is supported by the Apple Scholars in AI/ML fellowship.

References

Appendix A Advanced Signatures

When more control is desired, one can express signatures as Python classes to provide explicit instructions of the transformation and describe the format or role of each field more directly. For instance, the following signature generates search queries using context and an optional question:

Using the above, we can specify a complete system for the generation of a synthetic IR dataset where the queries are mediated by a question generated by the LM:

If questions are available, they can be supplied as shown: query_gen(context="Language typology", question="What are the primary language families of South America?").

As a work in progress feature, users can optionally specify the type of output fields as bool, int, float, list, or dict instead of the default free-form string type, as in contexts, question -> answer_found: bool.

Appendix B Comparison with existing libraries like LangChain and LlamaIndex

LangChain and LlamaIndex are perhaps the most popular library in the general space of prompting LMs. These libraries have a different focus compared to DSPy and they suffer internally from the prompt engineering challenges that DSPy aims to resolve. In particular, whereas the goal of DSPy is to tackle the fundamental challenges of prompt engineering for building new LM computational graphs, LangChain and LlamaIndex generally help application developers who need pre-packaged components and chains, e.g., implementations of popular and reusable pipelines (e.g., particular agents and specific retrieval pipelines) and tools (e.g., connections to various databases and implementations of long- and short-term memory for agents).

These off-the-shelf higher-level abstractions contrast with DSPy’s focus on introducing core composable operators. In particular, DSPy introduces signatures (to abstract prompts), modules (to abstract prompting techniques), and teleprompters to act as optimizers for arbitrary imperative code (DSPy programs) that chain modules together. Its goal is to help researchers and practitioners build new LM pipelines quickly and achieve very high quality through automatic compilation (self-improvement) instead of manual prompt engineering.

In contrast, typical existing research implementations and existing libraries like LangChain and LlamaIndex are implemented using manual prompt engineering, which is the key problem that DSPy tackles. We conducted an informal study to highlight this. In late September 2023, we found that the LangChain codebase contains 50 strings exceeding 1000 characters, which are generally prompts, compared to none at all in DSPy. Indeed, a substantial number of LangChain’s Python files are singularly dedicated to task-related templating and prompt engineering with 12 prompts.py files and and 42 prompt.py files. DSPy, on the other hand, provides a structured framework that automatically bootstraps prompts. The library itself does not contain a single hand-written prompt demonstration for any tasks at the time of writing, despite the very high quality with various LMs.

To review the typical forms of prompt engineering in existing libraries, we consider the following in LangChain. The LangChain Program-Aided Language Model Gao et al. 2023a chain program uses few-shot learning, leveraging a template that is 3982 characters long with 8 math word problems (Prompt 2) and corresponding outputted programs as learning examples for the language model. LangChain also contains a prompt for SQL query tasks for each of the databases like Oracle, GoogleSQL, DuckDB, Crate, and MySQL, with the average length of these prompts at 1058 characters. Other task areas such as QA with sources (Prompt B) and Graph_QA also have significantly lengthy prompt templates, with averages of 1337 and 722 characters, respectively. While expert-written prompts can be useful, we believe that LM- and task-adaptive prompts bootstrapped automatically can offer far more power (and are far more modular) than hard-coding a prompt per database provider inside the code base. The next appendix section contains a number of prompts copied from related research papers and existing libraries.

This section highlights a few popular existing frameworks that structure prompts with extensive prompt engineering templates. The primary objective is to capture how many words and characters are used for such large multi-line prompts defined for tasks or tools and present these example prompts retrieved from open-sourced papers and repositories. The formatting of these example prompts is adapted from Gao et al. 2023a.

Appendix D Modules

D.2 Chain of Thought

Appendix E Teleprompters

E.2 BootstrapFewShotWithRandomSearch

E.3 BootstrapFewShotWithOptuna

Appendix F Examples of the prompts automatically generated by DSPy

For GSM8K, we include the prompt bootstrapped by DSPy for GSM8K llama2-13b-chat for the vanilla program compiled with bootstrap×2\times 2 in Figure 9.

We also include a CoT prompt for GSM8K and a generate_query prompt from the multihop program for HotPotQA. All of these, particularly their demonstrations’ labels and their selection, are generated by DSPy automatically using llama2-13b-chat.