Meta-Prompting: Enhancing Language Models with Task-Agnostic Scaffolding
Mirac Suzgun, Adam Tauman Kalai
Introduction
The latest generation of language models (LMs)—notably, GPT-4 (OpenAI, 2023), PaLM (Anil et al., 2023), and LLaMa (Touvron et al., 2023)—have expanded the boundaries of natural-language processing and generation. These large-scale models can tackle a wide spectrum of tasks, ranging from writing Shakespearean sonnets about hedgehogs to summarizing intricate medical reports and solving competition-level programming puzzles. Despite their versatility, these models are not infallible; they sometimes generate responses that are inaccurate, misleading, or conflicting. As the operational costs of these models become more affordable, it becomes natural to ask whether one might use scaffolding systems and leverage multiple LM queries to not only refine but also to enhance the accuracy and robustness of these model outputs.
In this work, we introduce a new technique for enhancing the functionality and performance of LMs, called meta-prompting. It involves constructing a high-level “meta” prompt that instructs an LM to: (i) break down complex tasks or problems into smaller, manageable pieces; (ii) assign these pieces to specialized “expert” models with proper and detailed natural-language instructions; (iii) oversee the communication between these expert models; and (iv) apply its own critical thinking, reasoning, and verification skills throughout the process. When presented with a query, the LM, effectively prompted under meta-prompting, serves as a conductor. It produces a message history—a narrative, if you will—comprising the responses from various expert models. The LM is originally responsible for generating the conductor’s portion of this history, which includes the selection of experts and the formulation of specific instructions for them. However, the same LM doubles itself as these independent experts as well, generating outputs based on the expertise and information chosen by the conductor for each particular query.
This approach allows for a single, uniform LM to maintain a coherent line of reasoning while also tapping into a variety of expert roles. The use of dynamically selected contexts for prompting these experts introduces fresh perspectives into the process, while the conductor model retains a bird’s-eye view of the entire history and coordination. This method, therefore, enables a single black-box LM to function effectively as both a central conductor and a diverse panel of experts to produce more accurate, reliable, and coherent responses.
Our proposed meta-prompting technique combines and expands upon various prompting ideas introduced by recent studies—including, high-level planning and decision-making (Yao et al., 2023b; Sun et al., 2023; Hao et al., 2023a), dynamic persona assignment (Xu et al., 2023; Wang et al., 2023), multi-agent debating (Du et al., 2023; Zhuge et al., 2023), self-debugging and self-reflection (Schick et al., 2023b; Liu et al., 2023a; Gou et al., 2023; Madaan et al., 2023; Shinn et al., 2023). A key aspect of meta-prompting is its task-agnostic nature. Unlike traditional scaffolding methods that require specific instructions or examples tailored to each task, meta-prompting employs the same set of high-level instructions across various tasks and inputs. This universality is particularly beneficial for users who might find it cumbersome to provide detailed examples or specific guidance for every distinct task. For instance, in responding to a one-off request like “Write a Shakespearean sonnet about selfies,” the user would not need to supply examples of high-quality neoclassical poems. The meta-prompting approach elevates the utility of language models by offering a broad, flexible framework without compromising on specificity or relevance. Additionally, to demonstrate the versatility and integration capabilities of meta-prompting, we have enhanced our system with the functionality to invoke a Python interpreter. This allows for an even more dynamic and comprehensive application of the technique, further extending its potential to address a wide array of tasks and queries effectively.
We provide an illustrative visualization of a meta-prompting session in Figure 2. It depicts how the Meta Model—our technical term for the central controlling LM (a.k.a. the conductor)—intersperses its own output with inputs and outputs from various specialized expert models or code executions. Such a configuration makes meta-prompting a nearly universal tool. It allows for the consolidation of various LM interactions and computations into a single, coherent narrative. What sets meta-prompting apart is that it leaves the decision of which prompts to use and which code snippets to execute to the discretion of the LM itself.
In our comprehensive experiments, which primarily utilize GPT-4 as the foundational LM, we compare the efficacy of meta-prompting against other task-agnostic scaffolding methods. Our findings reveal that meta-prompting not only enhances overall performance but often leads to state-of-the-art results across a diverse range of tasks. Its flexibility is noteworthy: The conductor model has the capability to call upon expert models (basically itself, albeit with fresh instructions) for performing a variety of functions. These functions might include critiquing earlier outputs, selecting specific personas for certain tasks, refining generated content, and ensuring that the final outputs meet the desired criteria in both substance and form. This approach shows a marked improvement over several existing methods, as demonstrated in Figure 1.
The core contribution of this work is the introduction of a task-agnostic scaffolding system that leverages a single LM. This LM not only carries forward the thread of the task but also dynamically selects and instructs expert models appropriate for each specific task. The effectiveness of this system is showcased across various benchmarks, including the Game of 24 (Yao et al., 2023a), Checkmate-in-One from the BIG-Bench suite (BIG-Bench authors, 2023), and our novel task of “Shakespearean Sonnet Writing.” Overall, our empirical results underscore the versatility and robustness of meta-prompting in enhancing LM performance.
Meta Prompting
Intuition and Abstract Overview. The modus operandi of meta-prompting is to use a modelOur use of the term model refers to the application of an LM with certain prompt templates to play a specified “role.” We typically only use a single LM (e.g., GPT-4) to implement all the models in an execution. to coordinate and execute multiple independent inquiries and subsequently synthesize their responses to render a final response. This mechanism, in principle, endorses an ensemble approach, drawing from the strength and diversity of independent specialized models to collaboratively address and tackle multifaceted tasks or problems. We posit that while a single, general-purpose model might deliver valuable and useful insights into generic queries, combining the perspectives and conclusions of multiple domain-specific models (which we also refer to as experts) has the potential to yield more comprehensive, robust, and accurate solutions.
Central to our meta-prompting strategy is its shallow hierarchical configuration, where a single model—called the “Meta Model”—emerges as the principal entity of authority. This prompting structure is reminiscent of an orchestra, wherein the conductor’s role is mirrored by the Meta Model and each musician corresponds to a distinct domain-specific model. Just as a conductor harmonizes multiple musical elements to craft a beautiful melody, the Meta Model combines solutions and insights from a range of models to provide an accurate and comprehensive answer to an intricate problem or task.
Conceptually, a domain-specific expert within our framework can take diverse forms, such as a finetuned LM tailored to perform a particular task, a specialized API equipped to handle specific domain-related inquiries, or even computational tools like calculators or a Python interpreter that can perform arithmetic calculations or write and execute code. These experts, despite their varying functionalities, are directed and unified under the supervision of the Meta Model.
Under our setup, experts can be called only by the Meta Model. They cannot directly interact or communicate with each other, though the Meta Model can choose to share some text from or combine the insights of various experts when interacting with a new expert. This restriction is made to simplify the communication between the experts and to put the Meta Model at the center of the operation.
Algorithmic Procedure. Algorithm 1 provides pseudocode of our proposed meta-prompting approach. We further provide a conceptual overview of the procedure below:
Transforming the Input: Using the transformation function , the raw query is placed in a suitable template followed by initial instructions to the Meta Model.
Prompting the Meta Model: The current message list, namely , guides the Meta Model’s next action—either directly addressing the query or consulting a domain-specific expert.
Engaging Domain-Specific Expert Models: If the Meta Model does not return a result, it can conjure any expert and give it instructions, which are extracted from its output using . This process is isolated though: Each expert only sees what the Meta Model chooses to share with them, and responds accordingly. For instance, if a problem pertains to mathematics and history, the Meta Model might consult a mathematics expert for a calculation and a history expert for historical context. The output of the expert is extracted and additional instructions are appended, all using the template.
Returning the Final Response: If the Meta Model’s response contains a final answer (highlighted by distinct special markers), the solution is extracted using and returned.
Error Handling: In cases where the model response contains neither a final answer nor a call to an expert model, an error message appended to the message list . This ensures that our procedure is robust and can handle unexpected outputs.
Meta and Expert Model Specifications. In our setup, we employ the same LM, such as GPT-4, to function in both Meta and Expert capacities. Their roles are distinguished by their respective model instructions in their prompts, with the Meta Model adhering to a set of instructions provided in Figure 3, and the expert models following separate instructions dynamically determined by the Meta Model at inference time .
Experimental Setup
We compare meta-prompting with the task-agnostic, zero-shot versions of the following prompting methods:
Standard prompting: This represents our most basic baseline wherein an LM is asked to directly yield a response without any specific guiding input-output exemplars or any additional guiding instructions, besides the task description already included in the input query.
Zero-shot CoT prompting (Kojima et al., 2022): Drawing inspirations from the chain-of-thought method of Wei et al. (2022b), this zero-shot prompting approach simply appends “Let’s think step by step” to the input query, encouraging the model to have a more deliberative and iterative cognition before addressing the problem or task at hand.
Expert prompting (Xu et al., 2023): This prompting approach functions through a two-step process: It first crafts an expert identity tailored to align with the specific context of the input query. It then integrates this generated expert profile into the input to generate a well-informed and authoritative response. In our experiments, we consider two versions of expert prompting, namely (a) static (i.e., generic) and (b) dynamic (i.e., adaptive); the former uses a fixed and generic expert description, whereas the latter adaptively designs a new expert identity for each input query.
Multi-persona prompting (Du et al., 2023): Also known as solo-performance prompting (SPP), this method instructs an LM to perform the following: (i) Propose a small ensemble of “personas” to address the specific task or problem at hand; (ii) let these personas engage in a collective dialogue, collaboratively generating potential solutions while extending feedback to one another and refining their answers; and (iii) synthesize all the available information and deliver a final response.
2 Datasets and Tasks
To evaluate the efficacy of our proposed meta-prompting approach over other zero-shot prompting baselines, we consider a wide range of tasks and datasets that require various degrees of mathematical and algorithmic reasoning, domain-specific knowledge, and literary creativity. These include:
(a) The Game of 24 from (Yao et al., 2023a) where the goal is to form an arithmetic expression whose value is 24 using each of four given numbers exactly once,
Three BIG-Bench Hard (BBH; Suzgun et al. (2023b)) tasks—namely, (b) Geometric Shapes, (c) Multi-Step Arithmetic Two, and (d) Word Sorting—as well as one reasoning task directly obtained from the BIG-Bench suite (BIG-Bench authors, 2023), that is, (e) Checkmate-in-One;
(f) Python Programming Puzzles (P3; Schuster et al. (2021)), a collection of challenging programming puzzles written in Python—with varying difficulty levels;
(g) Multilingual Grade School Math (MGSM; Shi et al. (2023)), a multilingual version of the GSM8K dataset (Cobbe et al., 2021) with translations of a subset of examples into ten typologically diverse languages, including Bengali, Japanese, and Swahili;
(h) Shakespearean Sonnet Writing, a novel task we created where the goal is to write a sonnet with strict rhyme scheme “ABAB CDCD EFEF GG,” containing the three provided words verbatim.While all the other tasks and datasets were previously introduced by other studies, we present this task for the first time.
3 Answer Extraction and Evaluation Protocols
As shown in Figure 3, the system instruction in our proposed meta-prompting method encourages the Meta Model to present its final answer in a specific format. This format, designed for consistent and unambiguous extraction, requires that the final answer is wrapped within triple quotes and preceded by a distinct marker (namely, “>>FINAL ANSWER:”).
Once the final answer is extracted from the model and properly post-processed, we also need to evaluate its correctness.We have developed suitable pipelines for answer extraction and processing tailored to each task. Specific implementation details can be found in our codebase. Because we consider a wide range of tasks, there is not a single metric that allows us to measure accuracy across all. Depending on the nature and formulation of the task, we measure accuracy using one of the following three metrics:
Exact Match (EM): Under this strict metric, the correctness of an answer is determined by its precise alignment with the ground-truth label(s). An answer is deemed correct only if it is identical to a provided reference.
Soft Match (SM): This metric offers a more lenient approach than EM. For an answer to be deemed correct, it is sufficient for a ground-truth label to be present within the model’s output, regardless of any additional textual content.
Functionally Correct (FC): This metric ascertains whether the answer is functionally correct, meaning that it adheres to task-specific constraints.
We use EM for Geometric Shapes, Multi-Step Arithmetic Two, and Checkmate-in-One; SM for MGSM and Word Sorting,; and FC for Game of 24, Python Programming Puzzles, and Shakespearean Sonnet Writing.
4 Models and Inference
In our main experiments, we concentrate on GPT-4 (gpt-4-32k), which is accessible through Microsoft’s Azure OpenAI Service. Additionally, in our supplementary experiments, we include GPT-3.5 (gpt-35-turbo). Both GPT-3.5 and GPT-4 are models fine-tuned for following instructions, though GPT-4 has demonstrated significantly better reasoning and content generation abilities than GPT-3.5.In our preliminary experiments, we also tested other OpenAI models such as text-davinci-003 and code-davinci-002, but we discovered that our meta-prompting approach yielded consequential results when applied to GPT-3.5 and GPT-4.
In all of our experiments, we consistently applied the same parameters and system instructions to the Meta Model. We set the temperature value at , the top-p value at , and the maximum token count at .The temperature value, which usually ranges between 0 and 1, controls how much randomness or creativity the model exhibits. Ideally, a temperature of 0 should lead to the model producing the same output when presented with the same input. However, both GPT-3.5 and GPT-4 have shown a tendency to generate varied responses even at this setting. This means that reproducing our exact results might be challenging under identical experimental conditions. To address this issue, we are releasing all model inputs, interactions, and outputs in our GitHub repository.
Main Results and Discussion
Analysis of Expert Types Used in Meta Prompting. The Meta Model’s dynamic selection of expert types distinctly illustrates its adaptability and strategic alignment with specific task requirements. Analyzing tasks with and without a Python interpreter offers insightful contrasts in the model’s expert choices, influenced by the available tools and task characteristics. In scenarios where a Python expert is explicitly mentioned for code generation and execution, there is a noticeable preference for technical and computational expertise. For example, in Python Programming Puzzles, the Meta Model frequently utilizes Expert Python, Expert Mathematician, and several tiers of Expert Python Programmers. This pattern reveals a task-oriented strategy, highlighting a focus on programming and algorithmic problem-solving. Similarly, tasks such as Game of 24 and Word Sorting prominently feature Expert Python, reinforcing the model’s propensity to rely on computational expertise when Python capabilities are accessible.
In contrast, for meta-prompting without a specific Python expert, the spectrum of experts employed is more diverse. Tasks like Geometric Shapes predominantly involve design and geometry experts (e.g., Expert Graphic Designer and Expert Geometer), indicating a pivot towards visual and spatial problem-solving rather than computational approaches. This task illustrates where the Meta Model may have made a poor choice of experts, and in particular it might have been more preferable to use an expert in SVG visualizations. In Sonnet Writing, the Meta Model naturally leans on literary experts, notably Expert Poet and Expert Literary Critic, emphasizing creative and linguistic skills. This pattern demonstrates the Meta Model’s ability to dynamically tailor its expert engagement to the demands of the task, utilizing technical experts for computational challenges and a varied range of non-computational expertise for creative or abstract tasks.
Number of Rounds Taken to Reach a Solution. Examining the meta-prompting experiments involving a Python expert reveals that the average number of rounds required to reach a solution in the Meta Model varies significantly across tasks, indicative of their complexity and specific nature. Simpler tasks, such as Word Sorting (3.31 rounds) and Checkmate-in-One (3.48 rounds), typically necessitate fewer rounds, suggesting a more linear and straightforward resolution process, likely due to their clearly defined parameters. Conversely, more algorithmically challenging tasks like Python Programming Puzzles average a higher number of rounds at 6.07, reflecting the nuanced and multifaceted aspects of programming tasks that require extensive interactions for thorough clarification and iterative refinement. The Game of 24 and Multistep Arithmetic Two, with averages around 3.5 rounds, meld computational proficiency with logical reasoning, necessitating additional rounds for accurate and precise solutions. This observed correlation between the number of rounds and the task complexity underscores the Meta Model’s proficiency and adaptability. It efficiently manages simpler tasks with minimal interactions while skillfully handling the complexities of more challenging and heuristic-based problems, ensuring precision and efficacy in its solutions. This performance characteristic is particularly critical in environments where efficiency and interaction trade-off are key.
Enhancing Solution Reliability through Systematic Verification. The Meta Model’s systematic verification protocol strengthens the reliability and robustness of its solutions. Fundamental to this approach is the consistent practice of consulting an expert for validation before finalizing responses, a principle applied across diverse tasks. This method is further evidenced by the detailed interaction data. In tasks such as Checkmate in One, for instance, the Meta Model employs a two-step verification strategy. Initially, it consults an Expert Chess Player to come up with a solution, followed by a critical verification from an Expert Chess Analyst, ensuring strategic correctness. A similar approach is adopted in Sonnet Writing too, where an Expert Poet drafts the sonnet, and an Expert Poet Reviewer or Expert Essayist reviews it, making sure that the solution adheres to the strict rhyme scheme. This unsupervised but rigorous verification process extends to complex tasks like Game of 24 and MGSM, involving both external expert consultations and internal reviews. By integrating this dual verification mechanism, the model significantly enhances solution accuracy and reliability, essential for real-world applications where precision is paramount.
Navigating No-Solution Territories. Meta-prompting enables the Meta Model to acknowledge the absence or impossibility of a valid solution or its inability to find one more frequently than other prompting methods. In 100 examples of the Game of 24, the model reports no solution 9 times with Expert Python and 15 times without it, compared to the mere 2 instances under standard prompting. In Checkmate, across 250 examples, it admits to no solution 12 times without Expert Python and 10 times with it, a rarity in multipersona and standard prompting. While there were always solutions, it is arguably preferable to abstain from answering rather than provide an incorrect answer. Typically expressed as “No valid solution found” or more explicitly as “There is no solution to the 24 game with these numbers given the constraints,” these acknowledgments are likely the result of the model’s verification and feedback loop, emphasizing accuracy and confidence over speculative but incorrect responses.
Setting the Bar High: GPT-4’s Zero-Shot Task Solving Capabilities. Even without the enhanced capabilities of meta-prompting, GPT-4 stands out as an effective zero-shot task solver under standard prompting conditions. Its performance across various tasks, including Python Programming Puzzles and MGSM, is remarkable, particularly when compared to other LMs as highlighted by OpenAI (2023). GPT-4 excels as a task-agnostic solver, capable of processing and responding to diverse queries effectively. A significant attribute of GPT-4 is its proficiency in following instructions. Given clear and unambiguous natural-language instructions, the model demonstrates a high level of compliance and accuracy. This aspect of instruction-following is also a cornerstone of our meta-prompting framework, where we leverage GPT-4’s capabilities. Our experiments reinforce that GPT-4 excels in code generation, demonstrates impressive zero-shot reasoning, and engages effectively in role-playing, solidifying its position as a versatile and reliable LM.
Limited Performance Improvement with GPT-3.5. In comparison to GPT-4, GPT-3.5 demonstrates a more limited scope of performance enhancement across various tasks. Although it shows notable improvements in specific tasks such as Sonnet Writing and Checkmate-in-One, its capabilities do not consistently surpass baseline standards or zero-shot CoT prompting methods in other tasks, notably Word Sorting and Multiple Arithmetic Two. Our qualitative analysis suggests that GPT-3.5 may not be as effective as GPT-4 in simulating role-playing scenarios or managing extended context windows. This observation leads us to believe that factors such as the scale of the model, the quality and size of the instruction-following corpus may be significantly influencing the efficacy of the meta-prompting approach. Furthermore, it appears that the advantages offered by meta-prompting may even emerge more prominently at larger model scales.
2 Limitations and Failure Modes of Meta Prompting
The meta-prompting framework, despite its innovative approach, encounters several notable limitations, including cost efficiency, scalability, operational linearity, domain restrictions, information transfer challenges, and response patterns. A primary limitation is the elevated cost associated with multiple model calls. In our setup using GPT-4, the dual role of the Meta Model and the experts, distinguished by unique instructions, incurs substantial costs under the GPT-4 API pricing model. This cost factor diminishes the effectiveness of meta-prompting in smaller models like ChatGPT, which lack the comprehensive capabilities of GPT-4. Consequently, meta-prompting, though insightful, can become prohibitively expensive due to extensive model interactions and lengthy message histories. However, these costs will decrease as the costs of LMs decrease. Note that recent OpenAI API features announced after the experiments were run, namely the ability to run code in a sandbox directly through the API, could significantly decrease the costs of our system.
Another critical limitation is the requirement for substantial scale and a considerable context window. GPT-4 fits this criterion, but smaller models such as ChatGPT fall short. Meta-prompting’s design, characterized by extensive message histories, demands an LM capable of handling and retaining lengthy textual information, a feature not universally present in all LMs. Operational efficiency is also challenged by the linear (sequential) nature of meta-prompting. The framework, in its current form, processes steps one at a time, relying on the outcome of preceding calls. This dependency constrains the possibility of parallel processing, impacting the speed and efficiency of the system.
Additionally, our research confined meta-prompting within a closed-domain system. Nevertheless, the framework’s potential extends to incorporating external resources such as APIs, specialized finetuned models, search engines, or computational tools. More expansive implementations like AutoAgents (Chen et al., 2023a) and AutoGen (Wu et al., 2023), which include higher-level planning and diverse cooperation mechanisms, offer a glimpse into future directions. In subsequent versions, the Meta Model could benefit from refining or summarizing its history before advancing, optimizing the relevance and efficiency of the process. There is also untapped potential in concurrently summoning multiple experts or utilizing a single expert with varied temperature parameters to synthesize their outputs.
A practical challenge faced is the Meta Model’s occasional oversight in conveying necessary information to experts, forgetting that experts can only access data adhering to a certain format (within triple quotes in our system). This oversight can lead to unintended confusion and underscores the need for improved information management. Lastly, the Meta Model’s response pattern, particularly in tasks with lower performance, often includes apologies, such as “Apologies for the confusion in my previous response” or “I apologize for the previous incorrect solution.” This behavior likely stems from its training on instruction-following data.
This section seeks to contextualize our proposed meta-prompting approach amidst recent advancements in prompting strategies and scaffolding techniques. We provide a brief overview of these developments, highlighting their relevance and connections to our work.
Enhancing Reasoning in Language Models through Prompting. Recent efforts in LM scaffolding and prompting methods have significantly boosted the arithmetic and commonsense reasoning capabilities of LMs. The chain-of-thought (CoT) prompting (Wei et al., 2022b) and its variants—including least-to-most (Zhou et al., 2023), zero-shot CoT (Kojima et al., 2022), self-ask (Press et al., 2022), ask-me-anything (Arora et al., 2023), decomposed prompting (Khot et al., 2023), and auto-CoT (Zhang et al., 2023d)—have marked a paradigm shift in how LMs process complex queries. These methods encourage LMs to adopt human-like, sequential thinking processes, breaking down intricate questions into simpler subtasks and systematically solving them before presenting a final answer. Multiple studies (Wei et al., 2022a; Madaan and Yazdanbakhsh, 2022; Shi et al., 2023; Drozdov et al., 2023; Fu et al., 2023b; Suzgun et al., 2023b, inter alia) have shown the efficacy of these prompting methods across a broad set of tasks and benchmarks. More recent innovations such Tree-of-Thought (Yao et al., 2023a), Graph-of-Thought (Besta et al., 2023), Program-of-Thought (Chen et al., 2023d), and Skeleton-of-Thought (Ning et al., 2023), have further enriched this domain; these explore dynamic, non-linear reasoning pathways, broadening the computational and heuristic capabilities of LMs. However, they come with increased resource demands and greater time complexity, require multiple manual prompt crafting, and are often specialized for particular types of tasks.
Iterative Self-Feedback and Refinement Mechanisms. Recent instruction-following techniques and data-collection efforts have expanded the capabilities of LMs to follow instructions, emulate certain aspects of human behavior, and assist in tasks such as annotation and evaluation (Haluptzok et al., 2022; Aher et al., 2023). LMs such as GPT-4, PaLM, and Llama are now capable of effectively integrating self-feedback and refinement mechanisms through prompting and can leverage their own natural-language outputs to guide their behaviour and improve decision-making. SayCan (Ahn et al., 2022) and Inner Monologue (Huang et al., 2023) are early examples showcasing the benefits of inner dialogues in a closed-loop system for robotic control and action planning. Reflexion (Shinn et al., 2023) builds upon these studies and focuses on natural-language generation and reasoning tasks. It functions as a policy optimization mechanism through natural language feedback, using self-feedback and self-reflection to influence and correct behaviors in LMs, and has shown considerable success in preliminary experiments. In a more innovative vein, the Self-Taught Reasoner approach (STaR; Zelikman et al., 2022) iteratively trains an LM on its own outputs to refine initial rationales for more accurate solutions, leading to enhanced reasoning skills. Other notable methods such as Critic (Gou et al., 2023), Iterative Refinement (Chen et al., 2023b), RCI (Kim et al., 2023), Re3 (Yang et al., 2022), Refiner (Paul et al., 2023), Self-Critique (Saunders et al., 2022), Self-Correction (Welleck et al., 2023), Self-Eval, Self-Debug (Chen et al., 2023c), Self-Edit (Zhang et al., 2023a), Self-Evolve (Jiang et al., 2023b), Self-Taught Optimizer (SToP; Zelikman et al., 2023), and so forth, illustrate how verbal feedback, both internal and external, can significantly improve the accuracy, quality, and robustness of model outputs across various tasks and setups.
Exploring Role-Playing in Language Models. The integration of role-playing and self-collaboration concepts into LMs, grounded in cognitive psychology and developmental education principles, has emerged as a useful method for augmenting LMs’ problem-solving capabilities and optimizing their internal domain-specific knowledge and expertise. Recent studies (Park et al., 2022, 2023; Li et al., 2023; Xu et al., 2023; Fu et al., 2023a; Deshpande et al., 2023) have shown that endowing instruction-following LMs with “expert” personas or roles enhances the quality and accuracy of their output. In particular, approaches like CAMEL (Li et al., 2023) and Expert Prompting (Xu et al., 2023), which involve dynamically assigning personas to a single LM, have been shown to yield higher quality and more reliable responses than models without designated personas. Further investigations (Chen et al., 2023a, e; Du et al., 2023; Hao et al., 2023b; Liang et al., 2023; Liu et al., 2023b; Jiang et al., 2023a; Xiong et al., 2023; Zhang et al., 2023c) demonstrate that assigning multiple expert identities or roles to a single LM, tailored to specific tasks or problems, and prompting it to conduct multi-round internal dialogues—similar to a team of experts discussing and refining ideas—amplifies the reliability and comprehensiveness of the LM’s analysis; this leads to more well-rounded and thorough solutions. These studies advocate a complementary approach wherein multiple instances of an LM propose, debate, and refine their individual responses and reasoning in successive rounds, culminating in a unified final answer. This role-playing concept has shown to significantly improve mathematical and strategic reasoning across various tasks. Moreover, it improves the factual accuracy of the generated content, thereby reducing erroneous or fabricated responses.
Autonomous Decision-Making and Execution in Multi-Agent LM Systems. There has been a growing interest in using LMs for autonomous decision-making and task execution. Open-source projects like Auto-GPT, Agent-GPT, Baby-AGI, and LangChain are notable efforts developing agent protocols that are capable of planning, decision-making, and executing tasks end-to-end, with minimal or no human intervention. These systems highlight the potential and risks of LMs, which go beyond performing predefined tasks to adapting, learning, and autonomously executing decisions in real time. As discussed by Masa (2023), those autonomous models might be exploited by individuals with malicious intents and pose threats to humanity. There is also the dilemma of accountability: who bears responsibility when an LM-driven autonomous agent produces an inappropriate or criminal action? Ensuring safety and security with these agents is crucial, given its potential for mishaps or exploitation by malicious actors, and its vulnerability to cyber-attacks.
Integration of External Tools and APIs into Language Models. As LMs continue to evolve, the integration of external tools is becoming increasingly important. This tool-use integration, often achieved through in-context learning (e.g., Cai et al., 2023) or finetuning (e.g., Schick et al., 2023a), allows LMs to effectively engage with real-world scenarios and tackle a diverse range of dynamic tasks. Recent advancements (Cai et al., 2023; Gao et al., 2023; Gou et al., 2023; Hao et al., 2023c; Khattab et al., 2023; Lu et al., 2023; Qiao et al., 2023; Paranjape et al., 2023; Patil et al., 2023; Schick et al., 2023a; Yang et al., 2023; Yuan et al., 2023) have enabled LMs to perform accurate calculations, retrieve up-to-date information from search engines or databases, and interact with APIs, making them crucial for complex, multimodal real-world problems. OpenAI’s incorporation of predefined APIs and plugins into ChatGPT underscores the importance of external integration in developing a comprehensive LM ecosystem. However, most approaches often limit themselves to a select group of tools or domain-specific resources, posing challenges in adapting to new domains (Lu et al., 2023). Our meta-prompting approach, as detailed in Section 2, treats the LM as an independent tool and expert, available on-demand for specific tasks. Furthermore, incorporating a Python interpreter—through Expert Python—to execute and evaluate model-generated code has been instrumental in enhancing both accuracy and efficiency in various tasks.
In this work, we have introduced and examined meta-prompting, a simple yet powerful scaffolding technique that enhances the performance of language models in a task-agnostic manner. This approach leverages a language model to act as both a central conductor and a group of expert instances, thereby endowing traditional models with dynamic, multi-functional capabilities. A noteworthy aspect of meta-prompting lies in its proficiency to decompose complex tasks, engage distinct expertise for each component, and then integrate the varied outputs seamlessly. Demonstrating significant, double-digit improvements across a series of tasks, ranging from challenging arithmetic puzzles like the Game of 24 to the creative literary exercise of Shakespearean Sonnet Writing, meta-prompting promises to grow more potent and cost-efficient as language models continue to evolve, offering exciting prospects for future applications.
We would like to thank Federico Bianchi, Annabelle Carrell, Tayfun Gür, Dan Jurafsky, Suproteem Sarkar, Scott Duke Kominers, Lester Mackey, Neil Mallinar, Şule Kahraman, Deniz Keleş, Luke Melas-Kyriazi, Drew Pendergrass, Faiz Surani, Garrett Tanzer, Michael Wornow, and Eric Zelikman for their valuable comments, useful suggestions, and support.