Graph of Thoughts: Solving Elaborate Problems with Large Language Models

Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, Torsten Hoefler

Introduction

Large language models (LLMs) are taking over the world of AI. Recent years saw a rapid development of models primarily based on the decoder-only Transformer variant , such as GPT , PaLM , or LLaMA .

Prompt engineering is a resource-efficient approach for solving different LLM tasks. In brief, one includes the task description within the input sent to an LLM. If this description is appropriately formulated, the LLM solves the task using its autoregressive token-based mechanism for generating text. Such prompts may contain example tasks with solutions (few-shot prompting, also referred to as in-context learning (ICL)), or even no example tasks at all (zero-shot prompting). In recent years it was shown that this mechanism can be used to solve a broad set of tasks that involve mathematical, commonsense, or symbolic reasoning.

Chain-of-Thought (CoT) is an approach for prompting, in which one includes the intermediate steps of reasoning within the prompt (intermediate “thoughts”), besides the task input/output. CoT was shown to significantly improve the capability of LLMs to solve problems without resorting to any model updates. One major improvement over CoT, Self-Consistency with CoT (CoT-SC) , is a scheme where multiple CoTs are generated, and then the best one is selected as the outcome. More recently, CoT and CoT-SC were extended with Tree of Thoughts (ToT) , which models the LLM reasoning process with a tree. This facilitates using different paths of thoughts, and offers novel capabilities such as backtracking from non-promising outcomes. Unfortunately, the ToT approaches still fundamentally limit the reasoning abilities within a prompt by imposing the rigid tree structure on the thought process.

In this work, we argue that fundamentally more powerful prompting can be achieved by enabling LLM thoughts to form an arbitrary graph structure. This is motivated by numerous phenomena such as human reasoning, brain structure, or algorithmic execution. When working on a novel idea, a human would not only follow a chain of thoughts (as in CoT) or try different separate ones (as in ToT), but would actually form a more complex network of thoughts. For example, one could explore a certain chain of reasoning, backtrack and start a new one, then realize that a certain idea from the previous chain could be combined with the currently explored one, and merge them both into a new solution, taking advantage of their strengths and eliminating their weaknesses. Similarly, brains form complex networks, with graph-like patterns such as recurrence . Executing algorithms also expose networked patterns, often represented by Directed Acyclic Graphs. The corresponding graph-enabled transformations bring a promise of more powerful prompting when applied to LLM thoughts, but they are not naturally expressible with CoT or ToT.

We observe that these (and many other) thought transformations can be naturally enabled when modeling the reasoning process of an LLM as a graph. For this, we propose Graph of Thoughts (GoT), an approach that enhances LLMs’ capabilities through networked reasoning (contribution #1). In GoT, an LLM thought is modeled as a vertex, while an edge is a dependency between such thoughts. Using GoT, one can aggregate arbitrary thoughts by constructing vertices that have more than one incoming edge. Overall, the graph abstraction harnessed by GoT seamlessly generalizes CoT and ToT to more complex thought patterns, without resorting to any model updates.

Yet, putting GoT to practice requires solving several design challenges. For example, what is the best graph structure for different tasks? How to best aggregate thoughts to maximize accuracy and minimize cost? To answer these and many other questions, we carefully design a modular architecture for implementing GoT (contribution #2), coming with two design highlights. First, we enable a fine-grained control over individual thoughts. This enables us to fully control the ongoing conversation with the LLM, and apply advanced thought transformations, such as combining most promising thoughts from the ongoing reasoning into a new one. Second, we ensure that our architecture can be seamlessly extended with novel thought transformations, patterns of reasoning (i.e., graphs of thoughts), and LLM models. This enables rapid prototyping of novel prompting ideas using GoT, while experimenting with different models such as GPT-3.5, GPT-4, or Llama-2 .

We illustrate several use cases for GoT (sorting, keyword counting for summaries, set operations, document merging) and we detail how to implement them using the graph-based paradigm (contribution #3). We evaluate GoT and show its advantages over the state of the art (contribution #4). Overall, we observe that GoT is particularly well-suited for tasks that can be naturally decomposed into smaller subtasks that are solved individually and then merged for a final solution. Here, GoT outperforms other schemes, for example improving upon CoT and ToT by, respectively, ≈\approx70% and ≈\approx62%, in terms of the quality of sorting, while simultaneously reducing costs by >>31% over ToT.

We qualitatively compare GoT to other prompting schemesNote that we do not include a recent scheme called Graph-of-Thought because it is not a prompting scheme. While its name suggests close connections to ToT and CoT, as a fine-tuning scheme, it resorts to model updates, and is thus outside the focus of this work. Similarly, the graph-of-thoughts repository does not enable general graph-based reasoning and harnesses instead ToT with BFS. in Table 1. GoT is the only one to enable arbitrary graph-based thought transformations within a prompt, such as aggregation, embracing all previously proposed schemes.

Finally, we propose a new metric for evaluating a prompting strategy, the volume of a thought (contribution #5). With this metric, we aim to understand better the differences between prompting schemes. For a given thought vv, the volume of vv is the number of LLM thoughts, from which one can reach vv using directed edges. Intuitively, these are all the LLM thoughts that have had the potential to contribute to vv. We show that GoT, by incorporating thought transformations such as aggregation, enables thoughts to have fundamentally larger volumes than other schemes.

Background & Notation

We first outline background concepts and notation.

The conversation with the LLM consists of user messages (prompts) and LLM replies (thoughts). We follow the established notation and we denote a pre-trained language model (LM) with parameters θ\theta as pθp_{\theta}. Lowercase letters such as x,y,z,...x,y,z,... indicate LLM thoughts. We purposefully do not prescribe what is a single “thought”, and instead make it use-case specific. Hence, a single thought can be a paragraph (e.g., in article summary), a document (e.g., in document generation), a block of code (e.g., in code debugging or optimization), and so on.

We next describe specific prompting approaches.

The Input-Output (IO) prompting is a straightforward approach, in which we use an LLM to turn an input sequence xx into the output yy directly, without any intermediate thoughts.

Chain-of-Thought (CoT)

Second, in Chain-of-Thought (CoT), one introduces intermediate thoughts a1,a2,...a_{1},a_{2},... between xx and yy. This strategy was shown to significantly enhance various LM tasks over the plain IO baseline, such as mathematical puzzles or general mathematical reasoning .

Multiple CoTs

Third, one can generalize CoT into multiple CoTs by generating several (independent) kk CoTs, and returning the one with the best output (according to some prescribed scoring metric). It was introduced by Wang et al. in the scheme called Self-Consistency with CoT (CoT-SC) . This approach enhances CoT because it offers an opportunity to explore different reasoning paths. However, it does not offer “local exploration” within a path, such as backtracking.

Tree of Thoughts (ToT)

Finally, the Tree of Thoughts (ToT) scheme was introduced independently by Yao and Long (where it is referred to as Tree-of-Thought); it was used implicitly to a certain degree by other schemes such as thought decomposition . It enhances CoT-SC by modeling the process or reasoning as a tree of thoughts. A single tree node represents a partial solution. Based on a given node, the thought generator constructs a given number kk of new nodes. Then, the state evaluator generates scores for each such new node. Depending on the use case, the evaluation could be conducted using an LLM itself, or it can harness human scores. Finally, the schedule of extending the tree is dictated by the utilized search algorithm (for example BFS or DFS).

The GoT Framework

We now detail the GoT framework. We present it in Figure 1, and compare it to other prompting strategies.

Formally, GoT can be modeled as a tuple (G,T,E,R)(G,\mathcal{T},\mathcal{E},\mathcal{R}), where GG is the “LLM reasoning process” (i.e., all the LLM thoughts within the context, with their relationships), T\mathcal{T} are the potential thought transformations, E\mathcal{E} is an evaluator function used to obtain scores of thoughts, and R\mathcal{R} is a ranking function used to select most relevant thoughts.

We model the reasoning process as a directed graph G=(V,E)G=(V,E); VV is a set of vertices and E⊆V×VE\subseteq V\times V is a set of edges. GG is directed and thus the edges are a subset of ordered vertex pairs E⊆V×VE\subseteq V\times V. A vertex contains a solution to a problem at hand (be it an initial, intermediate, or a final one). The concrete form of such a thought depends on the use case; it could be a paragraph (in writing tasks) or a sequence of numbers (in sorting). A directed edge (t1,t2)(t_{1},t_{2}) indicates that thought t2t_{2} has been constructed using t1t_{1} as “direct input”, i.e., by explicitly instructing the LLM to use t1t_{1} for generating t2t_{2}.

In certain use cases, graph nodes belong to different classes. For example, in writing tasks, some vertices model plans of writing a paragraph, while other vertices model the actual paragraphs of text. In such cases, GoT embraces a heterogeneous graph G=(V,E,c)G=(V,E,c) to model the LLM reasoning, where cc maps vertices VV into their respective classes CC (in the above case, it would be C={plan,par}C=\{plan,par\}). Hence, any vertex vv can model different aspects of reasoning.

We associate GG with the LLM reasoning process. To advance this process, one applies thought transformations to GG. An example of such a transformation is to merge best-scoring (so far) thoughts into a new one. Another example is to loop over a thought, in order to enhance it. Note that these transformations strictly extend the set of transformations available in the CoT, CoT-SC, or ToT.

2 Transformations of Thoughts

GoT enables novel transformations of thoughts thanks to the graph-based model for reasoning. We refer to them as graph-enabled transformations. For example, in writing, one could combine several input articles into one coherent summary. In sorting, one could merge several sorted subarrays of numbers into a final sorted array. We illustrate examples of aggregation and generation in Figure 2.

Formally, each such transformation can be modeled as T(G,pθ)\mathcal{T}(G,p_{\theta}) where G=(V,E)G=(V,E) is the graph reflecting the current state of the reasoning, and pθp_{\theta} is the used LLM. T\mathcal{T} modifies GG usually by adding new vertices and their incoming edges. We have G′=T(G,pθ)=(V′,E′)G^{\prime}=\mathcal{T}(G,p_{\theta})=(V^{\prime},E^{\prime}), where V′=(V∪V+)∖V−V^{\prime}=(V\cup V^{+})\setminus V^{-} and E′=(E∪E+)∖E−E^{\prime}=(E\cup E^{+})\setminus E^{-}. V+V^{+} and E+E^{+} are new vertices and edges inserted into GG to model the new thoughts and their dependencies, respectively. To maximize the expressiveness of GoT – we also enable the user to explicitly remove thoughts, by specifying the corresponding vertices and edges to be removed (V−V^{-} and E−E^{-}, respectively). Here, it is the user’s responsibility to ensure that the sets V+,E+,V−,V^{+},E^{+},V^{-}, and E−E^{-} come with consistent transformations (i.e., for example, that the user does not attempt to remove a vertex that does not exist). This enables seamless incorporation of schemes where, in order to save space within the context, one can remove parts of reasoning that do not promise improvements.

The specific form of T\mathcal{T} and how it impacts GG depends on a specific transformation. We first detail the primary graph-enabled thought transformations, and then proceed to describe how GoT embraces the transformations from the earlier schemes. Unless stated otherwise, V−=E−=∅V^{-}=E^{-}=\emptyset.

First, with GoT, one can aggregate arbitrary thoughts into new ones, to combine and reinforce the advantages of these thoughts, while eliminating their disadvantages. In the basic form, in which only one new vertex is created, V+={v+}V^{+}=\{v^{+}\} and E+={(v1,v+),...,(vk,v+)}E^{+}=\{(v_{1},v^{+}),...,(v_{k},v^{+})\}, where v1,...,vkv_{1},...,v_{k} are the merged kk thoughts. More generally, this enables aggregating reasoning paths, i.e., longer chains of thoughts, beyond just individual thoughts. With the graph model, it is simply achieved by adding outgoing edges from the vertices v1,...,vkv_{1},...,v_{k}, modeling final thoughts in several chains, into a single thought v+v^{+} combining these chains.

Refining Transformations

Another thought transformation is the refining of a current thought vv by modifying its content: V+={}V^{+}=\{\} and E+={(v,v)}E^{+}=\{(v,v)\}. This loop in the graph indicates an iterated thought with the same connections as the original thought.

Generation Transformations

Finally, one can generate one or more new thoughts based on an existing single thought vv. This class embraces analogous reasoning steps from earlier schemes, such as ToT or CoT-SC. Formally, we have V+={v1+,...,vk+}V^{+}=\{v^{+}_{1},...,v^{+}_{k}\} and E+={(v,v1+),...,(v,vk+)}E^{+}=\{(v,v^{+}_{1}),...,(v,v^{+}_{k})\}.

3 Scoring & Ranking Thoughts

Thoughts are scored to understand whether the current solution is good enough. A score is modeled as a general function E(v,G,pθ)\mathcal{E}(v,G,p_{\theta}), where vv is a thought to be evaluated. We use the state of the whole reasoning process (GG) in E\mathcal{E} for maximum generality, because – for example – in some evaluation scenarios, scores may be relative to other thoughts.

GoT can also rank thoughts. We model this with a function R(G,pθ,h)\mathcal{R}(G,p_{\theta},h) where hh specifies the number of highest-ranking thoughts in GG to be returned by R\mathcal{R}. While the specific form of R\mathcal{R} depends on the use case, we most often use a simple yet effective strategy where hh thoughts with the highest scores are returned, i.e., v1,...,vh=R(G,pθ,h)v_{1},...,v_{h}=\mathcal{R}(G,p_{\theta},h).

Specific forms of E\mathcal{E} and R\mathcal{R} depend on the use case. We discuss the details in Section 5. For example, the score (or rank) for sorting corresponds to the count of elements correctly sorted (or incorrectly, when using the error as a score).

System Architecture & Extensibility

The GoT architecture consists of a set of interacting modules, see Figure 3 (the blue part). These modules are the Prompter (prepares the messages for the LLM), the Parser (extracts information from LLM thoughts), the Scoring module (verifies and scores the LLM thoughts), and the Controller (coordinates the entire reasoning process, and decides on how to progress it). The Controller contains two further important elements: the Graph of Operations (GoO) and the Graph Reasoning State (GRS). GoO is a static structure that specifies the graph decomposition of a given task, i.e., it prescribes transformations to be applied to LLM thoughts, together with their order & dependencies. GRS is a dynamic structure that maintains the state of the ongoing LLM reasoning process (the history of its thoughts and their states).

The Prompter prepares the prompts to be sent to the LLM. This module is responsible for the specifics of encoding the graph structure within the prompt. The GoT architecture enables the user to implement use case specific graph encodings by providing full access to the graph structure.

2 Parser

The Parser extracts information from LLM thoughts. For each such thought, the Parser constructs the thought state, which contains this extracted information. The thought state is then used to update the GRS accordingly.

3 Scoring & Validation

Here, we verify whether a given LLM thought satisfies potential correctness conditions, and then we assign it a score. Depending on how the score is derived, the module may consult the LLM. Moreover, depending on the use case, the score may also be assigned by a human. Finally, use cases such as sorting use simple local scoring functions.

4 Controller

The Controller implements a specific strategy for selecting thoughts from its GRS structure. It also selects what transformations should be applied to which thoughts, and then passes this information to the Prompter. It also decides whether the whole process should be finalized, or whether the next round of interaction with the LLM should be initiated. In our current design, this is dictated by the execution plan specified in the GoO.

5 GoO & GRS

The user constructs a GoO instance, which prescribes the execution plan of thought operations. The GoO is a static structure that is constructed once, before the execution starts. Each operation object knows its predecessor and successor operations. Then, during the execution, an instance of the GRS maintains the continually updated information about the LLM reasoning process. This includes which operation has been executed so far, the states of all the generated LLM thoughts, their validity and scores, and any other relevant information.

The above elements offer extensible APIs, enabling straightforward implementations of different prompting schemes. The APIs are outlines in the green part of Figure 3, and detailed in the documentation. We also provide examples of prompts used by these operations and a corresponding GRS in the red part of Figure 3.

Example Use Cases

We now describe several use cases of GoT. We detail one use case (sorting) and summarize the others.

We focus on the decomposition of the sorting use case and Graph of Operations, which are central for implementing and executing any workload within GoT.

We consider sorting numbers 0–9 with duplicates. The considered LLMs are unable to sort a sequence of such numbers correctly beyond a certain length consistently because duplicate counts do not match.

In GoT, we employ merge-based sorting: First, one decomposes the input sequence of numbers into subarrays. Then, one sorts these subarrays individually, and then respectively merges them into a final solution. Figure 4 illustrates this use case together with its graph decomposition. Here, an LLM thought is a sequence of sorted numbers.

To score an outcome, denote an input sequence with [a1,a2,...,an][a_{1},a_{2},...,a_{n}] and an output one with [b1,b2,...,bm][b_{1},b_{2},...,b_{m}]. We use the following score that determines “the scope” of errors:

where p∈{1,...,m}p\in\{1,...,m\}, q∈{1,...,n}q\in\{1,...,n\}, and

Here, XX indicates how many consecutive pairs of numbers are incorrectly sorted. If two numbers ii and i+1i+1 are incorrectly sorted (i.e., bi>bi+1b_{i}>b_{i+1}), then the expression within the summation returns 1, increasing the error score by one. For two numbers correctly sorted, this expression amounts to 0. Then, YY determines how well a given output sequence preserves the frequency of output numbers. Specifically, for each considered number xx (x∈{0,...,9}x\in\{0,...,9\}), we obtain the difference between the count of input elements being equal to xx, vs. the count of output elements equal to xx. For an output sequence perfectly preserving the frequency of xx, this would amount to 0. Any single “deviation” in this count, increases the “error scope” by 1. We then sum this over all considered values of xx. When plotting this score, to improve the clarity of plots, we additionally apply clipping min⁡(error-scope,n)\min(\text{error-scope},n), as some baselines (IO, CoT) result in large numbers of outliers with high error scope. Finally, to use a “positive score” describing “the scope of correctly sorted” elements, one can use the value max⁡(n−error-scope,0)\max(n-\text{error-scope},0).

2 Set Operations

Moreover, we also consider set operations, focusing on set intersection. They have numerous applications (particularly set intersection) in problems ranging from genome or document comparisons to pattern matching . Set intersection of two sets is implemented similarly as the sorting. The second input set is split into subsets and the intersection of those subsets with the first input set is determined with the help of the LLM. Afterwards the resulting intersection sets are aggregated for the final results. For the evaluation we use different set sizes of 32, 64 and 128 elements and we vary the number of elements found in both sets to be between 25% and 75%.

Our score indicates the total number of missing or incorrectly included elements in the final intersection. Specifically, denote two input sets with A=[a1,a2,...,an]A=[a_{1},a_{2},...,a_{n}] and B=[b1,b2,...,bn]B=[b_{1},b_{2},...,b_{n}], and the output set with C=[c1,c2,...,cm]C=[c_{1},c_{2},...,c_{m}]. Then,

where X1=∣C∖(A∩B)∣X_{1}=|C\setminus(A\cap B)| are the number of elements in CC that are not supposed to be there, X2=∣(A∩B)∖C∣X_{2}=|(A\cap B)\setminus C| are the number of elements missing from CC, and XdX_{d} is the number of duplicates in CC (because the LLM expresses the set as a list in natural language). Finally, to use a “positive score” describing “the scope of correctly computed” elements, one can use the value max⁡(n−error-scope,0)\max(n-\text{error-scope},0).

3 Keyword Counting

Keyword counting finds the frequency of keywords in a given category (countries in our example implementation) within the input text. GoT splits the input text into multiple passages, counts the keywords in each passage and aggregates the subresults. The number of passages is configurable and can also be left to the LLM, making it possible to treat each sentence as a separate passage. Here, to score a thought, we first – for each keyword – derive the absolute difference between the computed count and the correct one. We then sum all these differences to get the final score.

4 Document Merging

Finally, we also provide document merging. Here, the goal is to generate a new Non-Disclosure Agreement (NDA) document based on several input ones that partially overlap in terms of their contents. The goal is to ensure minimal amount of duplication, while maximizing information retention. Document merging is broadly applicable in, e.g., legal procedures, where multiple sources of information have to be combined into a single document or article. To score a solution, we query the LLM for two values (3 times for each value, and take the average). The first value corresponds to the solution redundancy (10 indicates no redundancy, 0 implies at least half the information is redundant), the second value stands for information retention (10 indicates all information is retained, 0 says that none is retained). We compute the harmonic mean of these values.

The Latency-Volume Tradeoff

We now show that GoT improves upon previous prompting schemes in terms of the tradeoff between latency (number of hops in the graph of thoughts to reach a given final thought) and volume. We define volume – for a given thought tt – as the number of preceding LLM thoughts that could have impacted tt. Formally, the volume of tt is the number of thoughts from which there exists a path to tt in the graph of thoughts. We assume that outputting a single thought costs O(1)O(1) time and fix the total cost to Θ(n)\Theta(n) for each prompting scheme.

The structure of the schemes is as follows. CoT-SC consists of kk independent chains originating from a single starting thought. ToT is a complete kk-ary tree. Finally, in GoT, a complete kk-ary tree is joined at its leaves with a “mirrored” kk-ary tree of the same size but with its edges reversed.

The analysis is detailed in Table 2. CoT offers a large volume of up to NN, but at the cost of a high latency of NN. CoT-SC reduces the latency by a factor of kk (which corresponds to its branching factor), but it simultaneously decreases the volume by kk as well. ToT offers a latency of log⁡kN\log_{k}N but also has low volume. GoT is the only scheme to come with both a low latency of log⁡kN\log_{k}N and a high volume NN. This is enabled by the fact that GoT harnesses aggregations of thoughts, making it possible to reach the final thought from any other intermediate thought in the graph decomposition.

Evaluation

We show the advantages of GoT over the state of the art. We focus on comparing GoT to ToT, as it was shown to consistently outperform other schemes. Still, for a broad comparison, we also experiment with IO, CoT, and CoT-SC. As our analysis results in a large evaluation space, we present representative results and omit data that does not bring relevant insights (e.g., CoT-SC).

We use 100 input samples for each task and comparison baseline. We set the temperature to 1.0 and use a 4k context size unless stated otherwise. For each experiment, we fix the numbers of thoughts in respective schemes to achieve similar costs in each experiment.

Parameters We experiment extensively with the branching factor kk and the number of levels LL to ensure that we compare GoT to cost-effective and advantageous configurations. We plot two variants of ToT: one with higher kk and lower depth (ToT), the other with lower kk but higher LL (ToT2). We usually aim to achieve a sweet spot in the tradeoff between sparser generation rounds (lower kk) vs. more rounds (larger LL). Usually more responses per round is more expensive (e.g., 80 vs. 60 total responses for Figure 7 but 6vs.6 vs.3 costs). We also try different problem sizes PP (e.g., in sorting, PP states how many numbers are to be sorted).

Used LLMs Due to budget restrictions, we focus on GPT-3.5. We also experimented with Llama-2, but it was usually worse than GPT-3.5 and also much slower to run, making it infeasible to obtain enough samples.

2 Analysis of GoT’s Advantages

The results of the analysis are in Figure 5 (sorting), 6 (set intersection), 7 (keyword counting), and 8 (document merging); see Section 5 for the description of specific use cases. Overall, GoT improves the quality of outcomes over all the considered baselines and it reduces inference costs compared to ToT.

GoT vs. ToT GoT improves upon ToT and ToT2 by a large margin over all the considered problem instances. ToT usually comes with somewhat higher quality than ToT2, but simultaneously much higher costs. GoT’s costs are always lower than ToT, and comparable (in some cases lower, in others higher) to ToT2. For example, it reduces median error by ≈\approx62%, thereby achieving a higher quality of sorting, for P=128P=128 in comparison to ToT while ensuring >>31% cost reductions. These advantages are due to GoT’s ability to decompose complex tasks into simpler subtasks, solve these subtasks independently, and then incrementally merge these outcomes into the final result.

GoT vs. IO and CoT GoT consistently delivers much higher quality of outcomes than IO/CoT. For example, for sorting (P=64P=64), GoT’s median error is ≈\approx65% and ≈\approx83% lower than, respectively, CoT and IO. Yet, the costs of GoT – and ToT – are much higher than in IO and CoT. This is mostly due to our configuration of CoT, where we do not artificially inflate the lengths of the chains of reasoning if this does not improve the outcomes. The higher costs of GoT and ToT are driven by kk new thoughts built for each Generate operation; these multiple thoughts are one of the reasons for GoT’s superiority in quality.

Increasing Complexity of Tackled Problems Most importantly, the advantages of GoT in the quality increase for all the baselines with the growing size of the problem PP. For example, in sorting, while for P=32P=32 GoT only negligibly improves upon ToT2, its median error count becomes lower by ≈\approx61% for P=64P=64 and ≈\approx69% for P=128P=128. The quartiles also become respectively better. The results for other schemes also follow the intuition; for example, IO becomes consistently worse with the increasing PP, which is expected as a single thought is unlikely to solve a large problem instance. Overall, this analysis illustrates that GoT is indeed well-suited for elaborate problem cases, as the execution schedules usually become more complex with the growing problem sizes.

3 Discussion on Task Decomposition

When splitting a task into subtasks and then solving these subtasks, the size of responses and the input (in tokens) are reduced proportionally to the degree of the task decomposition. However, the “static” part of the prompt (i.e., few-shot examples) may become a significant overhead (see GoT4 to GoT8 in Figure 7). Here, we observe that these few-shot examples can usually also be reduced in size (e.g., the passages used to demonstrate keyword counting can also be made smaller and still be indicative of the actual input size), thus actively working towards decreasing the cost (e.g., see the difference between GoT8 and GoTx in Figure 7).

The overall goal when conducting graph decomposition is to break down a task to the point, where the LLM can solve it correctly for the majority of time using a single prompt (or with a few additional improvement steps). This significantly lowers the number of improvement/refinement steps needed during the later stages of the graph exploration. Furthermore, as indicated by our results, combining or concatenating subresults is usually an easier task than solving large task instances from scratch. Hence, the LLM is often successful when aggregating the final solution.

Related Work

We summarize relations between GoT and related work.

We detail different prompting paradigms in Section 1 and Table 1. There are numerous other works related to prompting. We now briefly summarize selected most related ones; more extensive descriptions can be found in dedicated surveys . Wang et al. proposed Plan-and-Solve, an approach to enhance CoT with an explicit planning stage . Using complexity-based criteria to enhance prompting within a CoT was designed by Fu et al. . The self-taught reasoner (STaR) generates several chain of thoughts, and selects the ones that are valid. Similarly, a scheme by Shum et al. generates a pool of CoT candidates, and selects the best candidate based on whether the candidates match the ground truth and on a policy gradient-based method. Automatic prompt generation overcomes the issues of scaling in CoT . Zhou et al. propose to harness selecting the best prompt out of a candidate set . Skeleon-of-Thought generates at first a number of skeleton answers (brief bullet points of 3 to 5 words) and expands on these points in parallel in a second step.

Finally, in prompt chaining, one cascades different LLMs. This enables prompting different LLMs via different contexts, enabling more powerful reasoning . GoT is orthogonal to this class of schemes, as it focuses on a single context capabilities.

2 Self-Reflection & Self-Evaluation

Self-reflection and self-evaluation were introduced recently . They are used to enhance different tasks, for example for code generation or computer operation tasks . In GoT, we partially rely on self-evaluation when taking decisions on how to expand the graph of thoughts within a prompt.

3 LLMs & Planning

There are many works recently on how to plan complex tasks with LLMs . GoT could be seen as a generic framework that could potentially be used to enhance such schemes, by offering a paradigm for generating complex graph-based plans.

4 Graphs and Graph Computing

Graphs have become an immensely popular and important part of the general computing landscape . Recently, there has been a growing interest in domains such as graph databases , graph pattern matching , graph streaming , and graph machine learning as well as graph neural networks . The graph abstraction has been fruitful for many modern research domains, such as social sciences (e.g., studying human interactions), bioinformatics (e.g., analyzing protein structures), chemistry (e.g., designing chemical compounds), medicine (e.g., drug discovery), cybersecurity (e.g., identifying intruder machines), healthcare (e.g., exposing groups of people who submit fraudulent claims), web graph analysis (e.g., providing accurate search services), entertainment services (e.g., predicting movie popularity), linguistics (e.g., modeling relationships between words), transportation (e.g., finding efficient routes), physics (e.g., understanding phase transitions and critical phenomena), and many others . In this work, we harness the graph abstraction as a key mechanism that enhances prompting capabilities in LLMs.

Conclusion

Prompt engineering is one of the central new domains of the large language model (LLM) research. It enables using LLMs efficiently, without any model updates. However, designing effective prompts is a challenging task.

In this work, we propose Graph of Thoughts (GoT), a new paradigm that enables the LLM to solve different tasks effectively without any model updates. The key idea is to model the LLM reasoning as an arbitrary graph, where thoughts are vertices and dependencies between thoughts are edges. This enables novel transformations of thoughts, such as aggregation. Human’s task solving is often non-linear, and it involves combining intermediate solutions into final ones, or changing the flow of reasoning upon discovering new insights. GoT reflects this with its graph structure.

GoT outperforms other prompting schemes, for example ensuring 62% increase in the quality of sorting over ToT, while simultaneously reducing costs by >>31%. We also propose a novel metric for a prompting scheme, the volume of a thought, to indicate the scope of information that a given LLM output could carry with it, where GoT also excels. This provides a step towards more principled prompt engineering.

The graph abstraction has been the foundation of several successful designs in computing and AI over last decades, for example AlphaFold for protein predictions. Our work harnesses it within the realm of prompt engineering.

Acknowledgements

We thank Hussein Harake, Colin McMurtrie, Mark Klein, Angelo Mangili, and the whole CSCS team granting access to the Ault and Daint machines, and for their excellent technical support. We thank Timo Schneider for help with infrastructure at SPCL. This project received funding from the European Research Council (Project PSAP, No. 101002047), and the European High-Performance Computing Joint Undertaking (JU) under grant agreement No. 955513 (MAELSTROM). This project was supported by the ETH Future Computing Laboratory (EFCL), financed by a donation from Huawei Technologies. This project received funding from the European Union’s HE research and innovation programme under the grant agreement No. 101070141 (Project GLACIATION).

References

Appendix A Positive Score Evaluation

The following figures plot the same data as Figures 5 and 6 respectively, however use the ”positive score” described in Sections 5.1 and 5.2.

Appendix B Example Prompts - Sorting

We present the prompts only for the sorting of 32-element lists, as those for 64-element and 128-element lists are identical, except for the split_prompt where the number of elements in the one-shot example matches the problem size.

For sorting, we employ three distinct types of operations that interact with the LLM, each with its corresponding prompts. First, there is the Generate operation, utilizing the sort_prompt to guide the LLM in sorting a provided list of values, and the split_prompt to direct the LLM to split a specified list into a designated number of sublists. Next, the Improve operation employs the improve_prompt to instruct the LLM to refine a sorted list if it detects mistakes. Finally, the Aggregate operation leverages the merge_prompt to guide the LLM in merging two pre-sorted lists into a single sorted list.

First, we present the prompt stubs (Table 3), serving as templates to dynamically generate appropriate prompts at runtime. For clarity, we display their corresponding few-shot examples separately in Table 4. Following this, we outline the LLM interactions throughout the process of solving the sorting use case (Table 5 - Table 9).

Appendix C Example Prompts - Set Intersection

We present the prompts only for the intersection of two 32-element sets, as those for 64-element and 128-element sets are identical, except for the split_prompt where the size of the split is adjusted proportionally.

For set intersection, we employ two distinct types of operations that interact with the LLM, each with its corresponding prompts. First, there is the Generate operation, utilizing the intersect_prompt to guide the LLM in intersecting two input sets, and the split_prompt to direct the LLM to split a specified set into a designated number of distinct subsets. Second, the Aggregate operation leverages the merge_prompt to guide the LLM in combining two sets into one.

First, we present the prompt stubs (Table 10), serving as templates to dynamically generate appropriate prompts at runtime. For clarity, we display their corresponding few-shot examples separately in Table 11. Following this, we outline the LLM interactions throughout a complete set intersection process (Table 12 - Table 15).

Appendix D Example Prompts - Keyword Counting

We present the prompts only for GoT4 of the keyword counting task, as those used for GoT8 and GoTx are identical, except for minor differences in the split_prompt where the size of the split is adjusted.

For keyword counting, we employ three distinct types of operations that interact with the LLM, each with its corresponding prompts. First, there is the Generate operation, utilizing the count_prompt to guide the LLM in counting the keywords in a text, and the split_prompt to direct the LLM to split a given text into a number of passages. Next, the Aggregate operation leverages the merge_prompt to guide the LLM in merging two dictionaries of counted keywords into one. Finally, the ValidateAndImprove operation employs the improve_merge_prompt to instruct the LLM to correct mistakes that were made in a previous Aggregate operation.

We present the prompt stubs (Table 16 - Table 17), serving as templates to dynamically generate appropriate prompts at runtime. For clarity, we display their corresponding few-shot examples separately in Table 18 and Table 19. Following this, we outline the LLM interactions throughout a complete keyword counting process (Table 20 - Table 28).

Appendix E Example Prompts - Document Merging

We present the prompts only for GoT of the document merging task, as GoT2 only differs in the fact that it merges the 4 NDAs in 2 steps rather than 1. For document merging, we employ four distinct types of operations that interact with the LLM, each with its corresponding prompts. First, there is the Generate operation, utilizing the merge_prompt to instruct the LLM to merge the 4 NDAs into 1. Second, the Score operations instructs the LLM to score a given merged NDA using the score_prompt. Next, the Aggregate operation employs the aggregate_prompt to instruct the LLM to aggregate multiple merge attempts into a single, better one. Finally, the Improve operation leverages the improve_prompt to instruct the LLM to improve a merged NDA.

First, we present the prompt stubs (Table 29 - Table 30), serving as templates to dynamically generate appropriate prompts at runtime. Following this, we outline the LLM interactions throughout a complete merging process (Table 31 - Table 49). However, instead of displaying each input/generated NDA in every prompt/response, we present the 4 input NDAs in Table 31 - Table 33 and the final merged NDA in Table 49. Furthermore, as scoring is done using the LLM as well, we will present these interactions for the best performing merged NDAs (Tables 39 - 40 and Tables 47 - 48). Lastly, most responses are limited to a few lines only, as they don’t offer any further insights and would otherwise span multiple pages. However, we refer the interested reader to the results in the corresponding code repositoryhttps://github.com/spcl/graph-of-thoughts for full logs and further examples.

Appendix F Evaluation - GoT Configurations

We detail the concrete operations that GoT was configured with to solve the set intersection and sorting use cases.