Unleashing the Emergent Cognitive Synergy in Large Language Models: A Task-Solving Agent through Multi-Persona Self-Collaboration

Zhenhailong Wang, Shaoguang Mao, Wenshan Wu, Tao Ge, Furu Wei, Heng Ji

Introduction

Although large language models (LLMs) have demonstrated impressive performance as general task-solving agents, they still encounter challenges (Qin et al., 2023; Bang et al., 2023; OpenAI, 2023b; Bubeck et al., 2023) in various knowledge-intensive and reasoning-intensive tasks due to factual hallucination (Maynez et al., 2020) and a lack of slow-thinking (Sloman, 1996) capabilities. Unlike humans, who can leverage the power of collaboration and information integration among different cognitive processes and individuals (referred to as cognitive synergy (Curşeu et al., 2015; Goertzel, 2009, 2017)), current LLMs are akin to "jack-of-all-trades" with a vast mixture of knowledge and characteristics. Recent advancements, such as Chain-of-Thought (CoT) prompting (Wei et al., 2023; Kojima et al., 2022) and Self-refinement (Madaan et al., 2023; Shinn et al., 2023), have successfully enhanced the reasoning abilities of LLMs by simulating slow-thinking through the generation of intermediate steps or iterative revision. However, factual hallucination remains a major challenge for LLMs on knowledge-intensive tasks.

A cognitive synergist is an intelligent agent that collaborates with multiple minds to enhance problem-solving and efficacy in complex tasks. In this work, we aim to create a cognitive synergist based on a single LLM that can "split into" multiple personas and engage in self-collaboration to solve both knowledge-intensive and reasoning-intensive tasks. This idea is heavily inspired by the role of pretend play (Piaget, 1954; Pellegrini, 2009) in cognitive development and recent findings that assigning personas (Deshpande et al., 2023; Xu et al., 2023) to LLMs can elicit specific behaviors, improve answer quality, and potentially build an AI society Park et al. (2023); Schick et al. (2022); Li et al. (2023); Cai et al. (2023) with collaborative LLM agents. However, as shown in Table 1, previous works have limitations such as fixed or task-specific personas, the need for additional fine-tuning, and increased inference costs due to multiple LLM instances.

To unleash the potential of cognitive synergy for general task-solving, we propose Solo Performance Prompting (SPP), which prompts a single LLM to identify, simulate, and collaborate with multiple personas. Figure 1 provides a high-level overview of SPP. Here, a persona can represent either a domain expert, such as a movie enthusiast, or a target audience, such as a ten-year-old child. Through the dynamic identification of various personas, we empower a single LLM to acquire diverse domain knowledge accurately without additional retrieval systems. By facilitating multi-turn self-collaboration, we enable self-revision and self-feedback from various perspectives without requiring additional agents.

In real-world scenarios, such as those in creative industries, there is often a need to incorporate diverse information from different domains. Figure 2 presents a concrete example of how SPP operates on a challenging task that requires creative integration of information from various domains, such as the Legend of Zelda game, Harry Potter movies, and Jay Chou’s albums. Standard prompting fails to generate satisfactory output due to missing essential information and factual errors. In contrast, SPP produces informative and coherent answers by automatically identifying expert personas and engaging in a multi-turn self-collaboration. In this process, the AI Assistant persona iteratively writes drafts of the story, solicits feedback from other participants, and revises accordingly.

To explore the prevalence of cognitive synergy in different LLMs, we apply SPP to LLMs with varying scales and capabilities, including GPT-4, GPT-3.5-turbo, and Llama-13b-chat. Comparative results show that cognitive synergy only emerges in GPT-4 and not in less capable models. This draws an interesting analogy to human development, as children typically start engaging in role-playing at the age of 2 to 3 Piaget (1954), but not earlier. In summary, the key contributions of this paper are as follows:

We investigate whether LLMs can leveraging cognitive synergy for general task-solving. We introduce Solo Performance Prompting (SPP), which simulates multi-agent, multi-persona collaboration in a pure zero-shot manner.

We evaluate SPP across three challenging tasks: Trivia Creative Writing, Codenames Collaborative and Logic Grid Puzzle, spanning both knowledge- and reasoning-intensive domains. To our knowledge, SPP is the first zero-shot prompting method that can enhance both knowledge and reasoning abilities on GPT-4.

We present an intriguing finding regarding the emergent nature of cognitive synergy ability in LLMs, which only emerges in GPT-4 and not in less powerful models.

We conduct in-depth analyses of the impact of the identified personas and SPP prompt design, providing insights into why dynamic, fine-grained personas are necessary, as opposed to fixed, coarse-grained personas.

Solo Performance Prompting

To unleash the power of synergizing different personas to tackle complex problems, we propose Solo Performance Prompting (SPP) which instructs a LLM to perform the following the procedure for general task-solving: (1) Persona Identification: Identify multiple participants with special personas (including a leader persona: AI Assistant) that are essential for solving the particular task. (2) Brainstorming: The participants share knowledge and provide suggestions on how to approach the task based on their own expertise. (3) Multi-Persona Iterative Collaboration: The leader persona, AI Assistant, proposes initial solutions, consults the other participants for feedback, and revise the answer iteratively. Figure 2 shows a walking example of SPP during inference. Next, we formally describe the SPP procedure in detail.

Given an input sequence xx and a model M\mathcal{M}, let a prompt (including demonstration examples) prepended to the input to be pp and the final output to be yy. Denote an intermediate generation before generating the final yy as zz. Under this formulation, Standard Prompting and Chain-of-Thought (CoT) Prompting can be described as:

where pcotp_{cot} is the CoT prompt, e.g., "Solve the task step-by-step" and {z1,z2...,zn}\{z_{1},z_{2}...,z_{n}\} are the intermediate steps. In contrast, our proposed Solo Performance Prompting can be described as follows:

where the SPP prompt (psppp_{spp}) includes a high-level instruction and two carefully crafted demonstration examplesThe tasks we use in the demonstration examples do not overlap with the evaluation tasks. that showcase the expected task-solving procedure of SPP. We describe the design details of the prompt in §A.1. The corresponding intermediate generations (zz) of SPP are detailed below.

Given an input task, SPP first generates a list of participants with different personas. For example in Figure 2, the model identified a Jay Chou Fan persona to help answer "the last song in the second album by Jay Chou". We let the language model identify the personas dynamically instead of manually defining them. Given only two demonstration examples (detailed in §A), we observe that a state-of-the-art large language model, e.g., GPT-4 (OpenAI, 2023b), can identify accurate and meaningful personas for diverse tasks. We denote this part of intermediate generation as zpz_{p} in Equation 3.

Based on the brainstorming remarks, the AI Assistant persona generates an initial solution zs0z^{0}_{s}, then it consults each of the other participants for feedback {zfi}\{z^{i}_{f}\}. The participants are encouraged to critique the current generation and give revision suggestions. For example, the Jay Chou Fan persona checks whether the song "An Jing" ("Silence") is correctly included in the story. This process can be repeated for multiple times until every participant is satisfied with the current solution. In Equation 3, we denote the intermediate generations of the multi-turn dialogue as {zs0,zf1,...,zfm}j=1...n\{z^{0}_{s},z^{1}_{f},...,z^{m}_{f}\}_{j=1...n} where nn is the number of iterations before reaching the final answer. The final answer can be directly read out following user-specified output format.

In summary, SPP instructs an LLM to solve general tasks via multi-persona self-collaboration in a pure zero-shot manner. In contrast, as detailed in Table 1, previous prompting-based methods are either task-specific or require additional mechanism, e.g., searching (Yao et al., 2023), external tools Yao et al. (2022), memory component (Shinn et al., 2023), and fine-tuning (Xu et al., 2023).

Experiments

To explore the effectiveness of Solo Performance Prompting (SPP), we adopt an evaluation methodology similar to that of previous work Yao et al. (2023). We carefully design new tasks and select tasks from existing benchmarks Srivastava et al. (2022) that are challenging even for the most capable LLMs (OpenAI, 2023b). The evaluation aims to cover diverse types of tasks encompassing both knowledge-intensive and reasoning-intensive domains.

We invent the Trivia Creative Writing task (§3.1), which requires the model to internally acquire and integrate diverse information from various fields. We observe that even GPT-4 (OpenAI, 2023b) frequently exhibit hallucination and factuality errors in the Trivia Creative Writing task. We also propose the Codenames Collaborative task (§3.2), an extension of the Codenames task from the BigBench (Srivastava et al., 2022) that features a two-role collaboration setup. Codenames Collaborative demands creative reasoning across a broad range of related knowledge and challenges the model’s theory of mind skills. Lastly, we include a challenging pure-reasoning task, Logic Grid Puzzle (§3.3), from the BigBench (Srivastava et al., 2022) which necessitates complex multi-step reasoning.

Baselines.

We compare our approach with Standard Prompting, Chain-of-Thought (CoT) prompting methods (outlined in §2) and Self-Refine (Madaan et al., 2023). For CoT, a similar prompt design to Yao et al. (2023) is employed, where the model is prompted to generate a plan or a series of steps before producing the final output. For Self-Refine, we follow Madaan et al. (2023) to design feedback and refine prompts. We perform one self-refine iteration which requires three times more inferences than SPP. Full prompts for the methods can be found in Appendix A.2.

Models.

The default model we use is GPT-4 (OpenAI, 2023b). Detailed inference configurations, API versions, and full results can be found in Appendices C and F. In §3.4, we further investigate the prevalence of cognitive synergy in LLMs with different scales and capabilities, including GPT-3.5-turbo (OpenAI, 2023a) and Llama2-13b-chat (Touvron et al., 2023).

1 Trivia Creative Writing: A Knowledge-Intensive Task

As illustrated in Figure 3, Trivia Creative Writing asks a model to write a coherent story while incorporating the answers to NN trivia questions. Our preliminary experiments (Figure 10) show that a sufficiently large NN can effectively challenge GPT-4 to demonstrate factual knowledge across diverse domains. Thus, we mainly consider two evaluation settings, N=5N=5 and N=10N=10. We built a benchmark with 100 instances for each NN, covering a total of 1000 trivia questionsTo select difficult question instances that can pose challenges to GPT-4, we use a smaller open-source LLM, fastchat_t5_3b (Zheng et al., 2023), to obtain preliminary performance on the validation set, and then choose the failure cases as our question selection. extracted from the TriviaQA (Joshi et al., 2017) dataset. More details can be found in Appendix B.1.

Evaluation Metrics.

Evaluating GPT-4 level generation results can be challenging. Our preliminary experiments indicate that, even for humans, it is very difficult to identify which generation is better in terms of overall "quality" of the story from different prompting methods. Thus, instead of focusing on evaluating the coherence of the generation, which can be highly subjective, we employ an automatic metric which focuses on detecting factual hallucinations. As shown in Figure 3, we perform string matching with the ground truth target answers for each question on the output generation. For each question, a match to any of the answer aliases provided by the TriviaQA dataset is considered a correct mention. The metric score is computed as: # correct answer mentions# trivia questions\frac{\text{\# correct answer mentions}}{\text{\# trivia questions}}.

Results.

Table 2 presents the results of the Trivia Creative Writing task. The key observations are as follows: (1) Chain-of-Thought (CoT) does not outperform Standard prompting, indicating that CoT is ineffective in eliciting an LLM’s knowledge abilities. Qualitative examples in Figure 8 and 11 illustrate that although CoT generates reasonable plans for task resolution, the final generation still contains factual errors and hallucinations. (2) Self-Refine only brings marginal improvements over iterations. (3) SPP outperforms all baselines significantly. The improvement is more pronounced in the N=10N=10 setting compared to N=5N=5 (10% vs. 7%), suggesting that Solo Performance Prompting is particularly beneficial when the task requires incorporating knowledge from numerous domains.

2 Codenames Collaborative: A Knowledge+Reasoning Task

As illustrated in 4, Codenames Collaborative is a collaborative task that challenges a model’s knowledge, reasoning, and theory of mind abilities by assigning two player roles: the Spymaster and the Guesser. The Spymaster’s role is to provide a hint word related to the target words, excluding some other distractor words, while the Guesser’s role is to identify the target words based on the given hint and the full list of words. The same LLM (GPT-4 (OpenAI, 2023b)) is used for both roles sequentially, and a dataset with 50 instances is constructed based on BigBench’s (Srivastava et al., 2022) Codenames task data.

Evaluation Metrics.

The original Codenames task in the BigBench dataset has limitations due to its focus on the Guesser role and subjectivity in hint words. Our new task, Codenames Collaborative, resolves this by creating a self-contained evaluation setting that accurately measures the model’s capability without human annotation. As illustrated in Figure 4, we compute the overlapping ratio between the predicted words from the Guesser and the target words as the metric.

Results.

Table 2 shows the results on the Codenames Collaborative task. Similar to the Trivia Creative Writing task, we find that CoT does not bring positive gains compared with the Standard prompting. Interestingly, iterative self-refinement brings negative impact on this task, due to a high tendency changing the initial response even if it is already good. In contrast, SPP brings significant improvements (~5%), which indicates its effectiveness on collaborative tasks that require knowledge, reasoning, and theory of mind skills. Figure 12 provides further qualitative examples illustrating that SPP generates detailed and interpretable intermediate dialogues.

3 Logic Grid Puzzle: A Reasoning-Intensive Task

We utilize the Logic Grid Puzzle task from the Bigbench (Srivastava et al., 2022) dataset, which comprises 200 instances. Each instance describes a logic puzzle typically involving 2 to 5 houses, with each house inhabited by a person with specific characteristics, such as playing the piano. The objective is to answer questions about house numbers based on given clues, which requires multi-step reasoning and the selection of relevant information. An example input and output of the Logic Grid Puzzle task are illustrated in Figure 5. For evaluation metrics, we calculate the accuracy of the predicted house numbers by comparing them with the ground truth targets provided by the dataset.

Results.

Table 2 presents the results on Logic Grid Puzzle. In contrast to the previous two tasks, we find that CoT brings significant improvements compared to Standard prompting, verifying the observation from previous work that CoT elicits better reasoning abilities. Furthermore, we discover that SPP also achieves strong performance on this reasoning-intensive task.

4 The Emergence of Cognitive Synergy

We further discover that cognitive synergy can only be fully unleashed in LLMs with a certain level of instruction-following capabilities, akin to that of GPT-4. This can be intriguingly compared to human development, where children usually begin to participate in role-playing around the ages of 2 to 3 Piaget (1954), but not before that age.

As shown in Figure 6, the effectiveness of SPP is not seen in smaller and less capable models like GPT-3.5 and Llama2. Additionally, on Llama2, we identify a unique problem which we refer to as early-termination, where the model stops generating after identifying the participants, resulting in exceptionally low performance with SPP. The model behaves as if it were waiting for input from a user instead of following the demonstration examples to generate responses on its own. Detailed discussions and examples on the early-termination problem can be found in Appendix E.

Analysis

As demonstrated by the results in §3, Solo Performance Prompting (SPP) not only brings significant improvements to knowledge-intensive tasks such as Trivia Creative Writing and Codenames Collaborative without relying on external knowledge bases, but also achieves strong performance on reasoning-intensive tasks like Logic Grid Puzzle. To our knowledge, SPP is the first zero-shot prompting method that can enhance both knowledge and reasoning abilities on GPT-4.

LLMs can effectively identify useful personas in a zero-shot manner.

We are interested in investigating whether the identified personas are highly relevant to the tasks. We visualize the personas automatically identified by SPP using a word cloud for each task in Figure 7(a), where a larger font indicates a higher frequency. The key observations include: (1) The identified personas are closely correlated with the particular task. For example, in Logic Grid Puzzle, even though "logic puzzle" is not mentioned in the input, the LLM frequently identifies the persona "Logic Puzzle Expert." (2) On knowledge-intensive tasks, such as Trivia Creative Writing, SPP identifies more diverse and specific personas, while on reasoning-intensive tasks, such as Logic Grid Puzzle, the personas are more homogeneous.

We further investigate whether a detailed profile for each persona is needed for eliciting domain knowledge, as suggested by Xu et al. (2023). To this end, we design a variant of SPP, SPP-Profile, which involves generating profiles for each persona during the Persona Identification phase. The results in Figure 7(b) show that SPP-Profile does not outperform SPP. This suggests that a fine-grained persona name without a detailed description may already be sufficient for eliciting certain domain knowledge.

Dynamic personas v.s. fixed personas.

To further investigate the importance of dynamically identifying personas for each task instance instead of fixing a general persona, an ablated variant of SPP, SPP-Fixed-Persona, is introduced. For SPP-Fixed-Persona, we modify the prompt (Figure 17) to force the personas to be fixed as an "AI Assistant" and an "Expert". Comparing SPP and SPP-Fixed-Persona in Figure 7(b), we have the following insights: (1) SPP consistently outperforms SPP-Fixed-Persona across all tasks, suggesting that dynamic, fine-grained personas are more effective than fixed, general personas. Qualitative examples in Figure 8 and 13 shows that the fine-grained personas such as "Film Expert" and "Sports Enthusiast" correctly provide the answers, while the fixed persona "Expert" fails. (2) SPP-Fixed-Persona also suffers from the early-termination problem as defined in §3.4, where the LLM stops collaboration before providing the final answer as if it were waiting for external inputs.

Impact of the demonstrations in SPP prompt.

To investigate the effectiveness of the hand-crafted demonstration examples in SPP, we conduct an ablation study where we remove the second demo example and preserve the first one, which shows only a two-persona collaboration setting. As shown in Figure 9, we observe that (1) Adding the second example, which requires collaboration of more than two personas, effectively boosts the performance. (2) SPP is fairly robust to the prompt change and show good performance with only the first demo example.

Related Work

Recent research (Deshpande et al., 2023; Xu et al., 2023; Fu et al., 2023; aut, 2023; Li et al., 2023) demonstrates that assigning personas or roles to LLMs influences their generation behavior. AI societies with distinct personas or occupations have been explored for collaboration (Park et al., 2023; Schick et al., 2022; Li et al., 2023; Cai et al., 2023). However, limitations in persona assignment and multi-agent collaboration include single or fixed persona assignments (Xu et al., 2023; Fu et al., 2023; Schick et al., 2022; Li et al., 2023) and the need for multiple LLM instances, increasing inference cost. In contrast, SPP uses a single LLM to dynamically identify useful personas for general tasks. Our discovery on the emergent nature of cognitive synergy also aligns with related work (Olausson et al., 2023), which investigates the emergent ability of self-debugging in code generation.

Enhancing reasoning and factual knowledge in LLMs.

LLMs face challenges in complex knowledge-intensive tasks due to hallucination (Maynez et al., 2020) and reasoning-intensive tasks due to the lack of human-like slow thinking (Sloman, 1996; Kahneman, 2011). Approaches like Chain-of-Thought (CoT) and Self-Refinement encourage LLMs to solve tasks step by step or iteratively revise their answers (Wei et al., 2023; Kojima et al., 2022; Zhang et al., 2022; Fu et al., 2022; Xue et al., 2023; Yao et al., 2023; Madaan et al., 2023; Shinn et al., 2023; Gou et al., 2023; Chen et al., 2023; Huang et al., 2022; Yao et al., 2022). However, these methods do not necessarily reduce factual hallucination. Retrieval augmented LLMs (Borgeaud et al., 2022; Izacard et al., 2022; Wang et al., 2022; Shuster et al., 2021) enhance knowledge acquisition but do not improve reasoning abilities. We propose Solo Performance Prompting (SPP) to elicit both knowledge and reasoning abilities in LLMs, improving factuality while maintaining strong performance on pure-reasoning tasks.

Conclusion

Solo Performance Prompting unleashes the cognitive synergy abilities within powerful LLMs, significantly reducing factual hallucination while enhancing reasoning. The performance is assessed using newly proposed tasks, e.g., Trivia Creative Writing and Codenames Collaborative, demonstrating superior results compared to Standard, CoT and Self-Refine. The discovery of the emergent nature of cognitive synergy on different LLMs draws interesting analogy to human development.

Limitations

Although Solo Performance Prompting exhibits promising improvements in acquiring factually correct knowledge compared to Standard prompting, it has some limitations. For instance, even when a fine-grained persona is assigned, the answer may still be incorrect. It remains unclear to what extent assigning a persona can help enhance domain knowledge in a specific area. Dedicated diagnostic experiments and theoretical efforts are needed to quantify the impact of having a persona or not.

Furthermore, we currently adopt an identical SPP prompt with the same two demonstration examples for any given task inputs, which may be suboptimal. Future work investigating how to find better demonstration examples conditioned on each input could further improve the effectiveness of SPP.

Last but not least, if given sufficient computational budget, a natural variant of SPP could extend to a multi-agent cognitive synergist setup where a leader persona identifies several expert agents and forms a cabinet to collaboratively solve a task. The multi-agent setup allows for leveraging richer computation power, larger local memory, and more flexible human-computer interaction, which could be essential for deploying to real-world applications.

References

Appendix A Prompts

To prompt an LLM to behave as a cognitive synergist that follows the expected task-solving procedure as mentioned in §2, we carefully designed the structure of the SPP prompt as follows. The full prompts can be found in § A.2.We use the same prompt for any arbitrary tasks.

The first part of the prompt contains a high-level instruction: "When faced with a task, begin by identifying the participants who will contribute to solving the task. Then, initiate a multi-turn collaboration process until a final solution is reached. The participants will give critical comments and detailed suggestions whenever necessary."

Demonstration Examples.

Then, we include two manually crafted demonstration examples to showcase the expected task-solving behavior. The first example describes a Game of 24 task, where we only include two personas: an AI Assistant and a Math Expert. This task aims to provide an example of a reasoning-intensive task, where the AI Assistant needs to propose multiple proposals, and the other participants need to give fine-grained feedback on where the current solution went wrong and how to improve it. The second example describes a poem-writing task with diverse requirements, including lexical constraints, semantic constraints, and audience awareness. This task aims to provide an example of a knowledge-intensive task, where diverse personas are required to collaboratively solve the task. This example also demonstrates a case where it is important to assign a dedicated persona to the audience, e.g., a ten-year-old child.

Task Prefix.

The last part of the prompt reminds the model to "identify the participants and collaboratively solve the following task step by step." followed by task-specific format instructions and inputs.

A.2 Full Prompts

Figures 15, 16 and 17 show the full prompts for SPP, SPP-Profile and SPP-Fixed-Persona respectively. Figure 18 shows the prompts for Chain-of-Thought (CoT) prompting. Figure 19 shows the prompts for Self-Refine prompting.

Appendix B Task Details

Figure 3 shows a detailed illustration of the Trivia Creative Writing task. Additionally, we investigate how the number of the questions (N) and the ordering of the questions would affect the performance on the Trivia Creative Writing task. As shown in Figure 10, with a larger number of questions (N≥\geq5), Trivia Creative Writing effectively challenges GPT-4’s performance. While a single question (N=1) yields similar outcomes regardless of the prompting method, SPP approach is notably superior for larger Ns. The ordering of the questions has minimal impact to the task performance.

The topic list is automatically generated by prompting GPT-4 to provide 100 nouns from pop cultureThe full prompt for generating the topic list can be found in Figure 20. We performed further human curation to avoid potential harmful content..

Appendix C Inference Configurations

The main results in Table 2 are obtained from GPT-4. The GPT-4 API version we employ is Azure 2023-3-15-preview.There are rare cases when a generation triggers the content filter of the API. We exclude those instances from our results. The temperature is set to 0.00.0 (most conservative) and top_p to 1.01.0 for all generations to maximize reproducibility. Since even though the temperature is set to 0.00.0 the GPT-4 generation can still be non-deterministic, we conduct additional experiment to investigate its generation consistency under this configuration. As shown in Table 3, we perform three individual runs and compute the mean and standard deviation of the metric score on Trivia Creative Writing. We find that the variance is sufficiently small and Solo Performance Prompting enjoys lower variance than Standard and CoT prompting.

To evaluate the potential impact of initial persona assignment through a system message, we consider two inference settings: with or without the default system message, "You are an AI assistant that helps people find information". Divergent patterns are observed across various tasks and methods regarding the use of the system message. We report the average metric scores across both inference settings in Table 2. Full GPT-4 results for each setting can be found in Appendix F.

For GPT-3.5 results in Figure 6, we employ the same prompt, hyper-parameters and the best system message setting in terms of SPP’s GPT-4 performance. For Llama2, we leverage the Huggingface text-generation pipelinehttps://huggingface.co/blog/llama2 with greedy decoding.

Appendix D Additional Qualitative Analysis

Figure 11 presents examples of the Trivia Creative Writing task, illustrating that although CoT can generate plausible plans for task resolution, the final outcomes often contain factual inaccuracies and instances of hallucination. In contrast, SPP elicits precise knowledge with fine-grained personas.

Figure 12 displays examples of the Codenames Collaborative task, illustrating that SPP generates intermediate dialogues that are both detailed and interpretable, leading to superior performance compared to CoT.

Figure 13 shows additional qualitative examples on Solo Performance Prompting vs SPP-Profile.

Appendix E Early-termination with SPP-Fixed-Persona

Figure 14 shows an example of the early-termination problem (defined in § 4) where the generation stops before reaching the final solution as if the models is waiting input from an external user.

The problem is particularly severe on certain tasks, e.g., Codenames Collaborative, resulting in unexpectedly low performance as shown in Figure 7(b). The problem can be largely alleviated by removing the system message but cannot be entirely eliminated. Table 4 shows the statistics of the early-termination problem for each task and method. In contrast, we did not observe early-termination on SPP, SPP-Profile, Standard, or CoT prompting with GPT-4.

Appendix F Full Results

Full results of the three tasks: Trivia Creative Writing, Codenames Collaborative and Logic Grid Puzzle can be found in Tables 5, 6 and 7, respectively.

Appendix G Usage of AI assistants in writing

We used ChatGPT and GPT-4 solely for checking and correcting grammars.