Lemur: Harmonizing Natural Language and Code for Language Agents
Yiheng Xu, Hongjin Su, Chen Xing, Boyu Mi, Qian Liu, Weijia Shi, Binyuan Hui, Fan Zhou, Yitao Liu, Tianbao Xie, Zhoujun Cheng, Siheng Zhao, Lingpeng Kong, Bailin Wang, Caiming Xiong, Tao Yu
Introduction
Intelligent agents are broadly conceptualized as autonomous problem solvers with the abilities to sense their environment, decide, and act upon that enviorment (Sutton & Barto, 2005; Russell, 2010; Wilkins, 2014). Recent implementations of this concept in creating language agents (Yao et al., 2022b; Gravitas, 2023; Wang et al., 2023a) capable of utilizing natural language for varied and intricate tasks in diverse environments have demonstrated potential, particularly when built upon large language models (LLMs) (Brown et al., 2020; Chen et al., 2021; Chowdhery et al., 2022; OpenAI, 2023; Touvron et al., 2023). Such agents harness the human knowledge in LLMs and can think and communicate in human terms. This equips them to employ varied tools, operate in complex environments, engage in language reasoning, and create spontaneous multi-agent systems.
To effectively form the foundation of language agents, LLMs should not only master human interaction, reasoning, and planning but also ensure grounding in the relevant environments (Wei et al., 2022b; Huang et al., 2022a; Ichter et al., 2022). Human interaction, reasoning, and planning can be largely realized through the natural language capabilities of LLMs. On the other hand, the grounded execution in the environment is usually achieved by using general-purpose code or domain-specific APIs, such as controlling web browsers (Shi et al., 2017; Yao et al., 2022a; Deng et al., 2023; Zhou et al., 2023), interacting with OS CLI terminals (Yang et al., 2023), and manipulating robotic arms via primitive APIs (Ichter et al., 2022; Huang et al., 2022b; Liang et al., 2023). Therefore, we posit that for the construction of language agents, it is imperative for language models to possess harmonized capabilities in both natural language and programming languages. This balance ensures that models do not specialize exclusively in certain areas but can seamlessly integrate with environment contexts and generate controllable and valid actions. Presently, closed-source models like GPT-4 demonstrate such capabilities, which empower them to function as language agents. However, current open-source LLMs such as Llama 2 (Touvron et al., 2023) and CodeLlama (Rozière et al., 2023) have traditionally been tailored for either textual or code-related tasks, with limited ability to effectively balance both.
To address this need, we introduce Lemur and Lemur-Chat, cutting-edge, openly accessible models meticulously pre-trained and fine-tuned to harmonize text and code capabilities. We enhanced the base Llama-2-70B through thoughtfully designed pre-training and instruction fine-tuning stages. Specifically, we built a code-centric corpus based on The Stack (Kocetkov et al., 2022), comprising 90 billion tokens with a 10:1 text-to-code ratio, ensuring improved capabilities in coding ability while maintaining performance in natural language ability. We refer to this model as Lemur. After pretraining, we conducted instruction fine-tuning using about 300K examples from both text and code to build an instruction-following model, which we refer to as Lemur-Chat. Thorough assessments across 8 textual and coding benchmarks validate the superior performance of both Lemur and Lemur-Chat in multiple text and code evaluations, establishing them as the most well-rounded open-source models.
Moreover, this work embarks on assessing the vital capabilities of language agents across various scenarios, which we refer to as agent benchmarks. We place a particular emphasis on their tool-usage abilities, and abilities in grounding in the environment feedback and human feedback. We also explore the challenges posed by real and partially observable environments where the agent has to take actions based on limited knowledge and take additional actions to gather more information. Experimental results indicate that Lemur-Chat outperforms other open-sourced models in 12 of the 13 agent benchmarks. This underscores how the integration of natural and coding capabilities allows Lemur-Chat to exceed the current open-source models for language agents, markedly bridging the performance disparity between open-source and commercial alternatives. Our experiments show that in agent scenarios, there is a need for synergy between natural language and coding abilities. For models with strong natural language capabilities but weak coding abilities like Llama-2-70B-Chat, they can effectively use simple tools to assist reasoning (§4.2) because the action space is small and the difficulty of using tools is low. However, when facing complex decision-making scenarios such as web browsing and house navigation, the action space is usually large, and models with strong coding abilities have an advantage in generating complex executable action sequences (§4.5). Overall, Lemur has both strong natural language and coding abilities, enabling it to achieve better performance in both scenarios. This research provides insights into optimizing the synergy between natural and programming languages, laying a solid foundation for developing advanced language agents capable of operating efficiently in various environments. Our contributions can be summarized as follows:
Harmonized Language Models: We introduce Lemur, state-of-the-art open-source models that are meticulously pre-trained and instruction fine-tuned to demonstrate balanced language and coding capabilities.
Comprehensive Evaluations: We follow and extend extensive benchmarks in both classical and agent perspectives, demonstrate the advantage Lemur over other open-source models, and showcase its potential as a versatile language agent.
Insights on Synergy: Our research delves deep into the interplay between natural and programming languages, offering valuable insights that pave the way for the development of more advanced language agents adept at operating across a range of environments.
Pre-training and Instruction tuning of Lemur
This section will introduce the method to build the Lemur and Lemur-Chat models and their performance on commonly-used benchmarks for pre-trained language model evaluation. To build a more balanced model, the pipeline includes two stages: pre-training (§ 2.1) and instruction fine-tuning (§ 2.2). The training pipeline is shown in Figure 2.
In the pre-training stage, we choose the Llama-2-70B base model as the starting point, which is the cutting-edge open-sourced base model in most scenarios except the coding scenario. Our goal is to improve the coding ability while maintaining the reasoning ability of Llama-2. To this end, we built a corpus with a code-to-text ratio of 10:1. For the code part, we base it on The Stack (Kocetkov et al., 2022), a collection of source codes from GitHub with permissive licenses. Among all languages, we focus on scripting or interpreted languages (Python, SQL, Bash, Perl, etc.) because agent models are often executed in interactive scenarios. As for the text aspect, we use RefinedWeb (Penedo et al., 2023), Redpajama (Computer, 2023), as well as CommonCrawl, Wikipedia, Books, ArXiv, StackExchange and DM Mathematics (Saxton et al., 2019) to build the textual data corpus. Following previous works (Gao et al., 2021; Smith et al., 2022; Computer, 2023; Li et al., 2023), we do extensive deduplication after aggregating all data sources. The composition of the data is shown in Appendix § A.1.1. We train the Lemur-70B model initialized with Llama-2-70B using a TPUv4-512 pod. We train the model with sequence packing (Raffel et al., 2020; Chung et al., 2022) to improve the training efficiency. Please refer to Appendix § A.1.2 for more details.
2 Instruction Fine-tuning
During the instruction fine-tuning phase, we include four data sources to construct Lemur-Chat, including the Open Assistant crowdsourced annotated assistant-style dialogue corpus (Köpf et al., 2023), Orca data with chain of thought reasoning for human-written tasks (Lian et al., 2023; Mukherjee et al., 2023), ShareGPT & Chatlogs containing real user and ChatGPT chat history records (ShareGPT data, 2023), as well as Evol-CodeAlpaca data (Luo et al., 2023) consisting of complex coding tasks generated by ChatGPT along with their corresponding solutions. The statistics of these instruction datasets are shown in Table 1. After we collect these data, we additionally clean and deduplicate these instruction fine-tuning data. We conduct training on these data for 2 epochs. Please refer to § A.2 for more details.
From Language Model to Language Agent
This section introduces how we measure the language and coding abilities of a language model to guide the process of coordinating language and code. We further discuss the new challenges faced when connecting LLM to the environment and describe how we examine the necessary agent capabilities.
Previous work typically uses a variety of benchmarks to comprehensively reflect the performance of models on different types of tasks. Therefore, we evaluate the performance of various models across text and code benchmarks as follows.
Natural Language benchmarks: MMLU (Hendrycks et al., 2021a) to determine factuality, BBH (Suzgun et al., 2022) to check reasoning abilities, GSM8K (Cobbe et al., 2021) to gauge math reasoning.
Code benchmarks: HumanEval (Chen et al., 2021) and MBPP (Austin et al., 2021) to test Python writing abilities, Spider (Yu et al., 2018) to assess database query capabilities, MultiPL-E (Cassano et al., 2023a) to measure the multi-lingual coding capabilities, DS-1000 (Lai et al., 2023) for evaluation in data science scenarios
2 Connecting LLM Agents to Environment
While the measurements in §3.1 provide valuable insights on each domain, they may not fully encapsulate the models’ capabilities in language agent scenarios due to their focus on single-turn interactions in fully observable settings. In order to address this discrepancy, we scrutinize models in the context of multi-turn interactive scenarios to gauge their adaptability in practical situations. Our assessment is centered on various capabilities of language agents, which are encapsulated in the factors outlined in Figure 3 and Table 2. For a more comprehensive assessment of these capabilities, we reorganize existing multiple datasets into four sets to examine the diverse skills demanded by the agent scenario. Please refer to Appendix B.2 for details of each dataset.
Augment with Tools Tool usage (Schick et al., 2023; Hao et al., 2023; Mialon et al., 2023) is an important capability of language agents. For example, tasks like precise calculations or information retrieval can be offloaded to external modules (such as Python interpreters or search engines) using tools (Gao et al., 2023; Chen et al., 2022), thereby improving the reliability and interpretability of results. Generally speaking, practical tools involve calling operators in symbolic languages or directly invoking APIs using programming languages (Shen et al., 2023). This requires the model to decompose a complex task and ground its intention to external tools, obtaining the returned results into context for further actions based on them. This process relies on both natural language reasoning abilities and the use of programming languages (Cheng et al., 2022; Surís et al., 2023). To assess the ability of language agents to solve complex multi-turn problems using tools, we introduce the part of the MINT dataset (Wang et al., 2023b) which focuses on tool-utilization for reasoning. This part includes several adapted datasets for testing: MINT-{GSM8K, MATH} (Cobbe et al., 2021; Hendrycks et al., 2021b) is used to test the model’s ability to solve mathematical problems using a Python interpreter, and MINT-{HotpotQA, MMLU} (Yang et al., 2018; Hendrycks et al., 2021a) assesses the model’s capability to solve knowledge-based questions using Wikipedia searches. At the same time, for the MINT-TheoremQA (Chen et al., 2023a), the model needs to perform knowledge searches on Wikipedia and use a Python interpreter to calculate and draw conclusions.
Self-debug with Environment Feedback Self-debug is an important way to test whether a model can incorporate environmental feedback through multi-turn interactions (Jignasu et al., 2023; Olausson et al., 2023; Chen et al., 2023b). In the self-debug scenario, the model usually needs to complete a complex operation sequence, such as Python functions, database queries/modifications, robot action sequences, etc (Gur et al., 2023; Yao et al., 2023). These complex operations sometimes cannot be executed successfully and will return an error message. This requires the model to comprehend this kind of environmental feedback and engage in interactive error correction until it is correct, which tests the joint effect of natural language reasoning ability and programming languages. We use rich datasets from multiple benchmark tests to evaluate this performance, including multi-turn MINT-MBPP and MINT-HumanEval in MINT benchmark tests (Wang et al., 2023b), SQL and Bash in InterCode (Yang et al., 2023), as well as RoboCodeGen that calls robot APIs through code in Robotics scenarios (Liang et al., 2023). These environments require the model to complete a complex task and will provide error messages in case of execution errors. In these environments, the performance of the model will vary according to its self-debugging ability, reflecting the model’s ability to incorporate feedback. According to the original settings, we evaluate MINT-HumanEval, MBPP in a 5-round setting and Intercode-SQL, Bash in a 10-round setting. We evaluate RobotCodeGen in a 5-round setting.
Follow Natural Language Feedback Following natural language feedback is an important mechanism for agents to receive information from humans or other agents (Wang et al., 2022a; Ouyang et al., 2022; Gong et al., 2023). In scenarios where complex problems are solved through multi-turn interactions, the model not only needs to incorporate environmental feedback but also feedback from humans or other agents in order to improve. This mechanism requires the model to understand new instructions based on a context that combines natural language and code, and ground them into new action sequences. To evaluate the model’s ability to accept natural language feedback, we follow the approach of the MINT benchmarks: using a GPT-4 simulated user as a teacher to guide the model in problem-solving. This setup includes a series of MINT datasets (Wang et al., 2023b) to comprehensively evaluate performance after adding natural language feedback in various scenarios.
Explore in Partially Observable Environments Exploring partially observable environments is a unique and challenging factor in agent scenarios. All the settings mentioned earlier can be considered as fully observable environments, which means that agents can observe all the information of the current environment to plan, reason, and make decisions. However, in partially observable environments (also known as Partially Observable Markov Decision Process) (Kurniawati, 2021), agents can only partially observe the environmental information required to solve problems. This requires agents to collect information through exploration and continue making decisions. This process places high demands on various abilities of agents, such as natural language planning and reasoning, environmental interaction, etc., and is very close to real-world scenarios. To measure this ability, we use three datasets: InterCode-CTF (Yang et al., 2023) and WebArena (Zhou et al., 2023) in digital environments, as well as ALFWorld (Shridhar et al., 2021) in physical environments. InterCode-CTF provides a complete OS terminal for models to solve practical Catch the Flag (CTF) problems where agents need multiple rounds of exploration to obtain the flag. WebArena evaluates agents’ ability to control browsers for task completion through exploration. ALFWorld is a simulated home environment where agents need to explore navigation and complete specific tasks.
Experimental Results
The comprehensive evaluations of text and code benchmarks in Table 3 demonstrate the impressive capabilities of Lemur-70B and Lemur-70B-Chat models. Deviating from Llama-2-70B and Llama-2-70B-Chat, which are mostly pre-trained and fine-tuned on text data, the Lemur models augment coding abilities and thereby enhance the overall performance by 4.3% and 14.8% respectively. Alternatively, models like CodeLlama-34B and CodeLlama-34B-INST, which are primarily trained on code datasets, exhibit solid performance in code benchmarks. Importantly, the commendable increase in the overall performance of Lemur-70B and Lemur-70B-Chat models as compared to the CodeLlama models, by approximately 1.9% and 9.4% respectively, highlights the virtue of harmonizing textual and coding skills.
The synergic text and code abilities enable them to function as language agents. However, a disparity exists between natural language and coding capabilities in current open-source models. Such limitations obstruct these models’ abilities to act as language agents, leading to performance degradation in agent benchmarks. In subsequent sections, we meticulously evaluate various critical capabilities of agents, revealing the importance of synergic text and code abilities to language models.
2 Augment with Tools
In the realm of problem-solving, agents, akin to humans, employ various tools to augment their capabilities. This is exemplified by (Chen et al., 2022), who showcase that the mathematical reasoning prowess of LLM can be significantly enhanced with the aid of Python. As per the data presented in Table 4, it is evident that Lemur-70B-Chat outperforms both Llama-2-70B-Chat and CodeLlama-34B-INST, indicating its superior ability to effectively leverage tools.
3 Self-debug with Environment Feedback
The technique of self-debug gains considerable traction in the realm of code generation (Olausson et al., 2023; Zhang et al., 2023). This method consists of models using feedback information, such as interpreter error tracebacks and database observations, to rectify any existing errors. This adaptive capacity is essential for agents because they have to constantly receive and react to feedback from the environment during their interaction process.
As demonstrated in Table 5, the performance of the Lemur-70B-Chat significantly surpasses that of the Llama-2-70B-Chat and CodeLlama-34B-INST in interactive coding benchmarks. This underscores the importance of having balanced capabilities for interactive agents in such environments.
We further analyze the results with InterCode-SQL to understand how models incorporate environment feedback. In this setting, agents are provided with database schema and instruction guidelines for SQL game interactions. Acting as agents, the models are tasked with querying databases and responding to questions in a multi-turn interaction environment. Figure 4 shows the Growth in Success Rate across models with each interaction turn. Lemur demonstrates robust performance in the first round, showing an initial performance that is comparable to that of text-bison-001 and slightly lower than gpt-3.5-turbo. Nevertheless, Lemur improves consistently across ten interactive rounds, surpassing the performance of gpt-3.5-turbo eventually. In contrast, text-bison-001, which exhibits comparable initial performance with Lemur, does not show significant improvement. Llama-2-70B-Chat, while displaying consistent adaptability to feedback throughout the process, has a significant gap in initial performance due to inferior coding ability, hence its success rate remains relatively lower. CodeLlama-34B-INST hardly answers the questions correctly in the first round. We observe that this is because it blindly follows the advice in the game guide and stubbornly executes the show tables command first, instead of trying to understand the database structure that had already been provided. After an improvement in the second round, its performance returns to normal. However, the growth remained limited even after ten rounds of interaction, settling at par with Llama-2-70B-Chat and reflecting its relative weakness in adapting to environmental feedback.
4 Follow Natural Language Feedback
Following natural language feedback from users or agents in the environment is an important ability for language agents. It requires language agents to understand complex natural language instructions and convert the instructions into symbolic executable sequences based on the current contexts. To evaluate the models’ ability to follow natural language feedback, we follow the evaluation settings of MINT (Wang et al., 2023b), which measures language agents’ ability to leverage natural language feedback using performance improvement. To provide natural language feedback in multi-turn settings, MINT uses GPT-4 to simulate a user providing helpful feedback on the solutions from evaluated language agents.
We evaluate models on MINT-Reasoning and MINT-Code with and without GPT-4 feedback and the experimental results are in Table 6. MINT-Reasoning includes five modified benchmarks: GSM8k, MATH, TheoremQA, HotpotQA, and MMLU. MINT-Coding includes modified HumanEval and MBPP. We find that all models can benefit from GPT-4, which means powerful GPT-4 as a teacher can provide helpful feedback even without ground truth. We also calculate , which indicates the absolute improvement thanks to GPT-4 feedback. According to the results from Table 6, Lemur-70B-Chat model obtain 8.19 in , which is significantly better than Llama-2-70B-Chat and CodeLlama-34B-INST.
5 Explore in Partially Observable Environments
We evaluate language agents in varied environments and find that Lemur-70B-Chat exhibits balanced and commendable performance across all tested tasks. It scored in the complex InterCode-CTF, showcasing its proficiency in multifaceted cybersecurity skills and strategic adaptability. In WebArena, it achieved a score of , reflecting its adeptness in interpreting and executing advanced natural language commands in intricate web environments. In ALFWorld, it demonstrated superior planning and commonsense reasoning with a score of , successfully navigating and performing in simulated physical environments. While gpt-4 overall exhibits higher scores, Lemur-70B-Chat’s consistent performance across diverse and partially observable environments underscores its versatility and potential in handling real-world, multifarious applications.
We also conduct experiments to explore different output formats for actions in the WebArena environment. Instead of merely prompting the language model to directly produce predefined actions (e.g. type [id] [content] [press_enter_after]), we deterministically mapp the action space to Python representations (e.g. type(id:int, content:str, press_enter_after:bool)). We then prompt the language model to predict this representation and then parse the model’s Python prediction back into the predefined executable actions. Our findings, as shown in Figure 5 indicate that mapping to a Python representation leads to better performance than directly predicting actions. Such results suggest that carefully and rationally selecting intermmediate representation, based on the pre-training corpus, can effectively enhance the model’s performance as a language agent, aligning with prior findings (Hu et al., 2022).
Related Work
Transfer Learning on Code The advancement in expansive language models (Devlin et al., 2019; Radford et al., 2019; Raffel et al., 2020; Brown et al., 2020) has catalyzed progress in transfer learning for code-related tasks, further enriched by the continuous pre-training paradigm (Gururangan et al., 2020). Hernandez et al. (2021) offered insights into the interplay between model size, training data, and performance in transferring language ability to code tasks. Several models have been introduced, exhibiting enhanced program synthesis and infilling/completion performance, by undergoing continual pre-training on extensive code data (Feng et al., 2020; Wang et al., 2021; Chen et al., 2021; Rozière et al., 2023). However, their intense focus on code often results in a compromise on natural language capabilities. Lemur addresses this by moderately augmenting a large model with a balanced mixture of code and natural language data, maintaining proficiency in both domains.
Instruction Fine-tuning The process of aligning large language models to follow instructions, often referred to as instruction tuning, has been primarily directed towards NLP tasks (Wei et al., 2022a; Wang et al., 2022b). Recent studies have sought to broaden the use cases of instruction tuning to involve a wider variety of general tasks (Ouyang et al., 2022). Self-instruct method generates instructions using seed instructions, aiding the understanding of how to adapt language models by fine-tuning them on instructions garnered from ChatGPT (Wang et al., 2022a; Zheng et al., 2023; Xu et al., 2023b; Mukherjee et al., 2023). In this study, we adopt a similar approach in tuning the Lemur base model to follow instructions.
Language Agents Language agents are adept at following user instructions and engaging with environments to execute tasks. Recent trends in research and open-source communities have employed Large Language Models (LLMs) (Brown et al., 2020; Chen et al., 2021; Chowdhery et al., 2022; OpenAI, 2023) as the principal controllers for these agents (Yao et al., 2022b; Chase, 2022; Gravitas, 2023; Shinn et al., 2023; Wang et al., 2023a; Xu et al., 2023a; Lin et al., 2023; Yao et al., 2023). This is driven by the LLMs’ demonstrated abilities in reasoning, planning, grounding, and code generation (Wei et al., 2022b; Huang et al., 2022a; Ichter et al., 2022; Xie et al., 2023), crucial for comprehending user instructions, grasping the environmental context, and generating executable actions. Lemur seamlessly integrates capabilities in both text and code, enabling the generation of environment-grounded, executable actions essential for constructing language agents.
Agent Evaluation Benchmarks The rapid evolution of language agents requires an accurate and comprehensive assessment of agents. Recent works pose new dimensions that put language agents in web environments (Deng et al., 2023; Yao et al., 2022a; Zhou et al., 2023), interactive code environments (Yang et al., 2023), digital game (Fan et al., 2022), and household (Puig et al., 2018; Shridhar et al., 2020; 2021) to finish certain task under natural language instruction. Apart from collecting new tasks, these agent benchmarks can be established by re-purposing and transforming existing datasets to agent datasets (Wang et al., 2023b; Yang et al., 2023; Liu et al., 2023). Our research combines these tasks, conducts extensive evaluation, and evaluates the model’s capability to construct a language agent from multiple dimensions.
Conclusion
In conclusion, this research underscores the pivotal role of harmonizing natural and programming language proficiencies in the evolution of language models to sophisticated language agents. Through the development of Lemur and Lemur-Chat, we demonstrated that the meticulous amalgamation of these competencies allows for elevated performance in diverse environments and applications, narrowing the existent capability divide between open-source and proprietary models. We open-sourced both models with the intention of fostering further research in the field of language models for agents.
Acknowledgments
This project is an open research collaboration between XLang Lab of HKU and Salesforce Research. We extend our heartfelt gratitude for the generous research grants received from both Google Research and Amazon AWS. Our special thanks go to the Google TPU Research Cloud program for their provision of vital computational resources. Our research greatly benefitted from insightful discussions with esteemed collaborators: Pengcheng Yin, Yaqing Wang, and You Wu from Google Research; Peng Shi and Shuaichen Chang from AWS AI; Yizhong Wang from the University of Washington; Xinyang Geng from UC Berkeley; Ansong Ni from Yale University; Chen Henry Wu from Carnegie Mellon University; Zeyu Liu from the University of Texas at Austin; Lin Zheng and Jiacheng Ye from the University of Hong Kong.
References
Appendix A Lemur and Lemur-Chat
We performed pre-training on the top of Llama-2. The detailed statistics of the pre-training data corpus are presented below in Table 8.
A.1.2 Training Details
We train our model on a TPUv4-512 pod. Our codebase is based on Jax and EasyLM (Geng, 2023). Following the pretraining methodology of Llama 2 (Touvron et al., 2023), we used a batch size of 4M tokens. To improve training efficiency, we packed multiple shorter sequences into each batch entry when possible, an approach known as sequential packing (Raffel et al., 2020).
Optimization was performed with Adam using a peak learning rate of 4e-5 along with = 0.9 and = 0.95. Gradients were clipped at 1.0 to prevent exploding gradients. A cosine decay schedule was used for the learning rate, with a linear warmup of 2000 steps followed by decaying the learning rate to 10.% of its peak value at the end of training.
A.2 Instruction Fine-tuning
The OpenOrca dataset (Lian et al., 2023) is a collection of augmented FLAN Collection data. Currently, it includes around 1M GPT-4 and 3.2M GPT-3.5 Chat completions. It is tabularized in alignment with the distributions presented in the ORCA paper (Mukherjee et al., 2023). We randomly sampled 200K samples from the GPT-4 Chat completions in Table 1
OpenAssistant(OASST1) (Köpf et al., 2023) is a crowdsourced human-annotated assistant-style conversation corpus comprising 161,443 messages, enriched with 461,292 quality ratings. (Köpf et al., 2023). The data results from the collaboration of over 13,500 volunteers worldwide.
We curated English human instructions from ShareGPT(https://sharegpt.com/) and Chatlogs(https://chatlogs.net/) for instruction-tuning. Specifically, we use datasets from Huggingface dataset hub https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered and https://huggingface.co/datasets/winglian/chatlogs-en-cleaned. To filter English instructions, we utilized the langdetect package in Python and eliminated any instructions with non-English detection results. Additionally, considering that non-English instructions often contain consecutive non-English characters, we implemented a second filtering step, removing all sentences with three or more consecutive non-English characters. To ensure semantic diversity, we employed instructor(Su et al., 2023) to encode the filtered instructions, calculate cosine similarity, and remove instructions with a similarity score greater than 0.95. After deduplication, we obtained nearly 80K instances, with an average of about 6 rounds of high-quality data per instance.
We use two open-sourced Evolution-Instruct (Luo et al., 2023) datasets, i.e., Evol-Instruct-Code-80k-v1 (2023) and evol-codealpaca-v1 (2023), and an execution-verified Python dataset constructed by us. After applying the same deduplication method as the Text data, we obtained 46K examples.
A.2.2 Instruction Fine-tuning Details
We use Huggingface Transformers (Wolf et al., 2019) with Accelerate Library to fine-tune the Lemur model to obtain the Lemur-Chat model. We train on our data collection for two epochs. We use Adam optimizer with a learning rate of 2e-5 and a batch size of 128.
Appendix B Evaluation
Following Wang et al. (2023c); Rozière et al. (2023), we evaluated the results outlined in Table 3 utilizing greedy decoding for generation. The evaluation scripts were crafted based on the framework available at https://github.com/allenai/open-instruct. For a robust evaluation of chat/instruct models on code function completion tasks such as HumanEval, MBPP, MultiPL-E, and DS-1000, we furnished the incomplete code context as prompts to ensure reliable function completions, aligning with the practices of Wang et al. (2023c).
MMLU: The MMLU (Massive Multitask Language Understanding) (Hendrycks et al., 2021a) is structured to gauge the knowledge garnered during pretraining, by evaluating models in zero-shot and few-shot settings across 57 diverse tasks from fields like elementary mathematics, US history, computer science, law, and clinical knowledge. Adhering to MMLU’s original setup, the evaluation was performed using 5-shot examples.
GSM8K: Originated by Cobbe et al. (2021), GSM8K encompasses 8.5K linguistically varied grade school math word problems, targeting the evaluation of question answering on basic math problems necessitating multi-step reasoning. The evaluation was executed on GSM8K’s test set, employing 8-shot in-context examples within a chain-of-thought prompting method (Wei et al., 2022b).
BBH: The BBH (BIG-Bench Hard) (Suzgun et al., 2022) is centered on tasks believed to challenge the existing language models, particularly those necessitating multi-step reasoning. Following the guidelines laid out in the original paper, the evaluation was carried out with 3-shot in-context examples, albeit without employing chain-of-thought prompting due to computational resource constraints.
HumanEval: Introduced by Chen et al. (2021), HumanEval comprises 164 meticulously crafted programming challenges, accompanied by unit tests to ascertain the feasibility of the solutions proposed, aiming to challenge code generation models. The evaluation on HumanEval dataset from Chen et al. (2021) was conducted employing official zero-shot docstring prompts.
MBPP: The MBPP (Mostly Basic Python Problems) dataset (Austin et al., 2021), encompassing approximately 1,000 crowd-sourced Python programming problems, is directed towards entry-level programmers, covering fundamental programming concepts and standard library functionalities. The evaluation on MBPP dataset from Austin et al. (2021) was executed using official 3-shot prompts.
Spider: Designed for intricate and cross-domain semantic parsing alongside text-to-SQL tasks, Spider (Yu et al., 2018) aims to foster the development of natural language interfaces for cross-domain databases. It houses 10,181 questions and 5,693 unique complex SQL queries spanning 200 databases across 138 distinct domains.
MultiPL-E: MultiPL-E (Cassano et al., 2023b) emerges as a multi-programming language benchmark for appraising the code generation capability of large language models (LLMs). It extends the HumanEval dataset to include 18 additional programming languages, and offers a system for transposing unit test-driven code generation benchmarks to new languages, thus standing as a noteworthy multilingual code generation benchmark. The evaluation was undertaken using the official zero-shot prompt.
DS-1000: DS-1000 (Lai et al., 2023) is a code generation benchmark with a thousand data science questions spanning seven Python libraries such as NumPy and Pandas. The dataset reflects diverse, realistic, and practical use cases, providing a reliable metric for evaluation while defending against memorization by perturbing questions. The problems in the dataset are collected from StackOverflow, making it a practical and realistic benchmark for data science code generation tasks. We evaluate with the official zero-shot prompt.
B.2 Agent Evaluation
We construct our agent abilities evaluation suite based on the following datasets:
MINT (Wang et al., 2023b) is a well-rounded evaluation that covers a range of tasks repurposed for multi-turn evaluation. It consists of three types of tasks, namely reasoning (MMLU (Hendrycks et al., 2021a), GSM8K (Cobbe et al., 2021), MATH (Hendrycks et al., 2021a), TheoremQA (Chen et al., 2023a), HotpotQA (Yang et al., 2018)), code generation (HumanEval (Chen et al., 2021), MBPP (Austin et al., 2021)), and decision-making (ALFWorld (Shridhar et al., 2021)). To assess the proficiency of language models in employing tools, their reasoning process is reoriented to incorporate tool use. For instance, language models are prompted to utilize the Python calculator to work out mathematical problems, rather than supplying the answer outright. In code generation, the LLM is encouraged to incorporate Python interpreter messages to check generated code. To prevent any misunderstanding, we use the prefix “M-” to differentiate the original dataset from the MINT version in our paper.
InterCode-Bash/SQL (Yang et al., 2023) are two tasks that serve as an experimental platform that assesses the capacity of extensive language models to integrate feedback from the environment during interactive coding. This benchmark evaluates models in the way of generating a series of actions under user instruction and regards elements such as execution outcomes, and error backtracking, amongst others as environment observations.
RoboCodeGen (Liang et al., 2023) serves as a specialized evaluation framework focused on robotics-related tasks. It comprises three main types of questions that target spatial reasoning (e.g., identifying the closest point among a set of points), geometric reasoning (e.g., verifying if one bounding box is contained within another), and controls (e.g., PD control).
InterCode-CTF is a task in the InterCode evaluation suite. Capture the Flag (CTF) is a competitive cybersecurity game that requires LLMs to discover encrypted “flag” hidden within code snippets or file systems. Compared with Bash and SQL generation, CTF is much more complex, requiring agents to have knowledge of multiple coding languages, modularize higher-order objectives into sub-problems, create multi-step plans for solving each problem, and adjust strategies when a plan fails to provide any useful insights.
WebArena (Zhou et al., 2023) creates self-hostable websites of four popular categories by simulating real-world equivalent functionalities and data. To simulate human problem-solving abilities, WebArena also embeds tools and knowledge resources as standalone websites.
ALFWorld (Shridhar et al., 2021) is a synthetic environment benchmark adapted from Alfred (Shridhar et al., 2020) in the text-based interface where agents need to navigate in simulated households (e.g., go to coffee table 1, pick up paper 2, use desk lamp 1) and achieve high-level goals (e.g., check the paper under the desk lamp). Task instances may involve over 50 locations and 50 steps to solve, thus challenging the agent’s ability to plan, track sub-goals, and conduct systematic exploration.