LLM360: Towards Fully Transparent Open-Source LLMs
Zhengzhong Liu, Aurick Qiao, Willie Neiswanger, Hongyi Wang, Bowen Tan, Tianhua Tao, Junbo Li, Yuqi Wang, Suqi Sun, Omkar Pangarkar, Richard Fan, Yi Gu, Victor Miller, Yonghao Zhuang, Guowei He, Haonan Li, Fajri Koto, Liping Tang, Nikhil Ranjan, Zhiqiang Shen, Xuguang Ren, Roberto Iriondo, Cun Mu, Zhiting Hu, Mark Schulze, Preslav Nakov, Tim Baldwin, Eric P. Xing
Introduction
The landscape of Large Language Models (LLMs) has experienced a remarkable transformation in the past one year, witnessing an unprecedented surge in both the popularity and capabilities of these models. At the forefront of this evolution are proprietary LLMs such as GPT-4 and Claude , which have captured the attention of the AI community due to their power and versatility. At the same time, the recent emergence of openly accessible yet highly capable LLMs such as LLaMA , Falcon , and Mistral allow researchers and practitioners at large to easily obtain, customize, and deploy LLMs in more diverse environments and for more diverse use cases.
Despite the growing influence and accessibility of open-source LLMs, a notable trend has been to restrict visibility and access to their training, fine-tuning, and evaluation processes, including crucial components such as their training code and data. This practice limits the ability of the broader AI research community to study, replicate, and innovate upon advanced LLMs. A more transparent approach to sharing not just the final model but also training details and artifacts is crucial for fostering a more inclusive and collaborative research environment.
Motivated by the above, we note the following specific challenges in LLM research today.
Data Provenance. Understanding the origins and characteristics of the training data is crucial for assessing the reliability and biases inherent in LLMs. A lack of transparency about data sources and composition hinders the ability to identify and mitigate biases which can be perpetuated in model outputs. Simultaneously, data leakage—where training datasets overlap with benchmark datasets—can lead to misleading performance metrics that obscure a model’s general effectiveness (studied in ). These issues highlight the need for clear documentation of data origins and usage in LLM development.
Reproducibility. Even with full disclosure of data sources, the lack of access to complete training code, configuration details, and specific datasets can make it challenging to reproduce the results reported in studies. For example, although the training data mixtures are disclosed by LLaMA , the data processing and training code are not released. Yet, LLMs known to be trained using an open reproduction of LLaMA’s data (e.g., RedPajama ) still do not fully reproduce its benchmark evaluations , indicating that additional data processing or training procedures may be necessary.
Open Collaboration. The practice of only releasing final model weights not only leads to redundant efforts but also poses uniques challenges in conducting certain research. For instance, research into the emergent abilities of LLMs or the investigation of how different training data affects model behavior becomes more challenging without access to intermediate training checkpoints. Researchers are often forced to either work with the final model, which offers limited insights into its developmental nuances, or start from scratch, leading to unnecessary duplication of work and expenditure of compute.
LLM360The name LLM360 signifies open-sourcing LLMs from all angles, and that 360 data points (i.e., checkpoints, data chunks, evaluation results) are released for many of our models. aims to address the issues above through a comprehensive open-source LLM effort. Models in LLM360 are published with all training and model details (e.g., hyperparameters, schedules, architecture, and designs), all intermediate model checkpoints saved during training, and full disclosure of the exact pre-training data used.
We outline the LLM360 framework, focusing on its design principles and the rationale for fully open-sourcing LLMs. We detail the components of the framework, including datasets, code and configurations, model checkpoints, and training metrics. This framework provides a target for transparency that all present and future LLM360 models strive to meet.
We pretrain two new LLMs from scratch and release them under the LLM360 framework. Amber is a 7B English LLM pretrained on 1.3T tokens. CrystalCoder is a 7B English and code LLM pretrained on 1.4T tokens. We discuss the development details, preliminary evaluations, observations, and lessons we learned from Amber and CrystalCoder.
We release all training code, pretraining data, model checkpoints, and evaluation metrics collected during pretraining for both Amber and CrystalCoder. Notably, Amber is released with 360 model checkpoints saved during training, and CrystalCoder with 143.
We aim to make a continuous commitment to fully open-source LLMs by releasing multiple LLMs at various scales. As the first step, in this technical report, we discuss Amber and CrystalCoder, the first open-source LLMs in the LLM360 series. In the future, we plan to release more pre-trained LLMs that are larger in scale, exhibit better performance, and focus on various domains.
The rest of this report is organized as follows. In §2, we discuss related works and the predecessors that inspired LLM360. In §3, we provide a description of the LLM360 framework and the release artifacts that fall into its purview. In §4, we discuss the first two LLMs released under LLM360, Amber (§4.1) and CrystalCoder (§4.1.5), and preliminary analyses of both. §6 concludes.
Related Work
The closest project to LLM360 is Pythia, which also aims at full reproducibility of LLMs . The Pythia project provided 154 checkpoints for model sizes from 70M to 12B to better support research on the scaling behavior and learning dynamics of LLMs. While Pythia is a pioneering work, it no longer reflects many recent LLM practices, such as training over trillion-token datasets or training on language and code in different stages. On the other hand, LLM360 defines a release framework prioritizing transparency and reproducibility under which up-to-date models can continue to be released, and our 7B Amber model surpasses the 12B Pythia model in public benchmarks . Overall, Pythia set an early precedent for transparency and reproducibility of LLMs that we aim to perpetuate and expand in LLM360 to modern LLM pretraining regimes.
In general, open-source LLMs span a wide spectrum of transparency and reproducibility when it comes to their release artifacts. Many recent LLMs only release their final model architecture and weights, keeping their data sources and most training details undisclosed . Some are trained on publicly available datasets , whereas others disclosed their data mixtures but do not make training-ready data available to the public . Several LLMs of note have been released with substantially more transparent details and artifacts. For example, EleutherAI models such as GPT-J and GPT-NeoX included training code, datasets, and up to 150 intermediate model checkpoints. The value of the open-source GPT-NeoX training code was demonstrated by its use in subsequent LLM pretraining by others in the community . INCITE , MPT , and OpenLLaMA were released with training code and training dataset, with RedPajama also releasing 10 intermediate model checkpoints.
Overall, we observe a trend that more recent and capable LLMs are becoming more closed in their release artifacts. In contrast, the goal of LLM360 is to release modern and high-quality models while maintaining a high degree of release transparency.
The LLM360 Framework
In this section we present LLM360, a framework for releasing LLMs that promotes open-source transparency, reproducibility, data/model provenance, and collaborative research. LLM360 provides guidance and recommendations for release artifacts that are collected during LLM pre-training and subsequently made publicly available to the community.
As part of the launch of LLM360, we also release two new pre-trained LLMs, which we hope will foster immediate interest and collaboration in the open-source research community. First, Amber, an English language LLM with 6.7B parameters trained on 1.25 trillion tokens. Second, CrystalCoder, an English and code LLM, also with 6.7B parameters, trained on 1.4 trillion tokens. Details on Amber and CrystalCoder are reported in §4.
The pretraining dataset is the main ingredient of an LLM and significantly impacts its capabilities. Thus, it is important for users and adopters to have visibility into pretraining data to assess potential behavior issues and biases. For example, recent concerns about benchmark data leakage into LLM pretraining is much easier to study when pretraining datasets are available for exploration .
Furthermore, visible pretraining data improves the extensibility of LLMs in later fine-tuning and domain adaptation. Recent work suggests that training on repeated data disproportionately degrades final model performance . Given the breadth of data modern pretraining is performed on, visibility into the original pretraining data is essential for avoiding repeated data in downstream fine-tuning or continued pretraining on specialized domains.
LLM360 advocates for the public release of the data LLMs are pretrained on. When applicable, details about data filtering, processing, and training order should be released as well. Doing so equips the community with better tools to assess the capabilities and risks of LLMs and to reproduce and build upon existing LLMs for future use cases.
These code and settings have a significant impact on the performance and quality of LLM training, and are not always publicly disclosed. For example, we observed that a carefully balanced hybrid data-model-pipeline (3D) parallelism can outperform the standard FSDP in PyTorch by up to 15% on our Nvidia A100 clusters. Another example we observed is that it is essential to keep the inverse frequency matrix in RoPE positional embedding in FP32 , which aligns with the observation in Qwen .
In LLM360, we open-source all our LLM pre-training frameworks, hyperparameters, as well as the configurations. These include the entire training source code, training parameters such as learning rates and batch sizes, and system configurations such as parallelism dimensions.
It is typical during LLM training to periodically save checkpoints of the model to persistent storage. These checkpoints are not only crucial for recovery from faults during training, but also useful in post-training research such as studying different data and/or hyperparameter schedules, or reproducing infrequently-occurring training faults (e.g., loss spikes, NaN results). Recent research on model quantization and compression heavily relies on analysis of model weights and the dynamics during training .
LLM360 models are published with all intermediate checkpoints saved during their training, including model weights and optimizer states (when applicable, e.g., Adam moving averages). These checkpoints enable continued training from a range of starting points without training from scratch, making it easier to study and reproduce a wider variety of effects during training.
LLMs undergo training over weeks to months, and the trends and evolution patterns over this training period can offer valuable information. However, access to detailed logs and intermediate metrics for LLMs is currently limited to groups involved in pretraining, hindering a comprehensive study of LLMs. These statistics often contain key insights that cannot be directly derived otherwise, and even a simple analysis on the metrics, such as computing metric variances or norms, can reveal significant findings. For instance, the team behind GLM proposed an effective gradient shrinking algorithm for handling loss spikes and NaN losses by analyzing gradient norm behaviors .
Our aim with LLM360 is to alleviate this problem by completely open sourcing the logs and metrics we collect. This includes system statistics (e.g., GPU workload), training logs (e.g., loss, gradient norm), and evaluation metrics (e.g., perplexity, downstream tasks). Access to these logs may facilitate a deeper understanding of the whole training process, including how LLMs evolve during various training scenarios. We provide easy access to the figures by sharing directly on the LLM360 Weights & Biases pagehttps://wandb.ai/llm360/projects. A few example metrics include downstream evaluation results, training loss, gradient norm, etc.
In §4.3, we introduce how one can make use of the metrics, and illustrate an experiment tracking the memorization behavior of a model throughout training. The metrics are released in coordination with the data chunks and checkpoints for researchers to easily find their correspondence. Furthermore, we provide open access to the analysis and evaluation code used to foster reproducibility. The code and all the metrics can be found at an LLM360 repository: Analysis360.
Initial Model Release
In this section, we introduce Amber, the first model in the LLM360 family, as well as the finetuned models AmberChat and AmberSafe.
Below we review the details of our pre-training dataset, including data preprocessing, format, data mixing ratios, along with architectural details of our LLM model and specific pre-training hyperparameters. The exact setup of Amber can be found in the LLM360 code base.
We conduct the data preparation process similar to OpenLLaMAhttps://github.com/openlm-research/open_llama#dataset-and-training. Specifically, our pretraining data is a mixture of RefinedWeb, StarCoder, and RedPajama-v1. A slight difference with OpenLLaMA-v2 is our inclusion of C4, since we do not intend to introduce dupliciated documents after the deduplication process conducted by RefinedWeb. We simply put together all the original aforementioned datasets (without any further cleaning, filtering, or sub-sampling), conduct a global permutation, and partition them evenly into 360 data chunks. In total, we have 1.26 Trillion tokens. Table 2 presents the combination.
We used the exact same model architecture as LLaMA 7BThe architectural details are directly fetched from https://huggingface.co/huggyllama/llama-7b. Detailed LLM architectural configurations are summarized in Table 3, incorporating rotary positional embeddings (RoPE) at each layer of the network .
We followed the pre-training hyperparameters from LLaMA as closely as possible . Amber is trained using the AdamW optimizer with the following hyperparameters: . The initial learning rate is set to , following a cosine learning rate schedule that decreases to a final rate of . We apply a weight decay of and use gradient clipping at . The model is warmed up over steps. Differing from the LLaMA setup, based on our hardware setting with 224 GPUs, we use a pre-training batch size of () instead of .
1.2 Details on the Pre-training Infrastructure
Amber is trained on an in-house GPU cluster.
The GPU cluster consists of 56 DGX A100 nodes, each equipped with 80GB A100 GPUs. Each GPU is connected with 4 links NVLink. Cross node connection setting is 2 port 200 Gb/sec (4 HDR) InfiniBand. The throughput we manage to achieve with our distributed training framework is around 582.4 tokens per second.
Our pretraining framework is lit-llamahttps://github.com/Lightning-AI/lit-llama developed based on PyTorch Lightning. We used mixed-precision during pre-training with BF16 for activations and gradients and FP32 for model weights .
1.3 Finetuned Amber models
We also release a few finetuned versions of Amber, namely AmberChat and AmberSafe. AmberChat is trained on the evolved instruct training data as used by WizardLM . We use FastChat to finetune the model for 3 epochs on 8 A100s (80G) distributed by FSDP , the learning rate is , gradient accumulation steps is , warmup ratio is . We also finetune an aligned version of the model: AmberSafe, by conducting Direct Parameter Optimization (DPO) . AmberSafe is trained on ShareGPT 90KThe base model for this is checkpoint 355 instead of the last checkpoint, and further optimized on the SafeRLHF dataset . We set to 0.1, gradient accumulation steps to 4, and the learning rate to .
1.4 Results and Analysis
We use four benchmark datasets in the Open LLM Leaderboardhttps://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard as our evaluation on different aspects, i.e., ARC, HellaSwag, MMLU, and TruthfulQA, following the leaderboard settings. We run the evaluation on all 360 checkpoints, to observe the model ability across the pretraining process. As shown in Figure 4, we can see that the HellaSwag and ARC evaluation scores monotonically increase during pre-training, while the TruthfulQA score seems to decrease as the training proceeds. Another interesting trend is observed in the MMLU progress, where the score decreases in the initial stage of pretraining and then starts to increase.
In Table 4, we compare the final model performance of Amber to a set of models trained around similar time, namely OpenLLaMA, RedPajama-INCITE, Falcon, MPT. Many are inspired by the design of LLaMA. We found that Amber is relatively competitive in scores such as MMLU, but its performance on ARC is behind the curve. We also find that our finetuned Amber models are relatively strong, even compared with other similar models. In our early study, we note that AmberChat simply trained on ShareGPT 90K also demonstrates much higher performance than our base model, which is slightly different from the trends shown on other models in the table. We leave further investigation of this to future work.
1.5 Issues Encountered During Pre-training
In this section, we discuss several major issues encountered during the pre-training process of Amber. These issues could potentially impact our final model performance. We have addressed most of these issues in subsequent LLM pre-training efforts.
During the pre-training procedure, we encountered NaN loss in four out of 360 data chunks. Whenever we faced this issue, we tentatively skipped the entire data chunk. Initially our plan was to train on these four data chunks in later stage of the training, however, we found that these data chunks tend to cause NaN loss regardless of the position of training. We end up finishing our training by taking the first four chunks from the training sequence to complete our learning rate schedule.
In our pre-training framework, we did not manage to save the optimizer states; we only saved model checkpoints for each data chunk. This oversight might be the cause of the NaN loss issue observed in the four data chunks, as mentioned earlier. Each time we resumed pre-training from a previous model checkpoint, the optimizer state in the AdamW optimizer was re-initialized. This re-initialization could potentially affect model training stability.
In the initial phase of pre-training, our codebase had an issue where model checkpoints were saved with BF16 precision, despite our mixed precision training process maintaining model weights at FP32. This issue was later identified and rectified by our team, ensuring that all subsequent model checkpoints were saved with FP32 precision. We anticipate that the initial BF16 model checkpoints may have contributed to some degree of accuracy drop in the model.
2 CrystalCoder
This section provides a summary of the dataset and the model architecture utilized in CrystalCoder. For a detailed evaluation of results on benchmarks and a comparison with previous works on specific benchmarks, we refer readers to our future reports.
The pre-training dataset employed in CrystalCoder is a blend of SlimPajama and StarCoder data with around 1382B tokens in total. Diverging from previous approaches such as Code Llama , which strictly sequentially trains on English and coding data, we adopt a more gradual approach by seamlessly combining and training on both types of data, to provide a balance between code and general ability. The training process is divided into three stages. In the first stage, we train on half of the SlimPajama data, totaling around 345 billion tokens. Moving to the second stage, the remaining half of the SlimPajama data is utilized, along with two epochs of StarCoder data, resulting in approximately 927 billion tokens. In the third stage, we train on Python and web-related data, encompassing HTML, JavaScript, and CSS subsets from StarCoder, totaling 100 billion tokens. Additionally, we sample 10 billion tokens from the SlimPajama dataset in this stage. The preprocessed data and data mixing scripts are released in the Huggingface and Github repository of LLM360.
CrystalCoder employs a model architecture closely resembling LLaMA 7B, with the incorporation of maximal update parameterization (muP) . In addition to this specific parameterization, we have made several slight modifications, the application of RoPE restricted to the first 25% of hidden dimensions (similar to the implementation of GPT-NeoX ), and the use of a sequence length of 2048 with an embedding dimension of 32032. In addition, we simply use LayerNorm instead of RMSNorm since the CG-1 architecture supports efficient computation for vanilla LayerNorm.
CrystalCoder is trained on the Cerebras Condor Galaxy 1 (CG-1), a 4 exaFLOPS, 54 million core, 64-node cloud AI supercomputerhttps://www.cerebras.net/condor-galaxy-1.
We also benchmark this model on the four benchmark datasets in the Open LLM Leaderboard (similar to Amber), as well as coding benchmark datasets, including HumanEval pass@1, and MBPP pass@1. We show results in Figure 6.
3 Analysis360
Prior work such as Pythia has shown that an insightful study can be done by analyzing the intermediate checkpoints of a model. We hope LLM360 can also provide the community useful resources for both reference and research purposes. To this end, we release the initial version of the Analysis360 project, an organized repositories that analyze the model behavior on various aspects, including model characteristics and downstream evaluation results.
As an example of the analysis that can be performed over the set of model checkpoints, we conduct an initial study on memorization in LLMs. Recent work shows that LLMs may memorize a significant part of their training data, which can be extracted with appropriate prompting. Such memorization not only raises privacy concerns in leaking private training data, but also downgrades the performance of LLMs if the training data contains unintended duplicates or peculiarities. As we release all checkpoints and data, we can conduct a comprehensive analysis of memorization across the whole stage of training.
We adopt the memorization score introduced in , indicating the accuracy of tokens in the continuation of length with a prompt of length ,
where is the sequence from training data, while is the generated sequence with prompt . A memorized or -extractible sequence has a memorization score of . Following , we conduct our experiments with . We sampled sequence from each of the data chunks, and use the first tokens of each sequence to conduct the following experiments.
We group the data chunks according to the selected checkpoints, and plot the memorization score on each data chunk group for each checkpoint in Figure 8. We find that 1) Amber checkpoints memorize the latest seen data much more than previous data; 2) For each data chunk, the memorization score drops a bit with additional training, but keeps increasing afterwards.
We show the correlation between sequences in terms of memorization score or -extractible in Figure 9. We witness a strong correlation between the checkpoints.
Summary and Take-home Messages
In this section, we summarize the observations and a few take-home messages from our pre-training of Amber and CrystalCoder, our initial modeling efforts in the LLM360 series. We understand that pre-training is a computationally daunting task that many academic labs or small organizations cannot afford to conduct. We hope that LLM360 can provide comprehensive knowledge, allowing users to understand what happens during LLM pre-training (e.g., loss curve behaviors, how the evaluation metrics emerge, etc.) without the need to do so themselves. We also provide some potential use cases showing how researchers and developers can use LLM360 for their own projects.
Below we list a few of the lessons learned during our initial model training.
In the pre-training of Amber, NaN losses were periodically observed, which may have been caused by certain random states, the training precision, or data quality issues. Some solutions include switching to a different random seed or skipping those data chunks. We notice some “misbehaved” data chunks can cause NaN loss regardless of when they are trained. In a preliminary experiment, we move the “misbehaved” data chunks to the end of the training but still observe NaN losses.
In the pre-training of CrystalCoder and our subsequent LLM pre-training efforts, we observed that a hybrid and carefully tuned parallelism strategy—combining data, tensor-model, and pipeline (also referred to as 3D) parallelism strategies —achieves better system throughput than FSDP, especially in distributed clusters with limited intra-node bandwidth.
Data cleaning (and/or data quality filtering), along with data mixing ratios, are crucial aspects of LLM pre-training, as is the scheduling for various pre-training data categories (e.g., CommonCrawl, Books, StarCoder, etc.). In Amber pre-training, we attempted to adhere as closely as possible to the hyperparameters used in LLaMA; however, our performance still lags significantly behind LLaMA’s. A key omission in LLaMA’s technical report is a detailed description of their exact pre-training dataset. Our carefully crafted CrystalCoder pre-training dataset, which mixes English and coding data, achieves competitive performance with LLaMA on both the Open LLM Leaderboard and Code Evaluation benchmarks. We, along with the entire LLM open-source community, are diligently exploring the best approaches for data cleaning, data quality filtering, and determining optimal data mixing ratios, a pioneering effort exemplified by the DoReMi method .
We describe a few potential use cases of LLM360 below.
One can conduct experimental studies at any stage of model training. As previously mentioned, the optimal data mixing ratio remains a significant open problem in LLM pre-training. However, it is nearly impossible to verify a specific mixing ratio by conducting full LLM pre-training. A more feasible approach is to adjust the data mixing ratios on the fly, i.e., starting from an intermediate checkpoint, and either increasing or decreasing a specific data ratio from a particular category, e.g., increasing the data weight in Wikipedia.
For building domain-specific LLMs (e.g., medical, finance, law, etc.), one may not necessarily want to start from the last pre-trained LLM checkpoint (which would make it more akin to fine-tuning). Instead, one can always pick one of the LLM360 checkpoints (e.g., from 50% of the pre-training stage) and resume the pre-training to obtain a domain-specific LLM.
A lot of algorithmic approximation frameworks for efficient training require partially trained model weights . LLM 360 provides perfect model initializations for those methods.
Given the wide-ranging applicability and high performance of LLMs, applications powered by them have the potential to deeply influence various aspects of life. Consequently, it becomes essential for all parties involved in the chain of production of LLMs to carefully manage the potential impact and risks associated with them. All stakeholders need to be informed of these implications and take necessary actions accordingly.
We believe the transparent nature of the LLM360 initiative can help make the potential risks known to stakeholders. As one example, many risks associated with LLMs are related to certain forms of biases , such as the risk of social stereotypes, discrimination and exclusion, and the risk of under-representing certain languages or domains. By inspecting the exact training data and bias analysis (e.g. BOLD ) in Analysis360, stakeholders can have a thorough review of these risks before deploying the models. LLM360 can also help with risk mitigation. The project shares reproducible traces and exact data during LLM training, providing a reusable environment for researchers to conduct experiments to design better guardrails to contain potential risks.
We understand the importance of controlling the risk of LLMs and we are committed to further developing the LLM360 framework to foster responsible usage of LLMs. We would like invite the community to work with us, by sharing research results or by simply providing feedback.
Conclusion and Future Work
In this paper, we introduce LLM360, an initiative for comprehensive and fully open-sourced LLMs. Along with the first release of LLM360, we released two 7B LLMs: Amber (an English general-purpose LLM) and CrystalCoder (an LLM pre-trained specifically for code generation). In terms of artifacts, we released pre-training code, configurations, hyperparameters, intermediate model checkpoints, optimizer states, as well as the data sequence and data processing code. Our vision is to significantly advance and promote transparency within the open-source LLM pre-training community.
For future work, we are conducting a more detailed analysis on Amber and CrystalCoder’s base models as well as their fine-tuned models. Detailed results will be released and discussed in their respective technical reports. Our team is also pre-training a much larger LLM, which will be fully released as soon as the pre-training is complete. Additionally, we will explore the optimal ratios for mixing different subsets in the pre-training datasets.
We would like to thank Natalia Vassilieva, Joel Hestness, William Marshall, and Bhargav Kanakiya for their contribution to CrystalCoder and support on the LLM360 project. We would also like to thank the MBZUAI and Cerebras team for providing and managing the computing infrastructure.