AstroLLaMA: Towards Specialized Foundation Models in Astronomy

Tuan Dung Nguyen, Yuan-Sen Ting, Ioana Ciucă, Charlie O'Neill, Ze-Chang Sun, Maja Jabłońska, Sandor Kruk, Ernest Perkowski, Jack Miller, Jason Li, Josh Peek, Kartheik Iyer, Tomasz Różański, Pranav Khetarpal, Sharaf Zaman, David Brodrick, Sergio J. Rodríguez Méndez, Thang Bui, Alyssa Goodman, Alberto Accomazzi, Jill Naiman, Jesse Cranney, Kevin Schawinski, UniverseTBD

Introduction

The advent of Large Language Models (LLMs) has sparked interdisciplinary interest thanks to a confluence of factors: accumulation of massive datasets, leaps in computational power and breakthroughs in neural architectures. Flagship models like GPT-4 (OpenAI, 2023), PaLM (Chowdhery et al., 2022; Goo, ) and LLaMA (Touvron et al., 2023; Meta, 2023) have exhibited exceptional versatility in a variety of tasks from logical reasoning and comprehension to creative writing, often accomplished via methods like prompting, fine-tuning, and human-in-the-loop reinforcement learning.

The astronomy discipline presents both a unique challenge and a fertile ground for the application of LLMs. First, the corpus of scholarly texts in astronomy likely constitutes but a minuscule portion of the data on which generic LLMs are trained, resulting in limitations like hallucinations in favor of more “generic” responses. Second, the nature of astronomical research often involves cross-disciplinary insights due to universally applicable physical processes. When well-curated, LLMs could meaningfully assist in hypothesis generation.

Existing scales based on in-context prompting and instruction learning, primarily involving GPT-4, have already demonstrated significant potential for generating substantive hypotheses (Ciucă and Ting, 2023; Ciucă et al., 2023). Further, the astronomy community’s “open sky” policy, which grants public access to the majority of its datasets either immediately or after a brief proprietary period (Almeida et al., 2023; Fabricius et al., 2021), pairs well with the wealth of resources available in archives like NASA’s Astrophysics Data System (Accomazzi et al., 2015; Borgman and Wofford, 2021). Such an open-access policy can facilitate deep engagement with the astronomical literature.

Despite their general capabilities, LLMs frequently lag behind specialized, smaller models in domain-specific applications. This disparity stems from two primary factors: (i) the eclectic nature of the training datasets, which dilutes the focus on specialized subjects, and (ii) the design ethos of LLMs as “foundation models” meant for subsequent fine-tuning tailored to specific tasks. The existing landscape for fine-tuned LLMs in astronomy remains limited, however. To our knowledge, the only existing specialized model is astroBERT (Grezes et al., 2021), which has 110 million parameters, trained on nearly 400,000 ADS papers. But as an non-generative model, the utility of astroBERT remains limited to discriminative tasks.

Motivated by these gaps, we present AstroLLaMA, a state-of-the-art generative language model fine-tuned from LLaMA-2. Our model leverages a corpus of 300,000 astronomy abstracts from arXiv and boasts an architecture approximately 67 times larger than that of astroBERT. AstroLLaMA aspires to build upon astroBERT’s foundation by offering improved performance in generating specialized information.

AstroLLaMA

In this section, we discuss AstroLLaMA’s implementation, focusing on the curation of its dataset, base model architecture, and fine-tuning settings.

We derive our dataset from the arXiv repository, available on Kaggle.https://www.kaggle.com/Cornell-University/arxiv Our curated subset focuses on papers classified under the astrophysics category (astro-ph), resulting in a collection of 326,238 articles spanning from April 1992 to July 2023. We extract the these papers’ abstracts to form a corpus consisting of approximately 95 million tokens. The median length of these abstracts is 291 tokens. To enable effective model evaluation, we randomly designate 20% of this curated dataset for testing.

2 Base Model

Our base model is LLaMA-2, a 6.7 billion-parameter model developed by Meta (Meta, 2023). Originally trained on a corpus containing 2 trillion tokens, LLaMA-2 features a context window of 4,096 tokens. For tokenization, the model employs a bytepair encoding strategy (Sennrich et al., 2016; Kudo and Richardson, 2018), incorporating a vocabulary set of 32,000 unique tokens.

3 Fine-tuning Settings

For the fine-tuning phase, we rely on our curated training set described in Section 2.1, which includes 77 million tokens. Special [BOS] (Beginning Of Sequence) and [EOS] (End Of Sequence) tokens are prepended and appended to each training sequence. These sequences are then concatenated and divided into fixed-length chunks, each comprising 512 tokens.

The fine-tuning process follows the causal language modeling objective employed during the model’s pre-training phase. We use the AdamW optimizer (Loshchilov and Hutter, 2018) with hyperparameters β1=0.9,β2=0.95,ϵ=10−5\beta_{1}=0.9,\beta_{2}=0.95,\epsilon=10^{-5} and a batch size of 32. The learning rate follows a cosine schedule with a linear warmup to a peak value of 3×10−43\times 10^{-4} in the first 10% of the optimization steps and a final learning rate of 10% of its peak. Additional settings include weight decay and gradient clipping values of 0.1 and 1.0, respectively.

We fine-tune LLaMA over nearly three epochs, corresponding to about 230 million processed tokens, using four NVIDIA A100 GPUs, each equipped with 40GB of VRAM. To maximize resource efficiency, we employ 4-bit quantization and utilize LoRA, a technique based on low-rank matrix decomposition (Hu et al., 2021). We set LoRA’s hyperparameters α\alpha and dropout rate to 32 and 0.05, respectively. The entire process is facilitated through the Hugging Face Python library.

4 Fine-Tuning Evaluation

Fig. 1 depicts the performance of AstroLLaMA during its fine-tuning phase. Here, we present perplexity, a commonly used metric for evaluating causal language models. Perplexity is defined as the exponentiation of the training loss, with lower values indicating a better fit.

Our initial observations reveal that LLaMA-2 performs suboptimally on our dataset, with an average perplexity close to 10. By the conclusion of three epoch, AstroLLaMA achieves an average perplexity of 6.55. This represents a 32.5% reduction in perplexity compared to the base LLaMA-2 model, signifying a substantial improvement in the model’s predictive accuracy.

Results

As illustrated in the previous section, AstroLLaMA outperforms its non-fine-tuned counterpart, LLaMA-2, in terms of context-awareness during token prediction within astronomy abstracts. To delve deeper into the advantages of fine-tuning, we assess AstroLLaMA’s general abilities in two key aspects: text generation and embedding space quality. We compare its performance against multiple models, including LLaMA-2, GPT-4 and GPT-3 (ada-002) to provide a comprehensive evaluation.

Regarding text generation, we task AstroLLaMA, LLaMA-2 and GPT-4 with completing various astronomy-related abstracts, an example of which is presented in Fig. 2. Each model is given the first few sentences of an abstract as a prompt, allowing us to gauge its ability to comprehend the context and generate a meaningful continuation. For GPT-4, we utilize ChatGPT and specifically prompt it to limit the completion to a single paragraph. AstroLLaMA and LLaMA-2 are deployed using standard sampling methods, with the temperature set to 0.3 and a maximum new tokens limit of 1,024. We find that altering the temperature setting does not substantively improve LLaMA-2’s results.

Our observations largely echo the patterns depicted in Fig. 2. LLaMA-2 often deviates from the intended context after generating only a short and often off-topic continuation, resulting in inferior completions. While GPT-4 produces more coherent text, its responses are too generic to capture the nuanced understanding required in the astronomy domain. Even when explicitly prompted to focus on astronomy-related topics, GPT-4’s generated text remains largely off-target or generically applicable rather than domain-specific.

In stark contrast, AstroLLaMA exhibits remarkable context-awareness in its completions by showing a deep understanding of astronomical concepts. For example, in Fig. 2, AstroLLaMA comprehends that an effective search for stars in the Magellanic Stream involves a three-step process: initial wide-field imaging, followed by refinement using astrometric data from Gaia, and then further curation with spectroscopic data. The model also understands Gaia-ESO is surveying the southern sky and hence can observe (part of) the Magellanic Stream. It also demonstrates nuanced knowledge of the Magellanic Stream, understanding the importance of bifurcation within the stream. As a result, it appropriately completes the text by discussing the southeast stream and exploring metallicity differences to ascertain their origins.

Regarding embedding space quality, we assess models’ ability to reflect semantic similarities among astronomy texts. We randomly choose 10,000 abstracts from our dataset and embed them using AstroLLaMA and GPT-3. Specifically, we use OpenAI’s API to invoke the text embedding function for GPT-3 (ada-002). To get text embeddings from AstroLLaMA, we pass an input through the model and extract its final hidden states, which contain embeddings for all tokens in the input. Then, we omit the [BOS] token and take the average of all other tokens’ embeddings to get the final result. Finally, for each pair of abstracts we calculate their cosine similarity (the normalised dot product) between on their vector embeddings.

The top panel of Fig. 3 presents the distribution of these pairwise similarities for the two embedding methods. We find that the embeddings by GPT-3 are overly generic with similarities clustering around relatively high values of 0.7–0.9, suggesting a lack of discriminative power (most papers are embedded very similarly). AstroLLaMA’s embeddings, on the other hand, exhibit much higher variance within each bin. This suggests that our fine-tuned model is more adept at representing the specialized semantic variance inherent to the field of astronomy, which may enable a more granular representation of astronomical content and can facilitate better document retrieval and semantic analysis.

The bottom panel of Fig. 3 provides two representative examples where AstroLLaMA and GPT-3 classifications diverge. In the first example, GPT-3 fixates on the keyword ‘magnetized,’ resulting in an inflated similarity score, despite the contexts being markedly different. AstroLLaMA, on the other hand, successfully distinguishes between these disparate contexts. In the second example, AstroLLaMA accurately identifies that the study of Spitzer is closely related to star formation. GPT-3, however, fails to make this connection due to the absence of matching keywords.

Limitations and Future Directions

In this work, we introduce AstroLLaMA, a 7-billion-parameter language model fine-tuned on a dataset encompassing over 300,000 abstracts from astronomical research papers. Compared to its base model, LLaMA-2, and even GPT-4, a current state-of-the-art general LLM, AstroLLaMA exhibits marked improvements in generating high-quality abstracts with a competent grasp of relevant information in this literature.

AstroLLaMA is not without limitations, nevertheless. The most salient is the model’s knowledge gaps in certain areas of astronomy: in Fig. 2, AstroLLaMA’s estimation of potential star candidates from Gaia-ESO data is notably inaccurate. To address such issues, we are in the process of enriching AstroLLaMA’s training set with not just abstracts but the full LaTeX sources of existing astronomy articles, thereby expanding the token count by approximately two orders of magnitude. Another concern lies in the model’s tendency to generate hallucinated or fictitious numerical data, an issue likely attributed to our focus on reducing perplexity rather than explicitly steering the model towards factual accuracy. The release of AstroLLaMA aims to facilitate community engagement, both for addressing these inaccuracies and for refining its balance between “faithfulness” (respecting scientific evidence and accuracy) and “creativity” (being able to come up with interesting hypotheses).

AstroLLaMA stands as a compelling prototype for specialized LLMs in astronomy, showing superior context-aware capabilities compared to GPT-4 despite having much fewer parameters. It not only paves the way for improved performance in tasks like question-answering, scientific summarization and hypothesis generation but applies also to multi-modal models (Liu et al., 2023). We have made the AstroLLaMA’s weights and its training data publicly availablehttps://huggingface.co/universeTBD/astrollama for researchers interested in leveraging LLMs for astronomy-centric applications. Along with this, we are establishing various “playgrounds” on Hugging Face to invite interested readers to further adapt and refine this robust starting point for a variety of relevant downstream tasks.

Acknowledgments

We are deeply grateful to the Microsoft Accelerate Foundation Models Research Initiative for enabling us to fast-track our project. Thanks to advanced AI platform from Microsoft Research, we have been able to significantly expedite our efforts in using language models to analyze astronomical literature.

Ethics Statement

We obtain the pre-trained weights for LLaMA-2 from Meta, which offers these models for download on Hugging Face. The arXiv dataset used in this paper is publicly available on Kaggle. While we have demonstrated that AstroLLaMA is capable of generating high-quality, relevant abstracts for astronomical research papers, we have noted that it has the potential to generate inaccurate data and measurements. This should serve as a caution for researchers aiming to use this model for downstream tasks, and we invite the adoption of alignment strategies in future work to ameliorate this issue.

References