The GALAH+ Survey: Third Data Release

Sven Buder, Sanjib Sharma, Janez Kos, Anish M. Amarsi, Thomas Nordlander, Karin Lind, Sarah L. Martell, Martin Asplund, Joss Bland-Hawthorn, Andrew R. Casey, Gayandhi M. De Silva, Valentina D'Orazi, Ken C. Freeman, Michael R. Hayden, Geraint F. Lewis, Jane Lin, Katharine J. Schlesinger, Jeffrey D. Simpson, Dennis Stello, Daniel B. Zucker, Tomaz Zwitter, Kevin L. Beeson, Tobias Buck, Luca Casagrande, Jake T. Clark, Klemen Cotar, Gary S. Da Costa, Richard de Grijs, Diane Feuillet, Jonathan Horner, Prajwal R. Kafle, Shourya Khanna, Chiaki Kobayashi, Fan Liu, Benjamin T. Montet, Govind Nandakumar, David M. Nataf, Melissa K. Ness, Lorenzo Spina, Thor Tepper-Garcia, Yuan-Sen Ting, Gregor Traven, Rok Vogrincic, Robert A. Wittenmyer, Rosemary F. G. Wyse, Marusa Zerjal, the GALAH collaboration

Introduction

During the history of the Milky Way, the abundances of the different elements that make up the Galaxy’s stars and planets have continually changed, as a result of the processing of the interstellar medium by successive generations of stars. As a result, the study of the elemental abundances in stars provides a direct record of the galaxy’s history of star formation and evolution - a fact that has, in recent years, given birth to the science of Galactic Archaeology.

Until recently, however, observational limitations meant that the data available to answer the questions of how the Milky Way formed and evolved was restricted to a few hundred or thousand stars with high-quality element abundances in our Solar neighbourhood (Edvardsson et al. 1993; Nissen & Schuster 2010; Bensby et al. 2014, see e.g.). In the last decade, advances in multi-object observations made by spacecraft (such as Gaia) and ground-based facilities have brought about a revolution in the field of Galactic Archaeology. Where once the field was forced to focus on single-star population studies, it is now possible to carry out surveys that allow large-scale structural analyses.

Due to the intrinsic difficulty in determining the distances of stars, studies of the chemodynamical evolution of our Milky Way were previously restricted to nearby stars which were mapped by the Hipparcos satellite (ESA 1997; Perryman et al. 1997; van Leeuwen 2007). In the era of the Gaia satellite (Gaia Collaboration et al. 2016a; Gaia Collaboration et al. 2016b; Gaia Collaboration et al. 2018), we can now use astrometric and photometric observables and their physical relations with spectroscopic quantities to improve the analysis of spectra and thus the estimation of element abundances.

The connections between the chemical compositions and dynamics of stars across the vast populations in our Galaxy are a topic of significant ongoing research. Although we speak of the Milky Way in terms of the thin and thick disc (Yoshii 1982; Gilmore & Reid 1983), the bulge (Barbuy et al. 2018), and the stellar halo (Helmi 2020) as its main components (Bland-Hawthorn & Gerhard 2016), we understand that the Galaxy is more than a superposition of independent populations. With the data now at hand, we can analyse the Galaxy from a chemodynamical perspective, and use stars of different ages as time capsules to trace back the formation history of our Galaxy (Rix & Bovy 2013; Bland-Hawthorn et al. 2019, see e.g.). As one example, the most recent data release from Gaia has enabled significant leaps in our understanding of the enigmatic Galactic halo (Helmi 2020, for an overview see e.g.). 6-d phase space information from Gaia has revealed a large population of stars in the Solar neighbourhood that stand out against the smooth halo background as a coherent dynamical structure, pointing to a significant accretion event that is currently referred to as “Gaia-Enceladus-Sausage” (GES) a combination of “Gaia-Enceladus” (Helmi et al. 2018) and “Gaia Sausage” (Belokurov et al. 2018). Additionally, while we would expect the chemical composition of stars to be correlated with their ages and formation sites (Minchev et al. 2017, see e.g.), observations can now clearly demonstrate these connections (Feuillet et al. 2018; Buder et al. 2019, see e.g.), and can also demonstrate that stars within our Solar neighbourhood have experienced significant radial migration through their lifetimes (Frankel et al. 2018; Hayden et al. 2020, see e.g.).

Previous and ongoing spectroscopic surveys by collaborations like RAVE (Steinmetz et al. 2020a; Steinmetz et al. 2020b), Gaia-ESO (Gilmore et al. 2012), SDSS-IV APOGEE (Ahumada et al. 2020), and LAMOST (Cui et al. 2012; Xiang et al. 2019) have certainly shed light on several of these outstanding questions. Answering them completely, requires more and/or better data to map out the correlations between stellar ages, abundances, and dynamics. Upcoming surveys like SDSS-V (Kollmeier et al. 2017), WEAVE (Dalton et al. 2018), 4MOST (de Jong et al. 2019), and PFS (Takada et al. 2014) will certainly continue to broaden our capabilities and understanding surrounding our galaxy’s physical and chemical evolution. The data currently at hand, derived from spectroscopy, photometry, astrometry, and asteroseismology, provide high-dimensional information, and we must develop methods to extract the most accurate and precise information from them (Nissen & Gustafsson 2018; Jofré et al. 2019, for reviews on this see e.g.).

The recent growth in the quantity of available spectroscopic stellar data has delivered a new technique to galactic archaeologists - namely ‘‘Chemical Tagging’’, which allows the identification of stars that formed together using their chemical composition and an understanding of the astrophysics driving the dimensionality of chemical space. This technique is proving a vital tool, enabling us to observationally isolate and characterise the building blocks of our Galaxy. As a result, it remains a major science driver for the GALactic Archaeology with HERMES High Efficiency and Resolution Multi-Element Spectrograph (GALAH) collaboration https://www.galah-survey.org (De Silva et al. 2015). With the large variety of nucleosynthetic channels that can enrich the birth material of stars (Kobayashi et al. 2020, see e.g.), the hypothesis is that we should be able to disentangle stars with different enrichment patterns, provided we observe enough elements with different enrichment origins. The success of some chemical tagging experiments (Kos et al. 2018; Price-Jones et al. 2020, see e.g.) is challenged by the broad similarities in chemical abundance in populations like the low-\upalpha\upalpha disc (Ness et al. 2018, see e.g.), and by the small but real inhomogeneities even within star clusters (Liu et al. 2016a; Liu et al. 2016b). To put detailed chemical tagging into action, we will need a massive dataset (Ting et al. 2016, see e.g.) consisting of measurements made with outstanding precision (Ting & Weinberg 2021).

The publication of the previous second data release of the GALAH survey (Buder et al. 2018), entirely based on observations as part of GALAH Phase 1 with the HERMES spectrograph at the Anglo-Australian Telescope, has provided for the community abundance measurements of 23 elements based on 342 682 spectra. Observations for this phase have continued and we are able to publish all 476 863 spectra for 443 843 stars of the now finished Phase 1 observations as part of this third data release. In parallel, spectroscopic follow-up observations of K2 and TESS targets have been performed with HERMES and we are able to also include these observations (112 943 spectra for 99 152 K2-HERMES stars and 34 263 spectra for 26 249 TESS-HERMES stars) in our release. We further include ancillary observations of 54 354 spectra for 28 205 stars in fields towards the bulge and more than 75 stellar clusters. Given the significant contribution from these programs to this GALAH release, we will hereafter refer to the release as GALAH+ DR3.

For the previous (second) data release of the GALAH survey (Buder et al. 2018), we made use of the data-driven tool The Cannon (Ness et al. 2015) to improve both the speed and the precision of the spectroscopic analysis. This was performed almost entirely without non-spectroscopic information for individual stars, using a “training set” of stars with careful by-hand analysis. Although the data-driven approaches were successful for the majority of GALAH DR2 stars, we know that these approaches can suffer from signal aliasing (e.g. moving outliers closer to the main trends), can learn unphysical correlations between the input data and the output stellar labels, and that the results are not necessarily valid outside the parameter space of the training set. As part of the present study, we aim to assess how accurately and precisely the stellar parameters and abundances were estimated by the data-driven approaches. We have therefore adjusted our approach to the analysis of the whole sample and now restrict ourselves to smaller wavelength segments of the four wavelength windows observed by HERMES per star with reliable line information for spectrum synthesis instead and include even more grids for an accurate computation of line strengths when the conditions depart from local thermodynamic equilibrium (LTE; e.g. Mihalas & Athay 1973; Asplund 2005; Amarsi et al. 2020).

The publication of Gaia DR2 (Gaia Collaboration et al. 2018; Lindegren et al. 2018) provided phase space information up to 6 dimensions (coordinates, proper motions, parallaxes, and sometimes also radial velocities) for 1.3 billion stars, and having this information available for essentially all (99%) stars in GALAH has allowed us to make major improvements to our stellar analysis. By combining our knowledge of the (absolute) photometry and spectroscopy of stars, we can break several of the degeneracies in our stand-alone spectroscopic analyses, which arise due to the fact that absorption lines do not always change to a detectable level as a function of stellar atmospheric parameters. The data analysis process for this third data release from the GALAH collaboration makes use of fundamental correlations, and this quantifiably improves the accuracy and precision of our measurements.

As large Galactic Archaeology-focused surveys continue to collect data (like GALAH in its ongoing Phase 2), the overlap between them increases. This enables us to compare results when analysing stars in the overlap, which have the same stellar labels (stellar parameters, abundances or other non-spectroscopic stellar information), and it also allows us to propagate labels from one survey onto another (Casey et al. 2017; Ho et al. 2017; Xiang et al. 2019; Wheeler et al. 2020; Nandakumar et al. 2020, see e.g.). This label propagation makes it possible to combine these complementary surveys for global mapping of stellar properties and abundances, and we show an example of this in Section 8, placing GALAH+ DR3 data in context with the APOGEE and LAMOST surveys.

This paper is structured as follows: We describe our target selection, observations, and reductions in Sec. 2. While the target selection and observation of the several projects like K2-HERMES and TESS-HERMES were slightly different from the main GALAH survey, we have reduced and analysed all data (combined under the term GALAH+) in a consistent and homogeneous way. The analysis of the reduction products is described in Sec. 3, focusing on the description of the general workflow of the analysis group and highlighting changes with respect to the previous release (GALAH DR2). Secs. 4 and 5 address the validation efforts for stellar parameters and element abundances, respectively. These address the accuracy and precision of these labels as well as our algorithms to identify and flag peculiar measurements or peculiar stars. Based on experience with the data set, we stress the importance of the flags, but also how complex the flagging estimates are, with several examples of peculiar abundance patterns. We also highlight possible caveats (and possibly peculiar physical correlations) of our analysis in Sec. 6. We present the contents of the main catalogue of this data release in Sec. 7. In this section we also present the Value-Added-Catalogues (VACs) that accompany this release, namely a VAC (see Sec. 7.3.1) created by crossmatching our targets with Gaia eDR3 (Gaia Collaboration et al. 2020) and the distance estimates by Bailer-Jones et al. 2020, a VAC (see Sec.7.3.2) with estimates (such as stellar ages and masses) from isochrone fitting, a VAC (see Sec. 7.3.3) with stellar kinematic and dynamic estimates, a VAC (see Sec. 7.3.4) with radial velocity estimates based on different methods, and a VAC (see Sec. 7.3.5) on parameters of binary systems. While we made use of data from Gaia DR2 (Gaia Collaboration et al. 2018) for our spectroscopic analysis, we also provide a second version of each of our catalogues with new crossmatches and VACs with updated data making use of Gaia eDR3 (Gaia Collaboration et al. 2020), which was published shortly after our data release and supersedes Gaia DR2. We describe all changes of the catalogues between version 1 (based on Gaia DR2) and version 2 (with VACs now based on Gaia eDR3) in Sec. 7.1. We highlight the scientific potential of the data in this release in context by using the combination of dynamic information and ages together with the element abundances of the main catalogue in Sec. 8. We focus on Galactic Archaeology on a global scale and the chemodynamical evolution of our Galaxy. Along with the main and value-added-catalogues of this release, we publish the observed optical spectra for each of the arms of HERMES on the DataCentral https://docs.datacentral.org.au/galah/ and provide the scripts used for the analysis as well as post-processing online in an open-source repository http://github.com/svenbuder/GALAH_DR3 GALAH+ DR3 was timed to allow the scientific community to directly use abundances together with the latest Gaia eDR3 information. We have not yet incorporated the latter into our abundance analysis, but plan to do so in future data releases, as we outline in Sect. 9. In this section, we conclude and give an outlook to future observations as part of the ongoing observations of the GALAH survey (called phase 2 with an adjusted target selection) and our next data release.

Target selection, observation, reduction

The magnitude limited selection of the GALAH survey (see the magnitude distribution in Fig. 4a) causes a strong correlation between increasing distance (and decreasing parallax quality) with increasing luminosity. This tradeoff between luminosity and parallax uncertainty was also visible for the stars in common between Gaia DR1 and GALAH DR2 (Buder et al. 2019) and is still present with the use of Gaia DR2, as we illustrate in Kiel diagrams in Fig. 3b, showing that especially giants with larger distances suffer from large parallax uncertainties.

2 Observations

GALAH data are acquired with the 3.9-metre Anglo-Australian Telescope at Siding Spring Observatory. Up to 392 stars can be observed simultaneously using the 2dF robotic fibre positioner (Lewis et al. 2002) that sits at the telescope’s prime focus. The fibres run to the High Efficiency and Resolution Multi-Element Spectrograph (De Silva et al. 2015; Sheinis et al. 2015, HERMES;), where the light is dispersed at R∼28 000R{\sim}28\,000 and captured by four independent cameras. HERMES records ∼1000{\sim}1000 Å of the optical spectrum across its four non-contiguous channels (4713-4903, 5648-5873, 6478-6737, and 7585-7887 Å). Details of the instrument design and as-built performance of HERMES can be found in Barden et al. 2010, Brzeski et al. 2011, Heijmans et al. 2012, Farrell et al. 2014 and Sheinis et al. 2015.

Since HERMES was first commissioned, raw data it obtains has been contaminated by odd saturated points with vertical streaking, which was traced back to the choice of glass for the field flattening lens inside each of the four cameras (Martell et al. 2017). The original glass had been chosen for its high index of refraction, but uranium in the glass emitted \upalpha\upalpha particles that caused the saturated points and vertical readout streaks when they were captured by the HERMES CCDs (Edgar et al. 2018). In the first half of 2018, the original field flattening lenses were replaced with lenses made from a less radioactive glass, and the vertical streaks have almost stopped occurring in the data. The point spread function in the HERMES cameras changed as a result of changing the field flattening lenses, and is now larger and less symmetric in the corners of the detectors. As part of HERMES recommissioning, the GALAH team fed light from a Fabry-Perot interferometer into HERMES to characterise the new PSF across each detector, and this information has been incorporated into the data reduction procedure.

The observing procedure and targeting strategy for this data release are the same as for previous public GALAH data, including the selection of fields via the GALAH-internal obsmanager (keeping track of already observed fields and suggesting fields with lowest airmass at a given observing time for a given program) and the assignment of targets onto 2dF fibres via configure (Miszalski et al. 2006). For further information on the strategy of GALAH Phase 1, with the GALAH main and faint observations, we refer the reader to Buder et al. 2018. For the K2-HERMES observing strategy, the reader is referred to Wittenmyer et al. 2018 and Sharma et al. 2019, and for TESS-HERMES to Sharma et al. 2018.

GALAH+ DR3 expands the number of targets from DR2 significantly and includes all data taken between November 2013 and February 2019. The distribution of GALAH+ DR3 stars across VJKV_{JK} and signal-to-noise ratio (S/N) is shown in Fig. 4b and adds another perspective on the complex correlation of luminosity (or surface gravity log⁡g\log g) with S/NS/N for the observed stars, as shown in Fig. 3c.

GALAH, K2-HERMES, and TESS-HERMES observers choose from a database of available fields depending on conditions, limiting the hour angle to within ±2\pm 2 hours whenever possible. The standard observing procedure for regular GALAH survey fields is to take three 1200s exposures, with an arc lamp and flat lamp exposure taken at the same sky position as each field to enable proper extraction and calibration of the data. Bright-star fields are observed in evening and morning twilight, and when the seeing is too poor for the regular survey fields. They receive three 360s exposures and the same calibration frames as for the regular fields.

The median seeing at the AAT is 1\aas@@fstack′′51\aas@@fstack{\prime\prime}5, and the exposure time is extended by 33 %33\,\% if the seeing is between 2\aas@@fstack′′02\aas@@fstack{\prime\prime}0 and 2\aas@@fstack′′52\aas@@fstack{\prime\prime}5 and by 100 %100\,\% if the seeing is between 2\aas@@fstack′′52\aas@@fstack{\prime\prime}5 and 3\aas@@fstack′′03\aas@@fstack{\prime\prime}0. This exposure time was chosen to achieve a S/N of 50 per pixel (equivalent to 100 per resolution element) in the HERMES green channel (CCD 2). This is accomplished in nominal seeing when a star has an apparent magnitude of 14 in the photometric band matched to the camera (B=14 and CCD 1, V=14 and CCD 2, etc.) Mismatches between predicted and actual data quality are mainly caused by a combination of seeing, cloud and airmass. We show the distribution for the actual S/N per pixel as a cumulative distribution for all four HERMES channels in VJKV_{JK} in Fig. 4c. Depending on the spectral type, the S/N achieved for a given star in each CCD varies (i.e., a hot star will be brighter in the blue and green passbands and fainter in the red and infrared passbands).

3 Reductions

The main improvement is the wavelength solution, which is now more stable at the edges of the green and red CCDs, where we lack arc lines. This has been achieved by monitoring the solution and fixing the polynomial describing the pixel-to-wavelength transformation, if deviations from a typical or average solution are detected. The solution is described by a 4th order Chebyshev polynomial. We use IRAF’s identify function to find the positions of arc lines in each image and match them with our linelist. Fitting the solution, however, is now done in a more elaborate way. Initially, all spectra from the same image are allowed to have an independent solution. Then the four coefficients of the Chebyshev polynomial are compared. The first coefficient defines the zero-point. Because the 2dF fibres are not arranged monotonically in the pseudo-slit, the first coefficient is truly independent of the spectrum number (spectra being numbered 1 to 400 in each image). The values of the other three coefficients should be a smooth function of the spectrum number. If a coefficient for a specific spectrum deviates by more than 3\upsigma3\upsigma from a smooth function, it is corrected to lie on the smooth function. This successfully fixes the previous problems with incorrect wavelength solutions at the edge of the image.

Our improved reduction pipeline also features an improved parameterization of cross-talk. It can only be measured in larger gaps between every 10th spectrum. Cross-talk was previously represented as a function of the position in the image, but now each batch of 10 spectra (from one slitlet) is assigned the measured cross-talk without any interpolation. The cross-talk is still a function of the direction along the dispersion axis. The normalisation has been improved with a new identification of continuum sections (regions of a spectrum where the continuum is measured) and optimised polynomial orders. The pipeline has been actively maintained and adapted to perform well with the recommissioned instrument following the replacement of the field flattening lenses in 2018 May. Other minor improvements and computing optimisations were made.

As we lay out in the next Section, with the analysis approach via spectrum synthesis chosen for GALAH+ DR3, we will not make use of the full spectra of GALAH, but restrict ourselves to absorption features with reliable line information for spectrum synthesis, when estimating stellar parameters and abundances.

Data analysis

The two most important differences to the workflow of our analysis are the following: First, we are using astrometric information from the Gaia mission to break spectroscopic degeneracies. Secondly, we do not use data-driven approaches for the spectrum analysis in GALAH+ DR3, but only the spectrum synthesis code Spectroscopy Made Easy (Valenti & Piskunov 1996; Piskunov & Valenti 2017, hereafter sme), which had only been used for the training set analysis in DR2. We visualise the reasons for this step with the comparison of GALAH DR2 and DR3 in Fig. 5. We found in DR2 that stars at the periphery in stellar label space, e.g. high temperature (compare panels a) and d) or low metallicity (compare panels b) and e)) did not receive optimal labels from the data driven process.

2 The general workflow

Our general workflow follows the same approach as the spectrum synthesis analysis for DR2, with the aim to homogeneously and automatically analyse a large number of spectra that intrinsically look very different. The analysis is divided into two fundamental steps: first, we estimate the stellar parameters; and second, we keep the stellar parameters fixed while only fitting one abundance at a time for the different lines/elements in the GALAH wavelength range. For the stellar parameter estimation (first step), we first perform a normalisation and a first rough stellar parameter fit with one iteration, followed by a final normalisation and finer parameter fit that is iterated on until convergence. For this, we limit ourselves to 46 selected segments of the GALAH spectra which include the lines with reliable line data for the stellar parameter analysis, as described in detail in Sec. 3.3. For the abundance analysis (second step), we perform only one normalisation and iteratively optimise the abundance based on those data points of the lines/elements that we estimate to be unblended enough after comparing a synthetic spectrum with all lines with another one that only has the lines of the element in question. To achieve most accurate results, our analysis is based on the most recent line data, atmosphere grids, and grids to compute departures from the common LTE assumption during the line formation, as described in detail in Sec. 3.3. Below we describe this workflow in more detail, which illustrates the challenges of homogeneously analysing very different spectra:

Initialise sme (version 536) with choices of line data, atmosphere grid, non-LTE departure grids, observed spectrum (limited to the 46 segments These can be found in our online documentation. used for the parameter estimation) including selection of continuum and line masks, initial parameters for χ2\chi^{2} optimisation. Check if all external information is provided and then update the initial log⁡g\log g with this external information and the initial stellar parameters as outlined in the explanation of log⁡g\log g (Sec. 3.3).

Normalise all 46 segments individually with the chosen initial setup by fitting linear functions first to the observed spectrum (iteratively and with sigma-clipping) and then to the difference of the observed and synthetic spectrum.

Normalise all 46 segments again individually as in step 2, but with updated stellar parameters.

Normalise the segment(s) for the particular line (for the line-by-line analyses, e.g. Sr6550) or for all lines of the particular chemical species (e.g. Ca) with the chosen initial setup by fitting linear functions first to the observed spectrum. Improve this normalisation by fitting a linear function to the difference between the observed and synthetic spectrum to create a “full” synthetic spectrum.

Because the same line exhibits different degrees of blending in different stars, which are complex and difficult to predict ab-initio, perform a blending test by creating a “clean” synthetic spectrum only based on the lines of the element to be fitted. Then compare the “full” and “clean” spectra for elements other than Fe for the chosen line mask pixels and neglect those which deviate more than Δχ2>0.005\Delta\chi^{2}>0.005.

Collect stellar parameters and element abundances for validation and post-processing.

Calculate upper limits for each element/line for non-detections by estimating the lowest abundance that would lead to a line flux depression of 0.03 below the normalised continuum (see more detailed explanations at the end of this section).

Post-processing: apply flagging algorithms, calculate final uncertainties from accuracy and precision estimates, combine line-by-line measurements of element abundances weighted by their uncertainties.

For each star, the computational costs amount to between 50 CPU minutes (for the hottest stars with few lines), 2 CPU hours for the Sun, up to 6 CPU hours (for the coolest stars with most lines), with around 30-50% of that used for the stellar parameter step and the rest for abundance estimation for all lines. The total computational costs amount to 1.2 Mio CPU hours for the stellar parameter and abundance fitting, that is, neglecting data collection and post-processing.

3 Details of the spectroscopic analysis

are based on the corresponding compilation for the Gaia-ESO survey (Heiter et al. 2015a; Heiter 2020, Heiter et al. in press) with updated oscillator strengths (log⁡gf\log gf values) for some elements, in particular V i (Lawler et al. 2014), Cr i (Lawler et al. 2017), Co i (Lawler et al. 2015), Ni i (Wood et al. 2014) and Y ii (Palmeri et al. 2017). In addition, we “astrophysically” tuned (based on solar abundances and observations) the log⁡gf\log gf-values for approximately 100 lines that were not used for abundance measurements, but affected the continuum placement and blending fraction for the main diagnostic lines. The final compilation of the lines used for stellar parameter and element abundance estimation together with the most important line data are listed in Table 8.

are updated self-consistently with the other stellar parameters for each synthesis step via the fundamental relation of log⁡g\log g, stellar mass M\mathcal{M}, and bolometric luminosity LbolL_{\text{bol}}

from the 2MASS (Skrutskie et al. 2006) KSK_{S} band, a consistently calculated bolometric correction BC(KS)BC(K_{S}) for this band using stellar parameters for each synthesis step, distances Dϖ=r_estD_{\varpi}=\texttt{r\_est} from Bailer-Jones et al. 2018 as well as extinctions AKSA_{K_{S}} in the KSK_{S} band. If both 2MASS HH (Skrutskie et al. 2006) and WISE W2W2 (Cutri et al. 2014) band information have quality A (98% of our sample) we used the RJCE method (Majewski et al. 2011) to compute AKS=0.917⋅(H−W2−0.08)A_{K_{S}}=0.917\cdot\left(H-W2-0.08\right). For the remaining 2% of our sample we used AKS=0.38⋅E(B-V)A_{K_{S}}=0.38\cdot\texttt{E(B-V)} (Savage & Mathis 1979).

was selected based on the time and computation resources available. We have therefore measured only the following elements line-by-line (and report a value based on error-weighted means): \upalpha\upalpha (see next paragraph), Li, C, K, Ti, V, Co, Ni, Cu, Zn, Rb, Sr, Y, Zr, Mo, Ru, La, Ce, Nd, Sm, Eu. To estimate the abundance of the following elements, we ran combined all lines of the particular elements to fit a global abundance at the same time: O, Na, Mg, Al, Si, Ca, Sc, Cr, Mn, Fe, Ba. We want to stress that the use of non-LTE grids does not affect the computation time. We chose to run several elements in a combined way because of the ongoing developments of their line selection or non-LTE grids. During the development of the pipeline we have tested all individual lines for the elements run with non-LTE and only selected those with similar trends and absolute abundances to run combined. By using individual lines, we are less prone to unreliable line data, such as unreliable log⁡gf\log gf values. Incorrect oscillator strengths introduce a bias in the absolute abundance for each line. When the Solar abundance for these lines are however estimated independently from the others, they can still be used for the combined [X/Fe] abundance, after applying individual Sun-based corrections to the absolute abundances (see Table 9).

are calculated for all measured lines/elements if no detection was possible. In this case we estimate the smallest abundance needed to explain the strength of a line, that is the difference of line to continuum flux in the normalised spectrum of at least 0.03 or at least 1.5/(S/N)1.5/(S/N) in the line mask. We interpolate these values from precomputed estimates of line strengths for a set of stellar parameters and abundances. This approach was chosen and tested to estimate a larger number of upper limits for Li, but we want to caution the users to not blindly use them because we have not performed extensive tests for the other elements. They should only be used when essential for the science case and after thorough inspection of the observed and synthetic spectra.

Validation of stellar parameters

In this section, we describe the tests that we perform to validate the stellar parameters we obtain in terms of their accuracy (systematic uncertainties) and precision. In addition, we then describe several other algorithms that we have developed in order to identify peculiar stars or spectra - cases for which our standard pipeline might fail. The result of the performed quality tests are summarised in the stellar parameter flag flag_sp with flags that can be raised explained in Sec. 4.3. We do strongly recommend that all users take these flags into account, and make use of them unless the flags are explicitly not advisable for their particular science case. By default we recommend to use quantities with flag_sp 0, as this indicates that the stellar parameter estimates have passed all quality tests. Several influences on the accuracy, like unresolved binarity, as well as some possible systematics / caveats that we have not been able to quantify and therefore not flag, are addressed in Sec. 6.

To assess the quality of the stellar parameters we obtain, we resort to the commonly used comparison samples for accuracy, that is, the Sun (see our results for sky flat observations compared to literature in Table 11) and other Gaia FGK Benchmark stars (Heiter et al. 2015b; Jofré et al. 2014; Jofré et al. 2015; Hawkins et al. 2016; Jofré et al. 2018, GBS), photometric temperatures from the Infrared Flux Method (Casagrande et al. 2010, IRFM), stars with asteroseismic information, and open as well as globular cluster stars. For the precision assessment we use the internal uncertainty estimates and repeat observations of the same stars. We calculate the final stellar parameter errors for a given parameter PP via

The precision uncertainty eprecision2(P)e_{\text{precision}}^{2}(P) is usually quantified by either fitting uncertainty (efit2(P)e_{\text{fit}}^{2}(P)) In our case we use the square root of diagonal elements of the fitting covariance matrix to trace the fitting uncertainty (Piskunov & Valenti 2017). or uncertainty from repeated measurements (erepeats2(P)e_{\text{repeats}}^{2}(P)), which are typically expected to be of the same order. We compare and rescale these values to match in Sec. 4.2, so that we can use their maximum values for the individual measurements. Our repeat precision estimates are based on the behaviour with respect to our reference S/NS/N, that is snr_c2_iraf, and lead to our applied uncertainty estimation of

For the repeat observations, we make use of the 51 539 spectra of explicit repeat observations (typically at different nights) of stars, not the three individual observations scheduled for each star. Such repeat observations were mainly performed for the TESS-HERMES as well as bulge and cluster observations, but a smaller contribution comes from repeat observations of GALAH fields with bad seeing. Checks of the parameter distribution of the repeat observations and the overall sample suggest that they are representative of the sample.

Our effective temperatures are estimated from our spectra rather than photometry and because they correspond to the best-fit spectroscopic solution, we do report them rather than values calibrated to the photometric scale, but assess their accuracy.

1.2 Surface gravity

We find excellent agreement and negligible biases between our derived surface gravities and those from the GBS, as well as those obtained for stars with asteroseismic information. Due to the larger sample size of the stars with asteroseismic information, we apply the estimated scatter of this sample as an accuracy estimate for our log⁡g\log g.

1.3 Metallicity and iron abundance

1.4 Radial velocities

In contrast to the approach taken in GALAH DR2, we have estimated the radial velocities as a free parameter in the stellar parameter estimation and have thus been able to overcome a systematic trend of the reduction pipeline, which would otherwise have been overestimating the positive and underestimating the negative radial velocity by 1%1\%, respectively.

In version 2 of our data catalogues, we provide several different estimates for the community in a value-added-catalogue and add the values that we recommend to use for each spectrum in the column rv_galah We have updated this notation for convenience to allow the user to easily use our best estimate of radial velocities. In version 1 of the main catalogues, we reported only the SME based radial velocities, which used a erroneous barycentric correction, as we outline in this section..

We recommend users to consider the choice of radial velocity measurement for their specific science case. Our most accurate measurements are reported via rv_obst and are recommended (if available) for Galactic stellar dynamics studies. If the user wants to compare radial velocities with other studies, we recommend the use of rv_nogr_obst (if available) or rv_sme_v2 otherwise, as these do not correct for gravitational redshifts (and are currently the most commonly reported measurements by stellar spectroscopic surveys).

The source of the recommended, that is most accurate, values rv_galah is indicated via the flag use_rv_flag, as we outline in the catalogue description in Sec. 7.2.

2 Precision of stellar parameters

To estimate the precision of our stellar parameters, we use both internal sme covariance errors and repeat observations of the same star for all stellar parameters except for log⁡g\log g, for which we Monte Carlo sample the uncertainties.

This rescaled internal precision now allows for a combination of the individual estimate of the fit quality (through the internal sme-based uncertainty) with the general precision expected for a given S/NS/N, which could otherwise be underestimated when only using to the raw internal uncertainty.

We further note that we have changed our definition of S/NS/N in these figures compared to DR2 (Buder et al. 2018) to show the S/NS/N of the higher quality observation (DR3) instead of the lower one (DR2). The quantitative improvement of the precision from GALAH DR2 to GALAH+ DR3 is thus not necessarily an indicator of the decreasing precision, but of a more reliable precision estimate.

By construction and due to the exquisite astrometric and photometric external information available, this internal precision is significantly better than the previous spectroscopic estimates from GALAH DR2. We note, however, that these estimates do not take external influences like binarity or correlations of uncertainties into account.

Based on the open cluster membership analysis by Cantat-Gaudin & Anders 2020, we estimate that we have intentionally and unintentionally observed members of 75 stellar clusters. The eight open clusters with most observations are NGC 2682 (278 spectra, M 67), NGC 2632 (117, M 44, Praesepe), NGC 2516 (83), NGC 2204 (81), Ruprecht 147 (80), Melotte 22 (74), Blanco 1 (67), and NGC 6253 (50). Furthermore we have observed 10 of the 128 open clusters These are ASCC 16 (22 spectra), ASCC 16 (22), ASCC 21 (11), Berkeley 33 (8), Melotte 22 (74), NGC 2204 (81), NGC 2232 (20), NGC 2243 (8), NGC 2318 (2), NGC 2682 (278), and Ruprecht 147 (80). of the OCCAM survey (Donor et al. 2020), included as VAC from SDSS DR16 (Ahumada et al. 2020). The analysis of all open clusters observed with GALAH is addressed in the dedicated paper by Spina et al. 2020a. Here we only focus on three open clusters which cover a large range of stellar evolutionary stages and are also reported by the OCCAM survey The same plots as in Fig. 11 are available in our online documentation for all our clusters observed by GALAH..

3 Flagging of stellar parameters

After the stellar parameters have been estimated, we raise flags according to the individual criteria listed in Table 4. Fig. 14 shows the derived parameter values associated with all flagged spectra. The most used flags are 8 (8.6%), 1 (8.5%), 256 (8.0%), 4 (5.6%), 512 (2.4%), and 1024 (2.2%). Less than 2% of spectra have raised flags 2, 16, 32, 64, or 128.

Validation of element abundances

We validate element abundances in terms of accuracy and precision and discuss the two validations separately. Unlike for GALAH DR2, we are now not limited by the influence of the training set for this data release. We have also tried to estimate more upper limits and outline our approach in this section, followed by the description of our flagging algorithms with the aim to allow the community to make informed choices on the use of abundance measurements. As for the previous data releases, we want to stress that we discourage the use of flagged element abundances without consideration of the possible systematics that the inclusion of these flagged measurements can introduce.

For our accuracy studies (Sec. 5.1), the techniques we could adopt for comparisons with accurate benchmark are limited. Contrary to the stellar parameters, where multiple methods, and especially those which are independent of spectroscopy, are available for accuracy estimations, the available benchmarks for abundance accuracy are based on spectrocopy and - with exception of the Sun and Solar twins - also strongly limited in terms of accuracy (e.g. due to neglected 3D and non-LTE effects). We are therefore only using the Sun and Solar twins to validate the accuracy of our measurements to zeroth order (abundance zero points). We do not assess our abundance accuracy quantitatively beyond this, but we do also investigate systematic trends with respect to other spectroscopic studies presented in the literature. We want to stress that a proper estimation of the accuracy uncertainties would have to involve the systematic influence of the individual stellar parameters within their uncertainties, the uncertainties of the absolute abundances / zero points in terms of log⁡gf\log gf values and additional uncertainties from the fit to the sky flat, Arcturus, the comparison with the Solar circle sample, as well as the Solar twin comparison. For computational reasons we have not been able to quantify all of these influences, but report them if possible (see e.g. Tab. 9).

Our precision estimates for abundances show a similar behaviour to that of the fitting uncertainties efit(X)e_{\text{fit}}(X) and the S/NS/N-dependent repeat uncertainties erepeats(X,S/N)e_{\text{repeats}}(X,S/N), as we will explore in Sec. 5.2. Our reported final uncertainties are thus limited to a formula depending on element/line X and S/NS/N of CCD2

We estimate our element abundances by changing the absolute abundance for each element that is measured in the initially scaled-Solar chemical composition by Grevesse et al. 2007 of the model atmosphere. We convert the abundances to the customary astronomical scale for logarithmic abundances and report the raw values of these measurements, A(X1234) for the 1234 Å line of element X in the allspec catalogue (see Sec. 7.2). We subtract the Solar value A(X1234), that we define for this data release (see Tab. 9), from this measurement to estimate the ratio [X/H] and for elements other than Fe, we further subtract the iron abundance to estimate the ratio [X/Fe].

In addition to the definition of the abundance zero points, we also validated the accuracy of our element abundances by comparison to measurements of the GBS and Solar twin stars, as well as members of open and globular clusters.

Following the definition of the bracket notation, the Solar value A(X1234) should be strictly estimated from the measurement of the particular line in the Solar spectrum. For several lines within the GALAH wavelength range, however, we are facing difficulties in estimating the Solar A(X1234). Firstly, via 2df-HERMES, we can only perform sky (flat) observations rather than observing the Sun directly. Secondly, our observation setup is much shorter as for the normal setup of our observations. Thirdly, some lines are either not detectable (even within high S/NS/N spectra) or their equivalent width or line strength does not increase significantly with increasing A(X), that is, we perform a measurement at a plateau on the curve of growth. Contrary to many other studies or surveys, we choose to report the absolute abundances, and only use laboratory oscillator strengths (log⁡gf\log gf) rather than tuning these astro-physically based on the Solar spectrum with literature abundances. There are thus several solutions available to still estimate abundance zero points, which we will discuss subsequently:

Measure A(X) from the same line in Solar / sky flat / asteroid spectrum.

Use a different line in the Solar spectrum (because A(X) has to be the same).

Use another benchmark star (like Arcturus) via bridge measurements. To do this, one would measure the line that is weak in the Sun in Arcturus as well as another line of the element that is strong enough in both stars. In that case the difference in A(X) for the lines in Arcturus can be used to transfer them onto lines in the Sun.

Use the element abundance ratios of stars in the Solar circle, for studies suggest that the abundances should be Solar. APOGEE has followed this approach to estimate their zero points since their DR14 (Holtzman et al. 2018; Ahumada et al. 2020, see).

Compare with a literature study, e.g. via the estimates for Arcturus (Ramírez & Allende Prieto 2011, e.g.), Solar twins (Bedell et al. 2018, e.g.) or the overlap with a different survey, e.g. APOGEE DR16 (Ahumada et al. 2020).

For GALAH+ DR3 we try to use the first method whenever reliable and validate it using the other approaches. Whenever this approach was not advisable or the differences to the other methods were too significant (pointing to issues in the sky flat spectra), we used the fourth and fifth option to ensure the best overall consistency.

We want to stress that more work is needed to further scrutinise the line selection and abundance zero points. Due to time and computation restrictions during the implementation of the new non-LTE grids, we have only been able to run these elements combined, rather than line-by-line. However, we have found that the line-by-line analysis of element abundances is important for several elements (e.g. Al, Ca, and Ba), which are estimated from several different HERMES bands, and thus suffer from unreliable wavelength solutions in either band, and has to be done in future releases to improve the accuracy and precision of abundance measurements further.

For the validation of our abundance accuracy, we also turn to the best studied giant star, Arcturus, and the Gaia FGK benchmark stars. For Arcturus, we use the seminal study by Ramírez & Allende Prieto 2011, which is also used by APOGEE as reference. We list the values for Arcturus in Table 12, which in general show good agreement between our measurements and those of Ramírez & Allende Prieto 2011, both performed in the optical.

Similar to GALAH DR2, we could assess the element abundances for selected clusters with numerous observed members. These are, however, not useful in a straight forward manner to estimate the accuracy and precision of our measurements, due to internal processes like atomic diffusion and dredge-up changing the observed photospheric abundances for different evolutionary stages for open clusters (Gao et al. 2018; Bertelli Motta et al. 2018; Souto et al. 2018; Souto et al. 2019, see e.g.) as well as the presence of multiple populations in several globular clusters, leading to the spread in metallicities (Carretta et al. 2009b, see e.g.) and anti-correlations in several elements, like Na-O or Mg-Al (Carretta et al. 2009a; D’Orazi et al. 2010, see e.g.).

Since we lack good calibrators for these abundances, we choose to include overviews of several abundances for open clusters, as we expect the evolutionary effect within these to be seen, but predictable. Significant differences in abundances above the expected evolutionary effects are therefore indicative of systematic trends within our analysis. A more detailed analysis of abundances and trends in open and globular clusters will be performed in the studies by Spina et al. 2020a and D. M. Nataf et al. (in prep.) respectively, but we make the overview plots available in our online documentation. For the vast majority, our trends agree with the literature values, changesuch as the OCCAM survey (Donor et al. 2020), as shown in Fig. 16, where we plot the abundances of Si, Cr, Cu, and Ba for a selection of open clusters.

2 Precision of element abundances

We assess the precision of our element abundances by comparing the internal sme covariance uncertainties with those from repeat observations of the same star in Fig. 18. In contrast with the case for the stellar parameter estimation, we see that the covariance errors from the individual line measurements are typically in good agreement for almost all lines. The standard deviations of the measurements are also consistent irrespective of the fibre combination.

We note, however, that our final estimates of the internal sme-based uncertainties are lower than the those from the repeat observations, when a large number of lines is fitted in combination rather than line-by-line, in particular Fe, which we discussed in Sec. 4.2. This suggests that either the sme-internal method has problems in estimating realistic errors when many pixels are involved, or that our spectrum quality indicator (snr_c2_iraf) is not representative in those cases. With the exeption of Fe, we have however managed to estimate abundances of elements with many lines in a line-by-line manner, such that the fitting and repeat uncertainty show a similar behaviour. Unlike for the case of the stellar parameter estimation, we only report the final abundance uncertainty based on the maximum of the raw internal covariance error for the abundance fit and the S/NS/N-scaled repeat observation uncertainty.

3 Flagging of element abundances

As for all of our previous GALAH releases, we want to stress that we discourage the use of flagged element abundances without consideration of the possible systematic trends that these probably less reliable measurements can introduce. However, setting flags for element abundance measurements is extremely difficult and will never be perfect. In this section we lay out which flags we have implemented. We also point to identified possible caveats, where no flags are raised, but caution should be applied.

The major abundance flags are reported as flag_X_fe for each element X and are listed in Table 4. For the final element abundances, we only report measurements that have passed all our quality checks with no flag raised or are an upper limit (with flag 1, see our description of upper limits in Sec. 3.3). For ease of use we have raised a flag value 32 if the measurement is not reported, because it did not pass the quality checks.

As we elaborate in the next section, we have identified several possible caveats, for which we have not implemented an algorithm to raise a flag, either because we could not find a way to automatically identify the caveats, or we are unsure if the measurements should be flagged. For the abundances, we therefore advice to carefully consider the caveats described in Secs. 6.4 and 6.5 in particular.

Possible caveats: analysis shortcoming or physical correlation?

In the previous sections we have laid out the methods by which we flag unphysical results and spectra with peculiarities for which our pipeline is likely to underperform. However, we cannot visually inspect all of the more than 30 million measurements that have been performed for this data release. Furthermore, we aim to not follow up all of the possible correlations in full detail, because many of these pose problems to understand the possible astrophysical nature of these trends (as shown later in Sec. 6.3 for atomic diffusion causing systematic differences of surface abundances in open cluster stars). Instead, we choose to leave such efforts for future scientific follow-up.

2 Stacked spectra

3 Radial velocities in table version 1

4 Possible systematic trends

Below we present a list of possible systematic trends, as found up until the publication of GALAH+ DR3 during the validation. These do not appear in particular order.

For V, we caution the use of measurements with nr_V_fe 2 and 3, that is, using V i 4832. For Co, we caution the use of measurements with nr_Co_fe = 2 or 8, that is, measurements purely based on lines Co i 6490 and 7713. While we have not been able to narrow down the exact cause, we assume that measurements only based on these lines are caused by imperfect telluric corrections in CCD 3 for Co i 6490 and spikes or imperfect telluric corrections in CCD4 for Co i 7713. Such reduction issues introduce strong emission and absorption lines, which will either weaken the line beyond detectibility or strengthen it so that only high V or Co abundances can reproduce the observation. However, the resultant fits and their χ2\chi^{2} values will not appear suspicious and such reduction issues are therefore hard to identify and flag automatically.

While our long-term goal is to implement 3D non-LTE calculations, we believe that it is worth testing the implementation of vmicv_{\text{mic}} as a free parameter or the relations estimated by Dutra-Ferreira et al. 2016 for certain parts of the parameter space, if the advantages outweigh the loss in abundance precision. Using vmicv_{\text{mic}} as a free parameter showed for example significant improvements of trends with TeffT_{\text{eff}} for the APOGEE survey (Holtzman et al. 2018).

For computational reasons, we estimate the abundances of all elements independently, and assume scaled-Solar patterns for most other elements during that optimisation. However, our approach might introduce systematic trend for elements which are often correlated (e.g. C and O), surrounded by lines that are deviating from the scaled-Solar pattern, or when the abundance pattern in general differs from the scaled-Solar pattern, thus leading to differences in the continuum and molecular lines strengths. If computationally possible, it would therefore be preferable to fit all elements partially (Brewer et al. 2016) or fully self-consistent (Ting et al. 2019), which could also allow to estimate abundances not only via atomic lines, but also molecular features, which follow molecular equilibria (Ting et al. 2018).

Stellar clusters are not the main focus of our survey, and many of the observations that were performed for them are outside of the typical GALAH magnitude, distance, and age range. Most of the open and globular clusters targeted by our observations are much more distant, which leads to less reliable distance estimates, with implications for our distance-dependent log⁡g\log g estimates of their stars. Many of stars in the open clusters stars are typically younger than the GALAH targets, with astrophysical implications on additional features in their spectra.

A central assumption of our observations is that each fibre observes only one star. We try to ensure this by only selecting point sources from 2MASS with a sufficient separation from other bright neighbours. Our selection does however not exclude stars that are not extended within 2MASS, for example spectroscopic binaries.

As part of our spectroscopic analysis we rely on the quality of astrometric measurements, to infer reliable absolute photometry and then log⁡g\log g. While we flag stars with high RUWE values above 1.4 (Lindegren et al. 2018; Lindegren 2018), we caution the user to not blindly use all measurements, especially those of stars with uncertain astrometry.

We have used more elaborate distance estimates from Bailer-Jones et al. 2018 which infer more trustworthy distances based on a Galactic prior for stars with parallax uncertainties beyond 20%20\%. Especially for very distant stars, like some of our observations of LMC stars, this Galactic prior leads to an underestimated distance and thus likely overestimated log⁡g\log g (see Eqs. 1 and 2).

In general, we note that for stars with more constrained distance estimates, like open clusters (Cantat-Gaudin & Anders 2020), globular clusters (Baumgardt et al. 2019) and stars of the LMC (de Grijs et al. 2014), a reanalysis would be leading to more reliable stellar parameters and abundances, when using these distances instead of the ones solely estimated from Gaia parallaxes.

Potassium is estimated from the K i 7699 resonance line. This line is also a good tracer of interstellar potassium which leads to contamination of the stellar line in highly extinct regions. In the future we aim to estimate the extinction for example via diffuse interstellar bands and possibly use correlation of extinction and line strength of interstellar potassium (Munari & Zwitter 1997) to correct the spectra and measurements of stellar [K/Fe]. For this DR, we however caution the user to check the extinction of stars when using [K/Fe], as we measure this abundance without any corrections causing a rather hard to predict systematics (depending on the velocities of star and ISM) of [K/Fe].

While we report upper limits for advanced users, we strongly recommend that all users take great care in using these measurements. For all elements, but especially for neutron-capture elements, these estimates are pushing the limits of what we can be extracted from the data and are by definition only an upper limit, not a measurement. We therefore strongly recommend to check upper limit estimates against the data and inspect spectra when possible.

5 Peculiarities for certain groups of stars

For some groups of stars, we have found peculiar trends of abundances, for which we cannot exclude astrophysical reasons rather than influences of our analysis and suggest follow-up studies to disentangle those.

6 Uncertainties

For this data release, we include more accuracy and precision estimates than for GALAH DR2. However, for several stellar parameters and abundances, the means of accuracy estimation are limited, because there are no benchmark values available. We therefore want to caution the user that the accuracy uncertainties might be underestimated and also not complete in terms of their parameter dependence.

Catalogues included in this release

We have published two versions of our data catalogues, version 1 before Gaia eDR3 and version 2 after Gaia eDR3. We recommend the use of version 2, as it includes the following updates and fixes:

We have crossmatched our targets with the Gaia eDR3, parallax zero point corrections by Lindegren et al. 2020a and (photo-)geometric distance estimates from Bailer-Jones et al. 2020. We provide all this information in a new VAC (see Sec. 7.3.1).

We have propagated the information from Gaia eDR3 to estimate better added values for the VACs on stellar ages as well as dynamics.

2 Main catalogues

The main catalogues can be downloaded from the DataCentral https://cloud.datacentral.org.au/teamdata/GALAH/public/. and accessed via TAP https://datacentral.org.au/vo/tap..

We provide two main catalogues. The first one allstar includes one entry per star and is a cleaned version of the extended allspec, with each entry representing the highest S/NS/N measurement for each star and only the combined, final abundance estimates that are unflagged or upper limits.

The allspec catalogue includes an entry for each spectrum (multiple entries for some stars). It extends the allstar catalogue by also having the raw stellar parameter and abundance measurements, that is, raw measurements without zero point or bias corrections and uncalibrated fitting uncertainties (cov_*) for each stellar parameter and individual line abundances (ind_*), with more extensive flags.

The flags for both the main stellar parameters (flag_sp) and the final and raw abundance measurements are listed in Table 4 and explained in Secs. 4.3 and 5.3, respectively.

For illustration we plot the distribution of all element abundances in Fig. 21.

After the publication of Gaia EDR3, we have added dr3_source_id to the main catalogues and updated the reported rv_galah and e_rv_galah as well as rv_gaia and e_rv_gaia, as described in Sec. 7.3.4. The new versions of the catalogues can be found via their suffix v2.

We list the table schema of the catalogue in Table 13, but they can also be found in the FITS header or online at https://datacentral.org.au/services/schema/. It includes the following categories for the allstar and allspec catalogues:

Stellar parameter flags (both warning and flags)

Combined alpha-abundance (for unflagged/upper limit measurements, see Sec. 3.3)

Combined element abundances (for unflagged/upper limit measurements, see Sec. 3.3) and bitmask of the line selection

Most important products of the reduction pipeline

Crossmatches with Gaia DR2, Bailer-Jones et al. 2018, Gaia DR2 RUWE, 2MASS,

For the allspec catalogue, it also includes

Individual element abundances (including flagged measurements)

More products of the reduction pipeline WISE W2W2

3 Value-Added-Catalogues (VACs)

This data release of GALAH is accompanied by four Value-Added-Catalogues, one for a cross-match with Gaia EDR3, one for stellar ages and masses, one with kinematic as well as dynamic information for each star/spectrum, one for more elaborate radial velocity estimates, and a fourth one with additional estimates for double-lined spectroscopic binaries. After the publication of Gaia EDR3, we have updated all our catalogues and their latest versions have the suffix v2. Subsequently, we outline the important information used to calculate the additional values for each VAC and note the changes from v1 to v2.

As of December 2020, we provide the crossmatch of our allspec-catalogue with Gaia EDR3 (Gaia Collaboration et al. 2020; Fabricius et al. 2020), with improved astrometric (Lindegren et al. 2020b) and photometric (Riello et al. 2020) information. We performed this match via the dr2_source_id, specifically with the smallest angular distance In the future, we suggest to perform this crossmatch via GALAH’s 2MASS ID and the yet-to-come match of Gaia EDR3 and 2MASS identifiers. We also note that for 57 sources, a sorting via Gaia G magnitude would have yielded different identifiers. within the internal crossmatch catalogue of Gaia DR2 and Gaia EDR3 (Torra et al. 2020). For each star, we further list the parallax zero point as queried from Lindegren et al. 2020a to correct the parallaxes of the sources with 5 or 6-parameter astrometric solutions from Gaia EDR3 (Lindegren et al. 2020b). We further provide the photogeometric and geometric distance estimated by Bailer-Jones et al. 2020 which were matched via the dr3_source_id. The table schema of the VAC is included in the FITS file but can also be found at https://datacentral.org.au/services/schema/.

3.2 Stellar age and mass estimates

by Salaris & Cassisi 2006, with Z⊙=0.0152Z_{\odot}=0.0152 in accordance with the isochrones used. This was compared with the surface metallicity reported by the isochrones, which takes the evolutionary changes in surface metallicity ZZ into account. The code provides an estimate of age, actual mass, initial mass, initial metallicity, surface metallicity, radius, distance, extinction E(B−V)E(B-V), luminosity, surface gravity, temperature and the probability of being a red clump star. For each estimated parameter we report a mean value and standard deviation together with the 16th, 50th, and 84th percentiles. The isochrone grid consisted of 16 768 422 grid points. An 81×12181\times 121 grid spanning −2<log⁡Z/Z⊙<0.5-2<\log Z/Z_{\odot}<0.5 and 6.6<log⁡age/Gyr<10.126.6<\log{\rm age/Gyr}<10.12 was used. The mass dimension of the grid was resampled by interpolation, such that Δlog⁡Teff<0.004\Delta\log T_{\rm eff}<0.004 and Δlog⁡g<0.01\Delta\log g<0.01. For each parameter we report a mean value and standard deviation based on 16 and 84 percentiles. After the publication of Gaia EDR3, we have recalculated all entries in this VAC (provided in the table with suffix v2) based on new Gaia EDR3 parallaxes and their uncertainties, corrected by the zero point shifts by Lindegren et al. 2020a. The table schema of the VAC is included in the FITS file but can also be found at https://datacentral.org.au/services/schema/.

3.3 Kinematic and dynamic information

We provide a Value-Added-Catalogue with kinematic and dynamic information, that builds upon the 5D astrometric information by Gaia DR2 and radial velocities preferably from GALAH+ DR3 (97.3%) of all spectra and otherwise from Gaia DR2 (0.5% of all spectra). Where possible, we use distance that take into account uncertainties, preferably those estimated via isochrone matching as part of the BSTEP grid-based modelling (95.7% of all spectra, see description in Sec. 7.3.2), otherwise we use the prior-informed values (4.1% of all spectra) by Bailer-Jones et al. 2018.

For the calculation of orbit information we use version 1.6 of the python package galpy (Bovy 2015). More specifically, we use its orbit module for coordinate/velocity transformation as well as orbit energy computation. To estimate actions, eccentricity, maximum orbit Galactocentric height, and apocentre/pericentre radii, we use galpy’s actionAngleStaeckel approximations via the Staeckel fudge (Binney 2012) with a focus of 0.45 and the method by Mackereth & Bovy 2018.

For the computed phase space and dynamic properties, we report a variety of statistical values. In addition to the best-value, that is computed by using the best values as input, we also sample the distribution for each property within the uncertainties via Monte Carlo sampling with size 1000 and report the 5th, 50th, and 95th percentiles of these distributions. An example of the sampling of parameters for 100 randomly selected stars is shown in Fig. 23. We also provide the code to perform this sampling with different sampling choices. Whereas we currently sample the properties by assuming their input parameters are uncorrelated, we also provide the code to sample with the Gaia correlation matrices. The latter are currently not applying a distance prior and are thus problematic for large distances. However, we stress, that the vast majority of the stars from GALAH+ DR3 have very precise parallax measurements, for which the sampling choice is negligible (see Fig. 3).

3.4 Radial velocities

3.5 Double-lined spectroscopic binary stars

GALAH+ DR3 in context

To understand how we can use the available surveys on a global scale, two key points need to be considered. Firstly, if the surveys are complementary and secondly if the surveys are on the same scale.

While a detailed comparison of overlapping stars between the surveys would be very helpful to assess systematic trends, it should be performed in a dedicated study and include comparisons of stellar parameters as well as abundances for different parameter selections. Some basic comparisons can be found in our open-source repository, but a detailed study is beyond the scope of this paper.

In this section, we are rather aiming to give explanations why different surveys may not depict the same trends on first glance (based on different selection functions), while highlighting the potential of combining the different surveys.

GALAH+ DR3 includes 678 423 combined spectra of 588 571 stars, obtained at high-resolution (28 000) in 4 narrow optical bands (covering 1000 Å). APOGEE DR16 includes 473 307 combined spectra of 437 445 stars, obtained at high-resolution (22 500) in the H-band (15 000-17 000 Å). LAMOST DR5 VAC includes 8 162 566 combined spectra of 6 091 116 stars, obtained as low-resolution (1,800) in the full optical range (4000-9000 Å). It is important to note that this VAC is estimated by data-driven models trained on GALAH DR2 and APOGEE DR14.

The overlap of GALAH+ DR3 and APOGEE DR16 is 15 047 stars, that is 3% of the each survey. The overlap of GALAH+ DR3 and LAMOST DR5 is 47 118 stars, that is 8% and 1% of the respective survey. The overlap of APOGEE DR16 and LAMOST DR5 is 111 626 stars, that is 26% and 2% of the respective surveys.

These numbers show that the surveys are very complementary in the stars that they target, but also have a non-negligible overlap between them. Even more important, this overlap allows us to test if these surveys are on the same scale and even to cross-calibrate them to bring them on the same scale (Ho et al. 2017; Casey et al. 2017; Wheeler et al. 2020; Xiang et al. 2020; Nandakumar et al. 2020, see e.g.).

For the subsequent comparison we limit the samples to those stars with flag_sp = 0 , flag_fe = 0, and flag_alpha_fe = 0 for GALAH, ASPCAPFLAG = 0 for APOGEE, and FLAG_SINGLESTAR = 0, QFLAG_CHI2 = “good” as well as SNR ratios for at least 30 for either G, R, or I for LAMOST’s DR5 VAC.

The elemental abundances obtained by these surveys are data of particular interest for Galactic archaeology. A detailed comparison of those between the surveys is beyond the scope of this paper. Typically, different surveys operate at different resolutions and reach different S/NS/N in different wavelength regions, thus selecting different lines for their analyses. Different lines again, can form at different optical depths and may be blended differently; all possible factors for possibly different abundance measurements (Jofré et al. 2019).

2 Galactic and stellar chemical evolution

In this section we briefly aim to show the potential of GALAH+ DR3 for the exploration of Galactic and stellar chemical evolution, while leaving the true exploration to the scientific community. One would ideally like to take all abundance measurements into account for such an endeavour, including GALAH’s main goal of the chemical tagging experiment, but here we aim to show how much potential the exploration of a single element has to offer.

3 Chemodynamical evolution

To assess the potential of GALAH+ DR3 in terms of exploring the chemodynamic evolution of the Milky Way, we show the distribution of the data in plots that have been used in seminal studies for Galactic exploration.

Although halo stars are not the main target of GALAH, roughly 1% of all GALAH targets are expected to belong to the chemical or kinematic halo (De Silva et al. 2015). While the definition of halo stars is contentious, we at least aim to assess their rough number by looking at different kinematic and dynamic properties. For this, we look at the distribution of azimuthal / transversal velocity VTV_{T} with respect to the combined radial and vertical velocity VR2+Vz2\sqrt{V_{R}^{2}+V_{z}^{2}} in Fig. 22.

That these stars are not only different in their dynamics, can be seen in a chemical projection in Fig. 27c, where we follow up the distinct chemical signatures of accreted halo stars as found by Nissen & Schuster 2010 and Nissen & Schuster 2011 in the projections similar to those proposed by Hawkins et al. 2015 and Das et al. 2020. When assessing different nucleosynthesis channels via different elements, that is Al or Na, \upalpha\upalpha like Mg, and Cu or Mn, the accreted halo stars clearly stick out as a distinct overdensity because of their different chemical enrichment history compared to the majority of the Milky Way disc stars. A follow-up of these findings will be presented in the chemodynamical study of accreted halo stars by S. Buder et al. (in prep.).

Conclusions and outlook

With this third data release of the Galactic Archaeology with HERMES (GALAH) survey, we are providing the most complete set of information in terms of chemical composition, dynamics, and stellar ages to the public. The new data provides abundances for up to 30 elements and with the additional astrometric information provided by the Gaia satellite, we are able to estimate very precise orbits for almost all stars. This data is extremely valuable for different disciplines of astrophysics and will bring together observers with theorists.

In this manuscript, we describe the methodology behind the newly released data. This release incorporates data from GALAH’s partner surveys, namely the K2/HERMES and TESS-HERMES surveys, yielding a total sample of 678 423 spectra for 588 571 stars. Regarding the use of our data, we conclude:

Use version 2 of our data catalogues. They include fixes to previously reported erroneous values for the radial velocities rv_galah and a new VAC for the crossmatch with Gaia eDR3 (see Sec. 7.1 and references therein for more details) updates to the VACs based on the new Gaia data.

Use only unflagged measurements of our main stellar parameter quality flag (flag_sp=0\texttt{flag\_sp}=0).

Use only unflagged abundance measurements for each element X (flag_X_fe=0\texttt{flag\_X\_fe}=0). Even in this case, be careful of suspicious abundances and trends, which may not have been caught by our flagging algorithms. Systematic trends may be introduced through inherent limitations of the analysis pipeline. Please consult Sec. 6 on caveats, including high abundances of V, Co, Rb, Sr, Zr, Mo, Ru, La, Nd, and Sm.

We report both individual as well as stacked visit spectra. The latter result in higher S/NS/N and are generally the preferred values reported per star in the allstar catalogues. For studies of individual stars we recommend, however, to also compare all results of individual spectra of the stars from the allspec catalogues to confirm that the automatic stacking process was successful (see Sec. 6.2 for details).

2 Scientific avenues for the use of GALAH DR3

Since the advent of galactic archaeology (Freeman & Bland-Hawthorn 2002, and references therein), many large stellar surveys attempt to establish a narrative for the Galaxy by comparing vast amounts of stellar data (ages, kinematics, chemistry) to cosmological N-body ++ hydrodynamic simulations (Kobayashi & Nakasato 2011; El-Badry et al. 2018a; Buck et al. 2019, e.g.). These comparisons assume the present-day Milky Way to be an axisymmetric system in dynamical equilibrium where measurables can be expressed as a function of Galactocentric radius, RR (Sharma et al. 2011).

Astronomers have long suspected there is much to learn from examining dynamical perturbations and their dependence on the stellar properties (Minchev et al. 2009; Widrow et al. 2012). We refer this field as Galactic seismology and identify it as a subset of Galactic archaeology. Bland-Hawthorn & Tepper-Garcia 2020 provide a brief history of Galactic seismology dating back to its first use in 1985. Indeed, in the Gaia second data release (DR2) just two years ago (Antoja et al. 2018), a remarkable signature of incomplete phase-mixing was uncovered. If we consider a Galactic cylindrical coordinate frame defined by (R,ϕ,z)(R,\phi,z), with velocity components (VR,Vϕ,Vz)(V_{R},V_{\phi},V_{z}), the Gaia team discovered a “phase spiral” in the z−Vzz-V_{z} plane. The vertical (zz) oscillation frequency is anharmonic so this signal arises from a corrugated wave propagating across the Galactic disc. GALAH has been used to study this phenomenon in terms of stellar ages, actions and abundances (Bland-Hawthorn et al. 2019; Laporte et al. 2019), with further analyses already under way.

The exact formation of the halo and disc remains enigmatic, but the progress of cosmological simulations is now allowing us to by comparing properties like the chemical bimodality of the Milky Way’s stellar disc those of simulated galaxies (e.g. Buck 2020; Vincenzo & Kobayashi 2020, and references therein).

The large amount of stellar data provided by stellar spectroscopic surveys is bringing together expertise of previously independent research. Based on our data, the exoplanet community improves our understanding of exoplanets through their host stars with improved stellar parameters (Clark et al. 2020, e.g.), more realistic input for planet formation simulations (Bitsch & Battistini 2020, e.g.) and will be able to explore exoplanet host stars in a chemo-kinematic or -dynamic context (Carrillo et al. 2020, see e.g.).

With the publication of the reduced spectra, we are going another step towards an open data community. Using the spectra will allow other scientists to not only verify our results, but also apply their analyses techniques for parts of the parameter space, for which our own pipeline is not optimised, e.g. the analysis of very hot stars, emission line stars, or very cool stars, among others. Furthermore does the publication of the spectra allow scientists to apply machine learning or clustering algorithms onto the data (Price-Jones & Bovy 2019, see e.g.).

Starting from such studies of spatially and dynamically bound groups of stars or solar twins, we are just at the beginning of understanding the correlations for field stars between abundances and orbits (Coronado et al. 2020, see e.g.) as well as the abundances and ages for field stars (Morel et al. 2020; Hayden et al. 2020; Sharma et al. 2020, see e.g.) as well as their limitations (Feltzing et al. 2017; Ness et al. 2018, see e.g.).

3 GALAH DR4 and a sharper focus with GALAH Phase 2

We have learned several lessons in the analysis for this data release, which will help us to improve our analysis in the future, that is for the fourth data release of GALAH and thereafer. We have found several interesting trends, of which some are likely astrophysical, while others are not. We will follow these up in the future to hopefully minimise the unphysical trends. Several of these are likely to be addresses by improvements in the reduction of spectra with improved telluric corrections and improved stacking routines. While an in-depth comparison of the data-driven vs. model-driven approaches is still to be conducted, first results from our work indicates that a quadratic model reaches its limitations when used to describe a very high-dimensional space, covering the stellar parameters along A-M type stars, as well as 30 element abundances. With the introduction of more higher-order models or flexible models and methods, for example neural networks or Gaussian process regression in stellar spectroscopy (Ting et al. 2019; Wang et al. 2020), we believe that such limitations can be overcome and will allow to further overcome the significant computational costs of on-the-fly spectrum synthesis. A major limitation of all spectroscopic analyses remains with the immensely uncertain oscillator strengths used to create synthetic spectra, with significantly more effort needed. In the future, we aim to not only use improved synthetic model grids, based on 3D non-LTE computations, but also implement these more sophisticated interpolation routines combined with an Bayesian framework. The latter would allow us to include non-spectroscopic information like information from Gaia eDR3 and DR3 in a probabilistic way and help us assess the uncertainties of our estimates more reliably.

One of the most limiting bottlenecks of Galactic archaeology are the still significant uncertainties of stellar ages which can be estimated to no better than 10% (Soderblom 2010), but are typically significantly higher. With the start of GALAH Phase 2, for which we adjust our target selection to observe more main-sequence turn-off stars to get more reliable age estimates, we also adjusted our observing strategy with longer exposure time to achieve higher spectral quality (and thus higher accuracy and precision). These adjustments will help us to more efficiently collect high-dimensional data of stars in our Solar vicinity and provide the community with a promising data set of chemistry, dynamics, and reliable ages.

Acknowledgements

Based on data acquired through the Australian Astronomical Observatory, under programs: A/2013B/13 (The GALAH pilot survey); A/2014A/25, A/2015A/19, A2017A/18 (The GALAH survey phase 1), A2018 A/18 (Open clusters with HERMES), A2019A/1 (Hierarchical star formation in Ori OB1), A2019A/15 (The GALAH survey phase 2), A/2015B/19, A/2016A/22, A/2016B/10, A/2017B/16, A/2018B/15 (The HERMES-TESS program), and A/2015A/3, A/2015B/1, A/2015B/19, A/2016A/22, A/2016B/12, A/2017A/14, (The HERMES K2-follow-up program). We acknowledge the traditional owners of the land on which the AAT stands, the Gamilaraay people, and pay our respects to elders past, present, and emerging.

This work has made use of data from the European Space Agency (ESA) mission Gaia (http://www.cosmos.esa.int/gaia), processed by the Gaia Data Processing and Analysis Consortium (DPAC, http://www.cosmos.esa.int/web/gaia/dpac/consortium). Funding for the DPAC has been provided by national institutions, in particular the institutions participating in the Gaia Multilateral Agreement. This publication makes use of data products from the Two Micron All Sky Survey, which is a joint project of the University of Massachusetts and the Infrared Processing and Analysis Center/California Institute of Technology, funded by the National Aeronautics and Space Administration and the National Science Foundation. This work was supported by the Australian Research Council Centre of Excellence for All Sky Astrophysics in 3 Dimensions (ASTRO 3D), through project number CE170100013 and the Swedish strategic research programme eSSENCE. This work was supported by computational resources provided by the Australian Government through the National Computational Infrastructure (NCI) under the National Computational Merit Allocation Scheme (project y89).

The following software and programming languages made this research possible: IRAF (Tody 1986; Tody 1993), configure (Miszalski et al. 2006), topcat (Taylor 2005, version 4.4;); Python (version 3.7) and its packages astropy (Astropy Collaboration et al. 2013; Astropy Collaboration et al. 2018, version 2.0;), scipy (Jones et al. 2001), matplotlib (Hunter 2007), pandas (Mckinney 2011, version 0.20.2;), NumPy (Walt et al. 2011), IPython (Pérez & Granger 2007), and galpy (Bovy 2015, version 1.3;). This research has made use of the VizieR catalogue access tool, CDS, Strasbourg, France. The original description of the VizieR service was published in A&AS 143, 23. This research mad use of the TOPCAT tool, described in Taylor 2005. This publication makes use of data products from the Two Micron All Sky Survey, which is a joint project of the University of Massachusetts and the Infrared Processing and Analysis Center/California Institute of Technology, funded by the National Aeronautics and Space Administration and the National Science Foundation.

SB and KL acknowledge funds from the Alexander von Humboldt Foundation in the framework of the Sofja Kovalevskaja Award endowed by the Federal Ministry of Education and Research. SB acknowledges travel support from Universities Australia and Deutsche Akademische Austauschdienst. JK and TZ acknowledge financial support of the Slovenian Research Agency (research core funding No. P1-0188) and the European Space Agency (PRODEX Experiment Arrangement No. C4000127986). AMA acknowledges support from the Swedish Research Council (VR 2016-03765), and the project grants “The New Milky Way” (KAW 2013.0052) and “Probing charge- and mass- transfer reactions on the atomic level” (KAW 2018.0028) from the Knut and Alice Wallenberg Foundation. KL acknowledges funds from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation programme (Grant agreement No. 852977) KCF acknowledges support from the the Australian Research Council under award number DP160103747. LS acknowledges financial support from the Australian Research Council (discovery Project 170100521) and from the Australian Research Council Centre of Excellence for All Sky Astrophysics in 3 Dimensions (ASTRO 3D), through project number CE170100013. CK acknowledges funding from the UK Science and Technology Facility Council (STFC) through grant ST/ R000905/1, and the Stromlo Distinguished Visitorship at the ANU. DMN acknowledges support from NASA under award Number 80NSSC19K0589, and support from the Allan C. And Dorothy H. Davis Fellowship. YST is supported by the NASA Hubble Fellowship grant HST-HF2-51425 awarded by the Space Telescope Science Institute. GT was supported by the project grant “The New Milky Way” from the Knut and Alice Wallenberg foundation and by the grant 2016-03412 from the Swedish Research Council. GT acknowledges financial support of the Slovenian Research Agency (research core funding No. P1-0188 and project N1-0040). MŽ acknowledges funding from the Australian Research Council (grant DP170102233).

We thank Chris Onken for crossmatching GALAH and SkyMapper DR3 data (see Onken et al. 2019, for more information). We thank Kareem El-Badry for providing a vetted version of wide binaries from Gaia DR2 data and GALAH+ DR3 radial velocities. We thank Alvin Gavel and Andreas J. Korn for bringing the bug in SME 536 to our attention. We thank Scott Trager and many other colleagues who started a discussion about acknowledging the pioneering work by Beatrice M. Tinsley. We thank the SDSS collaboration and APOGEE team for their efforts in documentation, which greatly benefited our efforts in documentation and validation.

Data Availability

The data underlying this article are available in the Data Central at https://cloud.datacentral.org.au/teamdata/GALAH/public/GALAH_DR3/ and can be accessed with the unique identifier galah_dr3 for this release and sobject_id for each spectrum. For more information (including the single object viewer options and bulk downloads) we refer the reader to the Data Central documentation at https://docs.datacentral.org.au/galah/dr3/overview/.

References

Appendix A Linelist, reference values, and table schema of the main catalogue

Here we append all additional information used for the analysis.