The Payne: self-consistent ab initio fitting of stellar spectra

Yuan-Sen Ting, Charlie Conroy, Hans-Walter Rix, Phillip Cargile

I. Introduction

Large-scale multiplexing spectroscopic surveys are revolutionizing the quality and quantity of spectroscopic data for Galactic archaeology. Surveys such as APOGEE (Majewski et al. 2017), GALAH (De Silva et al. 2015) and Gaia-ESO (Smiljanic et al. 2014) are collecting high-quality spectra for 105−10610^{5}-10^{6} stars with a spectral resolution R≃25R\simeq 25,000000, orders of magnitudes more stars than previous samples. Lower-resolution spectroscopic surveys, e.g., RAVE (Steinmetz et al. 2006), Gaia-RVS (Recio-Blanco et al. 2016), and LAMOST (Luo et al. 2015), are collecting even larger samples. And upcoming spectroscopic surveys, such as DESI (DESI Collaboration et al. 2016), 4MOST (de Jong et al. 2014), WEAVE (Dalton et al. 2016), MOONS (Cirasuolo et al. 2014), SDSS-V (Kollmeier et al. 2017), will boost sample sizes at both high and low spectral resolution by another order of magnitude, towards ∼107\sim 10^{7} stars.

However, learning about Galactic archaeology and stellar physics from these spectra depends crucially on our ability to correctly and precisely infer numerous stellar labels from these spectra: stellar parameters and individual elemental abundances. This requires a rigorous method to extract the maximal information from these data, based on physical ab initio spectral models. This is the focus of this study.

A key to rigorous fitting of stellar spectra is the ability to fit all stellar labels (typically >20−50>20-50 for stellar spectra) simultaneously (Ting et al. 2016; Rix et al. 2016), principally for two reasons: the spectral features of many elements are blended in the spectrum, imprinting a covariant signature on the data. And for quite a number of elements, variations in their abundances not only affect the strength of their spectral features, but also alter the stellar atmospheric structure (Ting et al. 2016); this in turn affects the spectral features of other elements, especially in cooler stars. Therefore, spectral modeling should be based on self-consistently calculated models that take into account the dependence of the atmosphere structure on various element abundances. This dependence is widely implemented for changes in [Fe/H], but not other elements.

In practice, current spectral analyses often fit only small portions of the spectrum to determine any particular element abundance, holding the abundances of other elements fixed. And they often require subsequent recalibration of the basic stellar parameters (e.g., log⁡g\log g and TeffT_{\rm eff}) or abundance-TeffT_{\rm eff} trends inferred from the spectral fitting. This motivates the need for the development of a comprehensive approach to study these issues. Here we will present such a method, The Payne In appreciation of Cecilia Payne-Gaposchkin’s ground-breaking work on physical spectral models. in this study.

The Payne combines a number of important ingredients: a set of spectral models based on a state-of-the-art line list (Cargile et al., in prep.); models computed that are self-consistently calculated for each set of labels; a robust and flexible “interpolator” in the high-dimensional label space for spectral fitting that can precisely predict spectral model fluxes for arbitrary sets of labels; a well-defined and objective assessment and mitigation of the wavelength regions where the models have important systematic shortcomings; and a robust estimate of the label estimates from the entire remaining parts of the observed spectra. For modeling stellar spectra, The Payne is a fully automated, simple, transparent fitting machinery, given a set of ab initio synthetic spectral models. The codes for running The Payne are publicly available online. Moreover, the fitting is very efficient – e.g., fitting 25 labels for an APOGEE spectrum with The Payne takes less than 1 CPU second. The Payne differs from The Cannon (Ness et al. 2015; Casey et al. 2016) principally in two respects: it is based on physical instead of data-driven models, and it generalizes the “interpolator” beyond the quadratic polynomial implemented in The Cannon and Rix et al. 2016. In short, The Payne is an approach that for the first time combines all these ingredients, necessary for progress towards optimal modelling of survey spectra; and it leads to both precise and accurate estimates of stellar labels, based on physical models and without ‘re-calibration’.

This paper is structured as follows: we introduce The Payne and test the interpolator at its core in Section II. We apply The Payne to the APOGEE DR14 data set in Section III, and present the resulting catalog. We discuss the outlook of stellar spectroscopy in the light of The Payne in Section IV and conclude in Section V.

II. The Payne

Current approaches to modeling stellar spectra, either with physical or data-driven models, have important limitations that are well-documented in the recent literature (Boeche et al. 2011; Adibekyan et al. 2012; Bensby et al. 2014; Blanco-Cuaresma et al. 2014; Nissen et al. 2014; Holtzman et al. 2015; Ness et al. 2015; Boeche & Grebel 2016; Casey et al. 2016; García Pérez et al. 2016; Rix et al. 2016; Ting et al. 2016; Ting et al. 2017a; Ting et al. 2017b; Zhao et al. 2016; El-Badry et al. 2018a; El-Badry et al. 2018b). In this section we present our approach to addressing some of these limitations For example, to fully harness the information from spectra, a full spectral fitting method can be advantageous (Ting et al. 2016; Ting et al. 2017a, see detailed discussion in) than equivalent width based methods, as much of the spectral information is embedded in the subtle blended features.. At the core of The Payne is the ability to perform full simultaneous spectral fitting of all stellar labels through an efficient but precise way of “interpolating” a modest set of synthetic model spectra in high-dimensional label space.

The key idea for efficiently interpolating an ensemble of synthetic models is two-fold. First, we do not need to create a high-dimensional “grid” of model spectra, which would be computationally prohibitive for, say, 25 labels in this study; with an adaptive approach described below we only create models within the label space spanned by the data and “where needed”. Second, we resort to building a generative model for the spectra at arbitrary point (in a portion) of label space, as Ness et al. 2015 and Rix et al. 2016. If the model for the spectral flux at each pixel is forced to be a quadratic function of the NN labels, then only a few times N×(N+1)/2N\times(N+1)/2 ab initio spectral models are needed as a basis.

While quadratic models are simple and elegant, they limit the portion of label space over which precise (∼10−3\sim 10^{-3}) flux predictions are possible. For fitting a broad range of stellar labels (e.g., fitting dwarfs and giants or 3000 K≤Teff≤8000 K3000~{\rm K}\leq T_{\rm eff}\leq 8000~{\rm K} simultaneously), quadratic flux models appear too restrictive. Furthermore, for stellar labels such as vmacrov_{\rm macro}, at any given pixel, the variation of flux can be more complicated and often not monotonic. Such complex label-dependences of the flux are illustrated by three one dimensional examples in Fig. 1. In this figure we show the dependence of continuum-normalized flux as a function of TeffT_{\rm eff}, vmacrov_{\rm macro}, and C12/C13C_{12}/C_{13}. Here we assume the same Kurucz synthetic models as we will describe in Section II.4 convolved with the APOGEE averaged line spread function (LSF) to simulate the variations we expect from APOGEE. Clearly a quadratic model cannot capture the behavior of the flux over the entire parameter range, while a more flexible neural network can reproduce the variation in the model very well, as we quantify in greater detail below.

II.2. Neural networks for precise model spectrum prediction

The number of “neurons” ii in Eq. 1 is a free hyper-parameter to be optimized. Increasing the number of neurons enables the approximation a more complicated function, but at the risk of over-fitting the function. Besides adopting more number of neurons, one can also increase the complexity of the neural networks by increasing the number of “layers” by considering the composite of the current composite functions – i.e., fλ∼σ(⋯σ(⋯σ(⋯ ))f_{\lambda}\sim\sigma(\cdots\sigma(\cdots\sigma(\cdots))

Cross-validation experiments described below motivate the following choices. We adopt a two hidden layers model with 10 neurons. This choice was initially motivated by the fact that the number of free coefficients in this simple neural network model is comparable to that in a quadratic model. At least for stellar spectra, designing the neural networks to have roughly the same number of coefficients of simple polynomials seems to be a robust practical guideline. We checked that adopting a significantly more complex neural network model does not improve the qualitative results of this study, but does lead to over-fitting. We train the neural networks by minimizing the L2L^{2} loss, i.e., minimizing the sum of the Euclidean distances between the target (ab initio flux and the model-predicted (or, “interpolated”) flux at each pixel. We found no need for further imposing explicit L1L^{1} regularization (Casey et al. 2016, e.g.,) to the networks as it does not improve the results presented in this study. We limit ourselves to small networks precisely to avoid overfitting, as such regularization is not necessary.

Neural networks are of course not the only flexible model “interpolators”, as Gaussian processes or support vector regressions are also employed in related circumstances. For the case at hand, The Payne has the advantage of being much faster computationally. While the training of neural networks is more computationally expensive than the quadratic models (each wavelength pixel takes about 5 CPU minutes), once the neural networks are trained, the speed of inference is about same as the quadratic models, and is independent of the training set size, as we simply need to evaluate the composite functions. While Gaussian processes are powerful for full Bayesian inferences, predicting a model spectrum at a new label point through Gaussian processes can be extremely slow: it requires the inversion of a matrix, has a complexity of O(Ntrain3)\mathcal{O}(N_{\rm train}^{3}), and can be very memory intensive.

Finally, the fundamental idea of The Payne is different from some of other previous applications of neural networks in spectral analyses (Fabbro et al. 2018; Leung & Bovy 2019). These studies attempted to map spectrum to the stellar label through neural networks, but in this study, we are mapping stellar label to the spectrum. Summarizing the detailed pros and cons of these methods are beyond the scope of this study; here we will only briefly discuss the logic behind our choice. Direct mapping from spectrum to stellar label can be advantageous as the spectral-fitting component becomes trivial – evaluating stellar labels in this case only requires evaluating the mapping/function directly, which is extremely fast. On the other hand, mapping f:f: spectrum →\rightarrow label limits the ability to differentiate the function with respect to the label, unlike The Payne , which has f:f: label →\rightarrow spectrum. Differentiating the emulating function with respect to label can be useful in many cases – especially at low-resolution, comparing ∂f/∂\partial f/\partiallabel to theoretical line lists can be the key to enforcing that elemental abundance are derived from their corresponding absorption features instead of astrophysical correlations. It also allows us to impose theoretical prior as was done in Ting et al. 2017b (Leung & Bovy 2019, but see). This reason prompted our choice to map from stellar label to spectrum (Dafonte et al. 2016, see also). The downside of this approach, however, is that evaluating the label requires least squares minimizations, which is slower than simply evaluating a function. In short, both types of mapping have their own merits, and which method to use clearly depends on the applications.

II.3. The choice of stellar training labels for spectral model building

Beyond the choice of how to interpolate among a set of model grid points, another essential choice must be made: the training set size and the stellar labels at which the ab initio models are to be evaluated to provide training spectra. Formally, we require barely more training spectra than the number of free parameters in the neural networks, which would be 273 training spectra in the case at hand. However, uniformly distributing few hundreds training labels in a high dimensional (Ndim=25N_{\rm dim}=25) space would not be optimal because the distribution will be too sparse in the label space, and the interpolation will not be precise. But in generative models like The Payne we need not draw from a regular, uniformly spaced training labels.

As discussed in Ting et al. 2015, generating training spectra around the label space where real observed stars are expected to occupy can exponentially reduce the number of models needed. The volume of a hyper-ellipsoid in a high dimensional space is exponentially smaller than the volume of a hypercube where the training labels are uniformly distributed. In our illustrative application of The Payne, we fit 25 stellar labels, including all elemental abundance with entries in our line list within the APOGEE spectral range. As stellar parameters, we fit TeffT_{\rm eff}, log⁡g\log g, vmicrov_{\rm micro}, vmacrov_{\rm macro}, and C12/C13C_{\rm 12}/C_{\rm 13} along with the 20 elemental abundances (C, N, O, Na, Mg, Al, Si, P, S, K, Ca, Ti, V, Cr, Mn, Fe, Co, Ni, Cu, Ge). We consider a training set of 2000 training spectra. Rix et al. 2016 showed that adopting more training set than the free parameters will better constrain the flux variation, especially when the range of the parameter space is large. We found that adopting a 10 times larger training set does not change our results qualitatively in this study. For the 2000 training spectra, we adopt an adaptive refinement technique to decide on the training labels as described below.

We start with a “sparse” set of labels that samples TeffT_{\rm eff} and log⁡g\log g from the MIST isochrones (Choi et al. 2016) assuming [Z/H]=−1.5[Z/{\rm H}]=-1.5 to 0.5, Teff=3T_{\rm eff}=3,000−8000-8,000 000\,K and stellar age from 3−103-10 Gyrs, covering both dwarfs and giants. We consider stellar evolution states from the main sequence to the core helium burning at the red clump. We then use these labels to create two convex hulls for the giants (defined with log⁡g<4\log g<4) and the dwarfs (log⁡g>4\log g>4) separately in the Teff−log⁡gT_{\rm eff}-\log g space, i.e., minimum polygons that encompass the tracks from the MIST isochrones. Subsequently, we randomly sample TeffT_{\rm eff} and log⁡g\log g from a uniform distribution within these convex hulls. Analogously, we draw vmicrov_{\rm micro} uniformly from 0.1−3 0.1-3\,km/s and vmacrov_{\rm macro} uniformly from 0−30 0-30\,km/s with 2000 data points. We have found that this choice spans most of the derived APOGEE label space without requiring extrapolation. For C12/C13C_{12}/C_{13} we assume a weak prior. We adopt the isochrones value of C12/C13C_{12}/C_{13} given the stellar parameters of the training data. But we arbitrary spread out the C12/C13C_{12}/C_{13} values on the training set with a uniform distribution of ±35\pm 35. Finally, for this sparse grid, we randomly draw all elemental abundances [X/H][X/{\rm H}] from a uniform distribution with the condition −0.5<[X/Fe]<0.5-0.5<[X/{\rm Fe}]<0.5. Note that, here we train a single spectral model that encompasses both dwarfs and giants.

While the sparse grid is essential to make sure that we capture all cases, spanning a 25-dimensional space with only 2000 training data cannot constrain the flux variation to the needed precision. Therefore, we need to refine the label space from which we draw our training labels. To do that we train The Payne with the sparse grid, fit all APOGEE spectra, which results in an initial distribution of the sample in label-space. Then, we re-sample 2000 training data points with [X/H][X/{\rm H}] drawn from these initial APOGEE label values. We note that APOGEE data does not span the Teff−log⁡g−[Fe/H]T_{\rm eff}-\log g-{\rm[Fe/H]} space uniformly. Therefore, to avoid only fitting the variation of flux well at the at the bulk of the data, we do not resample main stellar parameters with the fitted values, but rather we sample TeffT_{\rm eff}, log⁡g\log g, [Fe/H], vmicrov_{\rm micro}, vmacrov_{\rm macro} and C12/C13C_{12}/C_{13} as before. But we adopt [X/H][X/{\rm H}] from the fitted APOGEE values that have consistent [Fe/H]{\rm[Fe/H]}. In other words, we bin the measured (using The Payne trained on the sparse grid) [X/H][X/{\rm H]} APOGEE values according to their fitted [Fe/H] values with a bin size of 0.2 dex. We only sample [X/H][X/{\rm H}] in the corresponding [Fe/H] bin consistent with the newly drawn [Fe/H] training label. And these 2000 resampled training points constitute the final training set. Our sampling scheme is summarized in Table 1.

We compute 1D LTE spectral models adopting the state-of-the-arts codes atlas12 and synthe maintained by R. Kurucz (Kurucz 1970; Kurucz & Avrett 1981; Kurucz 1993; Kurucz 2005; Kurucz 2013; Kurucz 2017, and reference therein). It is crucial to re-calculate the stellar atmospheric structure as we vary the stellar labels to obtain accurate stellar labels from APOGEE, instead of simply running the radiative transfer code. We calculate the stellar atmospheric structure by partitioning the stellar atmosphere into 80 zones of Rosseland optical depth, τR\tau_{R}, with the maximum Rosseland depth τR=1000\tau_{R}=1000. When generating synthetic models, we automate the inspection of numerical convergence for each layer of the stellar atmospheres. We adopt Solar abundances from Asplund et al. 2009 and the Arcturus stellar labels from Ramírez & Allende Prieto 2011 throughout this study. We assume a standard mixing length theory with no overshooting for convection. After the stellar atmosphere converges, we produce the synthetic model spectra through the radiative transfer code synthe at the nominal spectral resolution of R=300R=300,000000. The synthetic spectra are subsequently convolved to the APOGEE resolution assuming the APOGEE averaged LSF template. We normalize both the synthetic spectra and the APOGEE observed spectra following Ness et al. 2015. In this method, a set of wavelength pixels with the least response to stellar labels, based on the data-driven model The Cannon, are selected. A fourth order polynomial is fitted through the fluxes of these wavelength pixels and is used to approximate the continuum.

A crucial improvement of our ab initio models is the use of an updated line list (Cargile et al., in prep.), which will soon be made publicly available. Improving on the original Kurucz line list, the new line list tweaks three line parameters for every line stronger than 1% at R=300R=300,000000 in either the Sun or Arcturus: the central wavelength, oscillator strength and the dominant broadening parameter. These line parameters are simultaneously fit to the high resolution spectral atlas of the Sun and Arcturus in several angstrom segments in order to capture possible covariance between overlapping lines. We refer readers to the paper for more details. Fig. 2 shows a comparative assessment of the new line list. We synthesize spectra at the Solar and Arcturus stellar labels, convolve and normalize them to the APOGEE resolution with the methods described above. We then compare the models to the observed Arcturus and Solar spectra from APOGEE. There is a total of 7214 pixels in an APOGEE spectrum, and Fig. 2 shows the cumulative number of wavelength pixels as a function of the absolute deviation of the models from the observations at each pixel.

The model-data match based on the updated line list adopted in this study is shown by the blue line, while the match with models that use the standard un-tuned line list (available on R. Kurucz’s website) are shown by the green line. The shaded regions identify the pixels we mask and eliminate in the subsequent modeling – pixels that have normalized model fluxes deviating by ≥\geq2% at the APOGEE resolution from the observed spectra, either for Arcturus or the Sun. About 90% of pixels that we mask are due to disagreement with Arcturus especially in the middle chip of APOGEE, i.e., 1515,800−16800-16,400 400\,Å. The poorer agreement with Arcturus is not surprising because the line list is better calibrated to the Sun than to Arcturus, and because the cooler temperature of Arcturus results in more and stronger lines than in the Sun. The 2% cut is chosen to produce a satisfactory balance between the accuracy and the precision of our derived stellar labels. Imposing a more stringent cut will minimize the systematic errors of the spectral models, but at the expense of the precision we can achieve because we are excluding more spectral information. Also note that this binary spectroscopic mask only discards 12%12\% of the APOGEE spectra, and we are still performing full spectral fitting with all stellar labels simultaneously. This should be distinguished from the ASPCAP mask which APOGEE DR14 imposed, where individual abundances are determined with different filters.

Fig. 2 shows the comparisons of the APOGEE spectra of the Sun and Arcturus, compared to the convolved version of the very high S/N resolution R=300,000R=300,000 FTS spectra of the Sun and Arcturus, serving as “perfect model” templates. The convolved high-resolution observed Solar and Arcturus spectra do not match their APOGEE counterparts perfectly for several reasons. The APOGEE H-band suffers from severe telluric contamination which is imperfectly subtracted. Furthermore, the LSF and continuum normalization that we adopt are not perfect and could contribute to this discrepancy. Nonetheless, the convolved FTS spectra set the baseline for the best case scenario and show that the updated line list is closer to this limit than the original Kurucz line list. We also tested that making a spectroscopic mask at the FTS resolution and subsequently convolving it to the APOGEE resolution does not work. For The Payne, it is crucial to make the spectroscopic mask directly in the observable space. The mask is meant to capture both for theoretical imperfections (imperfect line parameters, non-LTE effects, etc.) and for observational problems (LSF, telluric absorption, etc.).

In Fig. 3 we further investigate which pixels are masked from the fit. The yy-axis quantifies how informative each pixel is, quantified by the rms of the model variations when sampling the training labels; the xx-axis shows the absolute deviation of the model from the observed spectrum for both Arcturus and the Sun. The rms is calculated with the refined synthetic model grid used in the final training. The shaded regions show pixels that are excluded from analysis. Fig. 3 shows that, overall, there is a weak correlation between the deviation and the spectral feature strength. This trend is expected because stronger lines are generally harder to model. But as shown, most of the spectral features are included in our fit, and only a minimal number of spectral features are masked.

Finally, we note that our method is completely general and can be applied to other spectroscopic models. We also tried to apply The Payne to the un-tuned Kurucz line list. We showed that, similar to the results using the new line list as we will present in this paper, the fits even with the old line list exhibit better agreements with the isochrones as well as a flat Teff-abundance trend for open clusters. However, the overall accuracy and precision with the old line list are not as good as the improved new line list. The worse precision is expected because, with the old line list, we need to mask out significantly more pixels (Fig. 2). The slightly worse accuracy (i.e., not as good an agreement with the isochrones) is a bit puzzling. It suggests that the H-band spectroscopic models are not consistent throughout all the APOGEE pixels. Checking how the results vary by restricting to different sub-ranges of wavelength could shed light on this issue, but this is beyond the scope of this paper. Moreover, a thorough comparison would also require us to apply the APOGEE’s ASPCAP pipeline to the new line list (instead of only applying The Payne to the old line list), something that we do not have the tool to perform ourselves. We will defer such detailed explorations to future studies.

II.5. Astrophysical verification of The Payne

In this section we present two important tests of The Payne: first, we compare newly generated ab initio models that were not included in the training step to models predicted from The Payne. This step directly tests interpolation errors in the training of the neural networks. Second, we fit noiseless models with The Payne to see how well we can recover stellar labels in the case of perfect synthetic models. This step tests how much any interpolation errors in flux space translates into uncertainties in determining accurate stellar labels.

Fig. 4 shows how well The Payne interpolates synthetic spectra. We trained on 2000 training spectra and test on the additional 850 validation synthetic spectra that are not used in training. The top left panel shows a small segment of wavelength range, comparing The Payne interpolation with the ab initio calculated spectra. The upper case illustrates a spectrum where the interpolation error is small (¡0.1%). Most of the validation spectra are in this category. The lower case shows one of the few extreme cases where the interpolation is poor (¿1%).

The top right panel shows the absolute interpolation errors for different synthetic spectra at different temperature ranges, taking the median over all wavelength pixels. For each synthetic spectrum, the median interpolation error is only about 0.1% with The Payne, more accurate than the typical S/N observed by APOGEE. Cooler stars have slightly larger errors because there are more spectral features in cool stars and the imperfectness of continuum normalization becomes more severe. We note however that in some cases, the errors can be >1%>1\%. We tested that including 10 times more training data and increasing (or decreasing) the number of neurons does not improve these cases. We will leave the fine-tuning of the network architecture and loss function as well as the tailoring of specific regularization to mitigates these extreme cases to future studies. Nonetheless, although not shown, we also tested that with a quadratic model, the interpolation errors are typically an order of magnitude larger, which is not surprising given the large range in TeffT_{\rm eff} and log⁡g\log g under consideration.

The bottom panels illustrate the pixel-by-pixel interpolation errors, averaging over validation spectra. Plotted on the bottom left panel is the median errors for a randomly selected wavelength segment. Typical pixel-by-pixel errors for The Payne are about 0.1%. The results over all wavelength pixels are summarized in the bottom right panel, which illustrates the cumulative number of wavelength pixels as a function of interpolation errors. The solid lines show the median as before, and the dashed lines indicate the 95 percentile (2σ\sigma) – i.e., pixel-by-pixel, more than half of validation spectra have interpolation errors smaller than the solid line with The Payne, and more than 95 percentile of the validation spectra are within the interpolation errors illustrated by the dashed line.

Having established that The Payne can interpolate models well, we will now investigate how much the interpolation error in flux space translates into accuracy error in determining stellar labels, i.e., in the limit of perfect spectral models with no noise, how well The Payne can recover the stellar labels. This will set a lower limit floor on how accurate (not precision) The Payne can recover stellar labels. Fig. 5 shows the recovery of stellar labels of the validation spectra by fitting (noiseless) validation spectra with The Payne. Throughout this study, we fit spectra by minimizing the χ2\chi^{2} of the interpolated model to the fitting spectra. The χ2\chi^{2} minimization is performed using scipy.optimize.curvefit. When fitting real observed APOGEE spectra, we also take into account the reported uncertainties for individual pixels; pixels masked out by spectroscopic mask are set to have infinite uncertainties. We have tested that initializing at different initial points for the χ2\chi^{2} minimization results in the same solutions. This is not surprising because, at the APOGEE’s resolution, most spectral features are resolved, and hence the degeneracy of stellar labels is not severe (Ting et al. 2017a). As such, we only run the optimization once for each spectrum. Since generating a spectrum to compare with the fitting spectrum requires only evaluating a function (the trained neural networks) which takes only milliseconds, the optimization typically only consumes one CPU second to fit for an APOGEE spectrum.

The top panel shows the 1σ1\sigma of the label recovery. As shown in the red line, for the bulk of the APOGEE spectra which have Teff≃4500−5000 T_{\rm eff}\simeq 4500-5000\,K, in the limit of perfect models, The Payne can recover labels to an accuracy of ≃0.02−0.1 \simeq 0.02-0.1\,dex for elemental abundances, 30 30\,K for TeffT_{\rm eff} and 0.05 0.05\,dex for log⁡g\log g. Some elemental abundances have larger accuracy problems, but these are elemental abundances that have rather weak signatures and/or with strong blends. In practice, almost all the 15 elemental abundances (C, N, O, Mg, Al, Si, S, K, Ca, Ti, Cr, Mn, Fe, Ni) that we will focus in the APOGEE example study have accuracy better than ∼0.05 \sim 0.05\,dex. The blue line shows the accuracy for stars cooler than 4500 4500\,K (e.g., Arcturus). Despite having more spectral features, the typical accuracy for cooler stars is two times worse due the larger interpolation errors, as already illustrated in Fig. 4. We also note that while the there might be biases for individual stars of 0.03−0.1 0.03-0.1\,dex, the bottom panel shows that, if the training sample is a fair representation of the global APOGEE chemical distribution, there is no strong overall biases due to the interpolation error. Plotted is the median deviation of the validation spectra fit to the assumed input. For all abundances, the overall biases is typically less than 0.01 0.01\,dex.

Importantly, we emphasize that the results show the accuracy of The Payne instead of precision because at a given stellar label, although The Payne could incur a bias, the differential recovery can still be very precise. As we will see in the APOGEE example application below, we achieve an elemental abundance precision of about 0.03 0.03\,dex for all elemental abundances by fitting the APOGEE spectra.

III. An illustration of The Payne: 25 Stellar labels from APOGEE data

As a specific application and illustration of The Payne, we fit the entire APOGEE DR14 data set, consisting of ∼270\sim 270,000000 spectra. We only consider the combined APOGEE spectra (instead of individual visits) throughout this study. We train The Payne with only 20002000 ab initio model spectra, and then fit for 25 stellar labels. We also fit for the radial velocity at the same time during the fit to avoid any radial velocity residual from the APOGEE reduction pipeline. When comparing to APOGEE DR14 values, we will refer to the official APOGEE pipeline, ASPCAP, values, instead of the values from The Cannon.

We start out by illustrating how well The Payne does in fitting Arcturus and the Sun at the resolution of APOGEE (Fig. 6). We generated 100 realizations of Arcturus’ and the Sun’s APOGEE spectra, just differing by Poisson noise of the spectra (S/N∼400\sim 400). The violin plots in Fig. 6 show the deviations of our fit of all 100 realizations from the Arcturus and Solar benchmark values adopted from Ramírez & Allende Prieto 2011 and Asplund et al. 2009. The solid black line shows the corresponding APOGEE DR14 values. Overall, The Payne shows comparable deviations from the benchmark values as APOGEE DR14, about 0.1 0.1\,dex for elemental abundances. Part of the deviations is due to the interpolation accuracy error described, but they are also partially contributed by the imperfect spectral models. For individual objects, performing full spectral fitting with The Payne can be more susceptible to model imperfection due to the covariant spectral features, especially with the lenient cut that we made which keeps almost the full APOGEE spectrum. If we were to make a more stringent cut for the spectroscopic mask, i.e. 0.5% error instead of the fiducial 2% error adopted, as shown in the red dashed lines, the accuracy can get better, with the exception of V which only has a very weak feature at the Solar temperature. But this comes at the expense of the precision of stellar labels for the overall sample because, as illustrated in Fig. 2 and Fig. 3, with a more stringent cut, we discard a significant portion of the spectra. Therefore, we adopt the fiducial spectroscopic mask throughout this study.

III.2. TeffT_{\rm eff} & log⁡g\log g

Fig. 7 shows how well The Payne can recover stellar parameters (TeffT_{\rm eff}, log⁡g\log g, [Fe/H]) for both giants and dwarfs with a single self-consistent training model. The left panel shows the values obtained by The Payne , and the right panel shows the APOGEE DR14 calibrated counterparts. APOGEE DR14 does not provide calibrated stellar parameters for dwarfs and sub-giants as they found that the current pipeline struggles to provide reliable estimates for non-giants (Holtzman et al. 2015, e.g.,). Overplotted in the both panels are the MIST isochrones, but at different stellar ages. The Payne derives TeffT_{\rm eff} and log⁡g\log g that are consistent with the MIST isochrones at 7 Gyrs, and the estimates show less scatter at the metal poor end for the giants compared to APOGEE DR14. The APOGEE calibrated their values with the photometric TeffT_{\rm eff} and the asteroseismic log⁡g\log g as we will discuss below, and the calibrated values are more consistent with 1.5 Gyrs old MIST isochrones, which might be too young for the bulk of the APOGEE data. It thus suggests that there is a discrepancy between the photometric TeffT_{\rm eff}, which the APOGEE values calibrated against, with the spectroscopic TeffT_{\rm eff} from The Payne, and the MIST isochrones at the 100 100\,K level. The figure shows that APOGEE DR14 calibrated values also generally favor more metal-rich estimates than The Payne. But this is largely due to their calibration with photometric temperature as we will discuss below.

The Payne does not perform as well for the cooler dwarf stars (Teff<4000 T_{\rm eff}<4000\,K) especially for metal poor stars ([Fe/H] <−0.5\,<-0.5). This could due to multiple reasons. For example, our adopted line list is only calibrated against hotter and more metal rich stars – Arcturus (Teff≃4300T_{\rm eff}\simeq 4300) and the Sun (Teff≃5800T_{\rm eff}\simeq 5800). Moving forward, spectral models built from an atomic line list that has been calibrated against a wider array of stars will be very valuable. The failure in the metal-poor dwarf regime could also be due to a breakdown of the assumptions of LTE.

As shown in Fig 7, the Teff−log⁡gT_{\rm eff}-\log g for dwarfs also exhibits a larger spread in the Teff−log⁡gT_{\rm eff}-\log g Kiel diagram than what is predicted by the stellar evolution models. Part of this larger spread could be due to the fact a non-negligible fraction of the main sequence stars could be unresolved binaries. Fitting single star models to binaries will incur biases which manifest itself as a thicker sequence in the Kiel diagram (El-Badry et al. 2018a). It is beyond the scope of this paper to fit for binaries, but we caution that the single star assumption can compromise the abundance precision that we obtain for dwarfs. For giants, the single star assumption is less a problem because the giant will outshine its dwarf companion, and giant-giant binaries are rare. We refer readers to El-Badry et al. 2018b where The Payne was adopted to fit for main sequence binaries by fitting a mixture of (data-driven) stellar models.

In Fig. 8 we compare The Payne estimates with TeffT_{\rm eff} and log⁡g\log g derived from other external means. In the left panel we compare the spectroscopic TeffT_{\rm eff} to the J−KJ-K color-TeffT_{\rm eff} relation from González Hernández & Bonifacio 2009. For this comparison we only consider giants that have small line of sight extinction, i.e., E(B−V)<0.02E(B-V)<0.02 from the SFD map (Schlegel et al. 1998), avoid the galactic disk ∣b∣>30∘|b|>30^{\circ} and have color 0.1<J−K<0.90.1<J-K<0.9 following González Hernández & Bonifacio 2009. In the right panel we compare spectroscopic log⁡g\log g for a subset of 3000 stars that have APOKASC v3.6.5 asteroseismic log⁡g\log g values. Without calibration, the log⁡g\log g estimates from The Payne agree with the asteroseismic log⁡g\log g value to about 0.07 0.07\,dex with only a weak metallicity dependence. Overplotted in red line is the best fit linear regression. We do not overplot the APOGEE DR14 values because, by definition, APOGEE DR14 log⁡g\log g are calibrated to match the APOKASC asteroseismic log⁡g\log g and the photometric temperature. As shown in the left panel, spectroscopic TeffT_{\rm eff} from The Payne however, is typically 100 100\,K cooler than the photometric TeffT_{\rm eff}, and shows a dependence with metallicity. It is hard to speculate what causes these trends, but it could either be inflicted by the inherent differences between H-band spectroscopic and photometric temperature, since APOGEE DR14 uncalibrated values also show similar offsets, or it could simply be due to the imperfect spectral model and line list. We found that imposing a more stringent spectroscopic mask does not resolve this issue, indicating that the cooler temperature is favored by our spectroscopic model and is not due to interpolation error. But as we will see, even without calibrating this relation, the derived stellar labels from The Payne seem to agree well in other plausibility tests that we will present below. So we choose not to calibrate the temperature and will leave the more detailed study of this discrepancy to future studies (Choi et al. 2018, e.g.).

One particularly interesting aspect of The Payne as shown in Fig. 7, is that besides deriving stellar parameters for the dwarfs, The Payne also yields reasonable TeffT_{\rm eff} and log⁡g\log g for the giants on the cooler end, around 3500 K to 4000 K. In fact, we found that fitting C12/C13C_{12}/C_{13} is crucial to get Teff−log⁡gT_{\rm eff}-\log g that are consistent with the isochrone at the cooler end for the giants. Since C12/C13C_{12}/C_{13} spectral features are highly blended with other features, C12/C13C_{12}/C_{13} can only be reliably derived with a full spectral fitting with all stellar labels simultaneously, an area where The Payne excels.

III.3. C12/C13C_{12}/C_{13} & C/N

The flux variation dependence on C12/C13C_{12}/C_{13} is a particularly difficult to model. As already shown in Fig. 1, the flux variation as a function C12/C13C_{12}/C_{13} has a sharp transition. Above C12/C13≃50C_{12}/C_{13}\simeq 50, the spectral dependence is very weak, and below ∼50\sim 50 the flux varies strongly with C12/C13C_{12}/C_{13}. Since carbon molecular features contribute significantly to the H-band APOGEE spectra, C12/C13C_{12}/C_{13} alters the spectra in a significant way. On the one hand, it implies that fitting C12/C13C_{12}/C_{13} is not only astrophysical interesting, it can also be crucial as part of the spectral fitting, without which the stellar parameters may be biased. But on the other hand, we found that, in the limit of imperfect models, if we do not impose a prior C12/C13C_{12}/C_{13}, the C12/C13C_{12}/C_{13} features can be wrongly adopted to adjust the global fit to get a lower χ2\chi^{2}. Therefore, as discussed in Section II.3, we assume a weak prior for C12/C13C_{12}/C_{13} from stellar evolution models.

Fig. 9 shows the C12/C13C_{12}/C_{13} values estimated with The Payne for all APOGEE stars. On the left, we show the C12/C13C_{12}/C_{13} values for dwarfs (with log⁡g>4\log g>4), and on the right for giants (log⁡g<4\log g<4). Overplotted in black lines are the MIST isochrones for the respective evolutionary states, assuming a stellar age of 7 Gyrs old, and metallicity [ZZ/H] ranging from -0.5 to 0.5. The C12/C13C_{12}/C_{13} values for dwarfs are less well constrained and have a larger scatter from the MIST prediction because the spectral response with respect to C12/C13C_{12}/C_{13} at C12/C13>50C_{12}/C_{13}>50 is weak and yields almost identical spectrum (see Fig. 1). As for the giants, the C12/C13C_{12}/C_{13} values roughly agree with the MIST isochrones, with a sharp transition around 5000 K due the the first convective dredge-up and follow by a second transition as the stars ascend in the HR diagram in the red-giant branch and reach a lower temperature. But the transition temperature seem to be smaller than the predictions from stellar evolution models.

We caution readers not to over interpret the C12/C13C_{12}/C_{13} results as we have assumed a prior for the C12/C13C_{12}/C_{13} in the training set. One of the current challenges of full spectral fitting is that, in the limit of imperfect models, one stellar label, such as C12/C13C_{12}/C_{13}, may in effect ”do the work” of another stellar label. As discussed, the reason to include C12/C13C_{12}/C_{13} is merely to ensure that the stellar parameters are robust at the cooler giant end since it contributes significantly at the cooler end due to the strong features as well as the second dredge-up. It also shows that C12/C13C_{12}/C_{13}, in principle, can be fitted simultaneously with all other labels employing The Payne.

Besides C12/C13C_{12}/C_{13}, the [C/N] ratio of stars will also be modified due to convective dredge up during the giant phase. In fact, the [C/N] ratio has been shown to excellent stellar mass indicators for giants (Martig et al. 2016; Ness et al. 2016; Ho et al. 2017); how much the dredge-up affects the [C/N] ratio depends crucially on the stellar mass. Since there is a tight correlation between stellar mass and stellar age (given a fixed metallicity), determining accurate [C/N] ratios from large spectroscopic surveys is particularly important because they are excellent age indicators for stars. In Fig. 10, we overplot the [C/N] ratios of the APOKASC sample, color-coded with their corresponding asteroseismic ages, with the predictions from the MIST isochrones. Since stellar evolution predictions depend on metallicity, we restrict the APOKASC sample with −0.1<[Fe/H]<0.1-0.1<{\rm[Fe/H}]<0.1 and assume Solar abundances for the isochrones. On the left-hand side, we show the results from The Payne, and APOGEE DR14 on the right-hand side. The Payne values agree better with the isochrones and show a reduced scatter and bias, especially for the older stars, indicating that our C to N abundances are likely more accurate. The excellent agreement between the stellar evolution models with spectroscopic indices also demonstrates that by fitting all stellar labels self-consistently and simultaneously, the improved spectral models and stellar evolution modes can be accurate enough to allow for a direct inference of stellar ages from spectroscopic indices, going beyond data-driven models.

III.4. Element abundance patterns

Elemental abundances are often derived from individual spectral lines, one element at a time. A key goal of The Payne is to demonstrate that all elemental abundances can be measured from stellar spectra directly from a simple χ2\chi^{2} fit by fitting all elemental abundances and stellar parameters simultaneously. In this study, we fit for 20 elemental abundances, namely C, N, O, Na, Mg, Al, Si, P, S, K, Ca, Ti, V, Cr, Mn, Fe, Co, Ni, Cu, Ge, all elemental abundances show visible absorption lines from our line list in the H-band. As already shown in Fig, 5, in the limit of perfect models and data, all of these elemental abundances can be extracted with The Payne.

However, we found that 5 elemental abundances (Na, P, V, Co, Ge) cannot be reliably derived with the current implementation of The Payne, an issue also well-diagnosed in APOGEE DR14 (Holtzman et al. 2015, e.g.,). These elements exhibit large scatter in an [X/Fe]−[Fe/H][X/{\rm Fe}]-[{\rm Fe/H}] diagnostic plot or a large scatter in the precision test (Section III.5.3). Elements like Na, P, V have only weak features (<1%<1\% change in flux for Δ[X/H]=0.05\Delta[X/{\rm H}]=0.05) in the H-band, and unfortunately, the features are also often blended with the telluric sky lines, an issue compounded by the current interpolation errors from The Payne. Although we derive estimates from these elemental abundances, we decided that they are not to be trusted. The reason for a large spread in Co, Ge in an [X/Fe]−[Fe/H][X/{\rm Fe}]-[{\rm Fe/H}] diagnostic is unclear because each of these elements does have a single strong feature in H-band, similar to K, and we have no problem getting reasonable K measurements as shown below. We will defer a more detailed study of the problems to a forthcoming paper. We will focus on the remaining 15 elements, 14 of which (except for Cu) have been reliably determined in APOGEE DR14 for comparison and only consider stars with a fitting reduced χR2<50\chi_{R}^{2}<50.

Fig. 11 shows the comparison of The Payne estimates with the calibrated values from APOGEE DR14, showing a generally good agreement to the level of 100 100\,K in TeffT_{\rm eff}, 0.1 0.1\,dex in log⁡g\log g and 0.1 0.1\,dex in [X/H][X/{\rm H}]. The Payne favors slightly metal-poor estimates, as already discussed in light of Fig. 8. The Payne spectroscopic estimates prefers lower temperatures, compared to the APOGEE DR14 values that are calibrated to photometric temperatures. As [Fe/H] and TeffT_{\rm eff} estimates are covariant (Ting et al. 2017a, e.g.,), this leads to more metal-rich estimates for elemental abundances. Another noticeable deviation is around log⁡g≃2.5\log g\simeq 2.5. Also shown in Fig. 7, the log⁡g\log g values for red clump stars from The Payne are overestimated compared to stellar evolution models. This discrepancy is also consistent with APOGEE uncalibrated values. The reason for this discrepancy is unknown; one possibility is the lack of fitting the helium abundance. It is conceivable that helium abundance differences between the RGB stars and the red clump stars could explain the log⁡g\log g discrepancy (Yu et al. 2018, e.g.,).

Fig. 12 shows the [X/Fe]−[Fe/H][X/{\rm Fe}]-[{\rm Fe/H}] derived with The Payne. The background demonstrates the elemental abundances estimated by The Payne of the giant stars (log⁡g<4\log g<4). Overplotted in white symbols are the literature values. We consider Bensby et al. 2014 to be the main reference literature which provides, in this plot, abundances for O, Na, Mg, Al, Si, Ca, Ti, Cr and Ni. This main sample is complemented by C abundances from Nissen et al. 2014, K abundances from Zhao et al. 2016, Mn abundances from Battistini & Bensby 2015, Cu abundances from Mishenina et al. 2011. For [Fe/H], we adopt [Fe/H] from the same catalog to avoid systematics across different surveys. The Payne attains reasonable [X/Fe]−[Fe/H][X/{\rm Fe}]-[{\rm Fe/H]} without any external calibration. The separation of the high-α\alpha versus the low-α\alpha sequence is clearly visible across all α\alpha-elements. Notably, we attain a Ti trend that is consistent with the literature values – resolving one of the persistent problems in APOGEE (Holtzman et al. 2015, e.g.,). There is a 0.1 0.1\,dex discrepancy between the literature values and The Payne estimates for Si, K, and Ni. But we note that the K abundances from Zhao et al. 2016 adopts NLTE models. The Payne also favors a flat [Mn/Fe] trend, which is at odd with the literature [Mn/Fe] trend.

One important improvement coming from The Payne, as already demonstrated in Fig. 7, is the determination of stellar labels for APOGEE dwarf stars. Fig. 12 and Fig. 13 demonstrate that The Payne recovers consistent abundances for both dwarfs and giants with a few key differences. First, the carbon abundances for the dwarfs are higher than the giants, and at the same time, the nitrogen abundances are lower as expected due to convective dredge up. Second, since dwarf stars are dimmer, they are dominated by stars that are closer to the Sun, and hence the dwarfs show a more prominent low-α\alpha sequence and have relatively fewer high-α\alpha stars. The dwarf abundances also seem to agree better with the literature values for Al, Si, K, Mn and Ni. Since most of the literature values are derived from main sequence dwarf stars, this agreement is encouraging and might suggest that the discrepancies between Fig, 12 and Fig. 13 might partially due to the imperfect spectral models, or could also be astrophysical related, such as atomic diffusion in dwarf stars (Dotter et al. 2017). Interestingly, The Payne produces upward trends for both Cr and Mn for dwarfs, and thus the dwarf Mn abundances agree with the literature values but the Cr abundances do not. Disentangling the discrepancies in Cr and Mn requires an careful investigation of the line list which we will postpone to future studies.

Finally, the dwarf abundances as illustrated in Fig. 13 show a marginally larger spread than the giants, suggesting that the precision of the dwarfs stars might be inferior than the giants stars. This might not be surprising, as a large fraction of main sequence stars could be unresolved binaries. Fitting a single star model to binaries can affect the precision (El-Badry et al. 2018a). Finally, in a companion paper, El-Badry et al. 2018b adopted the dwarf abundances in this study to build a data-driven model and successfully fit for the unresolved binaries spectra, indirectly verifying that the dwarf stellar parameters and metallicities in this study are internally consistent and robust.

III.5. Testing The Payne with open and globular cluster data

In this section we explore the stellar labels derived with The Payne for stars in open and globular clusters with APOGEE spectra. These stars serve as strong tests of The Payne owing both to extensive literature data and also to the fact that open clusters are believed to be at least approximately chemically homogeneous. The latter fact allows us to empirically test the measurement precision of The Payne and also to test for any systematic behaviors in the derived labels as a function of e.g., TeffT_{\rm eff}.

In Fig. 14, we compare [Fe/H] from The Payne with the literature values for 11 known clusters (open clusters and globular clusters) with more than 5 identified members in APOGEE and with metallicity [Fe/H] >−1.5\,>-1.5, where our training set truncates. The open cluster members in APOGEE are identified in Mészáros et al. 2013. We adopt the median of all members of individual clusters to be the estimate of the cluster metallicity, and the shaded regions show the 1σ\sigma metallicity range of all cluster members. Plotted are the differences of The Payne and the APOGEE calibrated metallicity estimates compared to the literature values. By definition, the APOGEE metallicities show no global trend because they are calibrated against these literature values. The deviations of estimates from The Payne shows a weak dependence with metallicity. The trend is similar to the APOGEE metallicity deviations before calibration. In fact, this behavior is likely traced back to the TeffT_{\rm eff}-metallicity biases that we see in Fig. 8. As the origin of these discrepancies is unclear, we choose not to calibrate our TeffT_{\rm eff} to the APOGEE scale. While we do not conform to the standards, as we have discussed in Section III.2 and in various accuracy tests throughout the paper, the APOGEE-Payne scale seems to be more consistent with the MIST isochrone models.

Interestingly, going beyond the global trend, the APOGEE estimates and The Payne estimates show similar relative offsets across various clusters. Since APOGEE and this study adopt very different methods (including different line lists), this suggests that the local correlated deviations from the literature values may be due to the difference between optical spectroscopy (literature values) and H-band spectroscopy (APOGEE spectra). Finally, while there is a discrepancy in metallicity, since this is temperature related, as shown in Fig. 12 and 13, it does not affect the study of [X/Fe][X/{\rm Fe}] since the differences in the two abundances roughly cancels out as they are both caused but the differences in TeffT_{\rm eff}.

III.5.2 Testing the abundances

In Fig. 15, we show the [X/H]−Teff[X/{\rm H}]-T_{\rm eff} trend of three largest open clusters in APOGEE. Open clusters are found to be very chemical homogeneous (Bovy 2016; Ness et al. 2017). Therefore, apart from secondary effects like dredge-up and atomic diffusion (Dotter et al. 2017), their chemical abundances should be independent of their evolutionary state, and hence, TeffT_{\rm eff}. This property is usually used to calibrate out any systematic behavior of [X/H][X/{\rm H}] with TeffT_{\rm eff}. As shown in Fig. 15, The Payne estimates have no significant [X/H]−Teff[X/{\rm H}]-T_{\rm eff} trend for both clusters, showing that our abundances display no strong systematic error as a function of TeffT_{\rm eff}. However, we caution 95% of the members from these three clusters are giants. More follow up studies of dwarf stars in these open clusters are therefore needed to test the stellar labels in the dwarf regime.

Furthermore, as discussed Section III.3, the C and N abundances of stars are sensitive to stellar ages. Since open clusters have well established ages, they can also be used to check the accuracy of our C to N abundances. In Fig. 16, we show the [C/N] ratios of the same three open clusters: NGC 6819 (Kalirai et al. 2001; Anthony-Twarog et al. 2014, 2.5 2.5\,Gyr, e.g.,), M67 (Richer et al. 1998; Sarajedini et al. 2009, 4 4\,Gyr, e.g.,) and NGC 6791 (Grundahl et al. 2008, 8 8\,Gyr, e.g.,).The top panels in Fig. 16 show the measurements from The Payne, and the bottom panels show the calibrated abundances from APOGEE DR14. Overplotted are the predictions from the MIST isochrones, taking into account the metallicities of each cluster – [Z/H]=0[Z/{\rm H}]=0 for NGC 6819 and M67; [Z/H]=0.25[Z/{\rm H}]=0.25 for NGC 6791. The thick black dashed line in each panel shows the MIST prediction for individual clusters given their corresponding stellar ages. As shown, The Payne [C/N] ratios agree better with the isochrones, and there is less spread indicating that our C to N abundances are likely more accurate.

III.5.3 Abundance precision

Fig. 17 shows the elemental abundance dispersion of three open clusters discussed in the previous section. Since open clusters are chemically homogeneous, at least to the level of 0.03 0.03\,dex (Bovy 2016; Liu et al. 2016), their elemental abundance dispersion gives an independent estimate of the measurement precision. Fig. 17 demonstrates that The Payne obtains a precision of 0.03 0.03\, dex for almost all elemental abundances, more precise than APOGEE DR14 calibrated values, especially in the metal rich end (NGC 6791). We caution however that the precision achieves for individual stars clearly depend on the stellar parameters of the stars. The open clusters only probe precision at the metal-rich end. Interestingly, we found that fitting C12/C13C_{12}/C_{13} is the key to get more precise abundances at the metal rich end, presumably also due to a higher contribution from C12/C13C_{12}/C_{13}, especially for the members of NGC 6791 that are, on average, cooler than the other two clusters (see Fig. 15). This might be the reason why APOGEE DR14 is performing somewhat worse in precision at the metal-rich end.

In this cluster precision test, we only consider cluster members that have median S/N  =100−300\;=100-300, the typical S/N of the global APOGEE sample. About 80%80\% of the APOGEE sample has S/N  >100\;>100. The black solid line shows the Cramer-Rao bound for a typical APOGEE K-giant with S/Npix=200{\rm S/N}_{\rm pix}=200, i.e., the best precision one could in principle achieve if there is no systematics from spectral models and interpolation (see Ting et al. 2017a, for a detail discussion on the Cramer Rao bounds). When calculating the Cramer-Rao bounds, we assume the APOGEE LSF as well as the same spectroscopic mask that we impose on the real data. We caution that while this should mimic the instrumental effect and bad telluric regions, there might be other minor instrumental/observation effects that are not being accounted in the Cramer-Rao bound. The Payne allows us to get closer to the Cramer-Rao bound, but we are not yet reaching this fundamental limit.

We also tested how our achieved precision varies as a function of S/N by adding noise to the cluster member spectra, and found that the achieved precision is not very sensitive to the S/N. The precision consistently hovers around 0.03-0.05 for S/N  >50\;>50, and only grow as (S/N)-1 at S/N  <50\;<50. Almost all APOGEE spectra have S/N  >50\;>50. This is also consistent with both the theoretical expectation and previous empirical studies (Ness et al. 2016; Casey et al. 2016; Ting et al. 2017a) which have demonstrated that spectra are generally information rich. Even at S/N∼50\sim 50, through full spectral fitting, precise abundances can be readily achieved.

However, why there is a precision ceiling of ∼0.03\sim 0.03 dex at higher S/N is unclear. This result is in line with previous studies (Bovy 2016; Liu et al. 2016; Ness et al. 2017) illustrating that open clusters are indeed chemically homogeneous to the level of at least 0.03 0.03\,dex. The limits derived in Bovy 2016 are plotted in black dashed line as a reference. These previous studies arrive at this conclusion either from employing a statistical argument (Bovy 2016), a data-driven approach (Ness et al. 2017) or a more careful line-by-line differential analysis (Liu et al. 2016), while our result is based on direct full spectral fitting of physical spectral models to the data. It is interesting that we are not attaining the Cramer-Rao bound. Some argue that open clusters have intrinsic chemical spreads (Liu et al. 2016) and are inhomogeneous at this level. This might well be the reason we are not reaching the best limit. But we also note that due to spectral model and interpolation imperfections, it is possible that the spread we are measuring is due to systematic errors. A further improvement of The Payne will hopefully shed more light on the chemical inhomogeneity of open clusters.

III.6. A catalog of stellar labels for APOGEE DR14 stars from The Payne

We present all stellar labels (TeffT_{\rm eff}, log⁡g\log g, vmicrov_{\rm micro}, vmacrov_{\rm macro}, C12/C13C_{12}/C_{13} and 15 elemental abundances) in this study in an electronic form with this paper. The catalog is summarized in Table 2. We remove duplicated stars in the APOGEE DR14 catalog and exclude stars that have determined stellar labels that are close to the TeffT_{\rm eff}, log⁡g\log g or [Fe/H] boundaries of our training set; we only present stars that have 3050 3050\,K <Teff<7950 \,<T_{\rm eff}<7950\,K, 0<log⁡g<50<\log g<5 and −1.45<[Fe/H]<0.45-1.45<[{\rm Fe/H}]<0.45. We also further exclude dwarf stars that have Teff<4000 T_{\rm eff}<4000\,K because as shown in Fig. 7, our current models cannot determine stellar labels reliably for dwarf stars cooler than this temperature. This leaves a total of 222,707 stars in our catalog.

We caution that in this catalog we keep stars that have large χR2\chi^{2}_{R} in the fitting for completeness, but we recommend readers to only use stars that show “good” in the ”quality_flag” column. This flag excludes all stars with χR2>50\chi^{2}_{R}>50, a fiducial cut we adopt in this study. It also excludes fast rotators with vmacro>20 v_{\rm macro}>20\,km/s (mostly hot stars with Teff>6000 T_{\rm eff}>6000\,K). We found that some rapidly rotating stars yield unreliable abundance patterns. But this is expected because here we do not properly account for stellar rotation vsin⁡iv\sin i and our training grid truncates at vmacro=30 v_{\rm macro}=30\,km/s, a broadening that is still too small for typical fast rotators. We will explore the inclusion of rapid stellar rotation in the future.

IV. Discussion

The Payne provides a straightforward way to perform full spectral fitting with a minimal number of spectral models required; in our case, we only generated 2000 synthetic ab initio spectra for 25 stellar labels. The Payne does not require a boutique spectroscopic mask (García Pérez et al. 2016, e.g., APOGEE/ASPCAP,), but only a simple spectroscopic mask, constructed algorithmically from the comparison of the synthetic and observed spectra of two standard stars. This appears to be sufficient to attain stellar labels that are more precise and broadly consistent with stellar evolution models. But it is important to emphasize that the main goal of this paper is to lay out this new fitting methodology, using APOGEE merely as a sample application. There are several limitations in the current APOGEE-Payne catalog.

Despite the improvement going beyond the quadratic models and a small median interpolation errors of 0.1%0.1\%, the interpolation error can be larger than 1%1\% in some extreme cases (see Section II.5), and can still prohibit obtaining absolute abundances to the level better than 0.05−0.1 0.05-0.1\,dex, especially for the cooler stars. Elements with only very weak and blended features may be more susceptible to the interpolation error, and the absolute abundances for individual stars could be biased upto 0.2 0.2\,dex. Another limitation of this catalog is that we do not fit for stellar rotation, vsin⁡iv\sin i, but rather adopt vmacrov_{\rm macro} as an approximation. We found that for some hot stars with Teff≳6500 T_{\rm eff}\gtrsim 6500\,K, their vmacrov_{\rm macro} values reach the boundary (vmacro=30 v_{\rm macro}=30\,km/s) of our training set and exhibit seemingly spurious abundance patterns. We create a flag in the catalog for these fast rotators and defer a proper accounting of vsin⁡iv\sin i to future studies; ultimately, this can just be another (two) labels to fit. Furthermore, as discussed in Section III.2, there is a 100 K inconsistency between our spectroscopic TeffT_{\rm eff} and external photometric TeffT_{\rm eff}, which appears to favor more metal poor estimates at the high metallicity end and more metal rich estimates at the low metallicity end, with a discrepancy up to 0.1 dex. The reason for this discrepancy in unknown, but it seems to agree with the APOGEE uncalibrated TeffT_{\rm eff}. It thus calls for a more careful analysis of the spectral models adopted in this study and H-band spectral models in general. We also truncate our training set at [Fe/H] =−1.5\,=-1.5 and do not analyze metal poor globular clusters or halo stars. Further, we do not fit for unresolved stellar binaries that might affect the abundances for dwarf stars. Such an analysis was done separately in El-Badry et al. 2018b using The Payne. This illustrates that The Payne is not a panacea for stellar spectroscopy — is only a new methodological framework, and we mention a few areas for future improvement below.

Our analysis of the APOGEE also made simplifying assumptions about the experimental set-up. First, we convolve all synthetic spectra with a fixed averaged LSF template from APOGEE, assuming that the averaged LSF is an accurate representation for all APOGEE spectra. This is not the case because the instrumental dispersion can vary from fiber-to-fiber and observation-to-observation. In this application of The Payne, we do not fit for the LSF, but use vmacrov_{\rm macro} as a free parameter instead to compensate part of the LSF variation. Second, we normalize synthetic spectra the same way as we normalize observed spectra; but even with a self-consistent normalization, the normalization scheme can be still problematic at low temperatures. In particular, Ness et al. 2015 derived a set of “continuum pixels” for APOGEE giant spectra with Teff≃4T_{\rm eff}\simeq 4,000−5000-5,000 000\,K. At lower temperatures, this set of pixels that we adopt might no longer be valid reference points, and the systematics between models and observations can still skew the continuum. In the long run, fitting the LSF and continuum along with the stellar labels might mitigate some of the remaining systematics seen in this study.

The success of The Payne relies on further key ingredients. A robust spectroscopic mask, which we derived here from only the Sun and Arcturus. It will be crucial to have a set of standard stars that all large-scale spectroscopic surveys will observe: due to the subtle combined effect from the instrumental profile, telluric skylines, and normalization, we found that a robust mask must be made based on observations from the same instrument. It is, for example, not sufficient to make a spectroscopic mask using the highest resolution FTS spectra and convolve the mask to the observable space. Second, The Payne must rely on ab initio spectral models that span a broad range of the Teff−log⁡g−[Fe/H]T_{\rm eff}-\log g-{\rm[Fe/H]} space. Again, a limitation of The Payne’s current application derives from the fact that the line list underlying its ab initio models is only calibrated to two stars, the Sun and Arcturus. Models for cooler stars (Teff<4,000T_{\rm eff}<4,000 K) and more metal poor stars therefore remain problematic. One future step will be extending the calibration of the line list beyond the Sun and Arcturus as well as constructing a spectroscopic mask beyond using these two stars. It is also essential to explore other more sophisticated options, such as 1D or 3D NLTE models. With The Payne, only a few hundreds or thousands of spectral models are needed to fit for all stellar labels. Therefore, self-consistent NLTE elemental abundances should be possible with The Payne (see Kovalev et al., in prep.).

V. Summary

We present The Payne as a new framework for fitting stellar spectra with physical models which allows for the simultaneous determination of numerous stellar labels. The key ideas and components of The Payne are summarized below:

The Payne builds generative models that predict spectral fluxes at each pixel for any set of stellar labels, based on a modest discrete set of physical, ab initio model spectra. Only O(1000)\mathcal{O}(1000) suitable ab initio spectra are needed to create a model for fitting Ndim=25N_{\rm dim}=25 labels. This opens up new avenues for determining stellar labels drawing on computationally expensive synthetic models, such as 3D NLTE models.

The generative models for each pixel are built with neural-networks-like function, a flexible extension of the quadratic models.

At any combination of stellar parameters and abundances, the underlying spectral models for the training step consistently calculate the atmosphere structure and the radiative transfer, to yield the ab initio spectra.

An auto-refinement technique iteratively determines the points in stellar label space where ab initio spectra are needed to construct a sufficiently high-quality generative model.

For fitting of actual survey spectra, we apply a straightforward and algorithmic way to construct a pixel mask that excises wavelengths where observed spectra systematically and prominently differ from the model predictions.

The code for The Payne is minimal and transparent, and the fitting is very efficient. Fitting 25 labels simultaneously to an APOGEE spectrum, and accounting for radial velocity at the same time, takes no more than a CPU second per spectrum. One significant advantage of this speed is the ability to explore many detailed options and choices in the fitting.

The Payne is a general method that can be applied to any set of synthetic models and survey data. As a sample application, we applied The Payne to the full APOGEE DR14 catalog. By directly fitting all stellar labels simultaneously, The Payne delivers the state-of-the-art results that appear both accurate and precise, without any external calibration. We demonstrate this by comparing the APOGEE labels derived from The Payne to APOGEE DR14 results and subjecting them to a variety of physical plausibility tests, illustrating the following points:

Fitting all data only with a single generative model, we obtain stellar parameters (TeffT_{\rm eff}, log⁡g\log g, [Fe/H]) for both giants and dwarfs that are broadly consistent with external constraints, although modest [Fe/H]-dependent systematic offsets are evident when comparing to photometric temperatures and asteroseismic log⁡g\log g values.

We obtain detailed elemental abundances for APOGEE dwarf stars, which had not been done within APOGEE before. Broadly, the abundance patterns we see in the dwarfs are consistent with those in the giant sample.

The derived elemental abundances for stars within open clusters show no systematic trends with TeffT_{\rm eff}.

We resolve the problem of the “peculiar” [Ti/Fe]−[Fe/H][{\rm Ti/Fe}]-[{\rm Fe/H}] abundance distribution that has been persistent in APOGEE data releases and obtain a physically plausible abundance pattern.

Our derived [C/N] ratios agree well with predictions for surface abundance changes due to dredge-up in stellar evolution models. The agreement between stellar evolution models and our [C/N] ratios vividly demonstrates that it is possible to infer stellar ages directly from ab initio spectroscopic labels.

We achieve an elemental abundance precision of ∼0.03 \sim 0.03\,dex for all elements. On this basis, we confirm that open clusters are chemically homogeneous to a level of at least 0.03 0.03\,dex.

The Payne offers a very powerful framework to perform accurate, efficient comparison of large data sets to sophisticated spectral models. Work is currently underway to improve the training of The Payne among the cool giants, to employ NLTE spectral models, and to apply The Payne to a diverse variety of existing spectroscopic and photometric data sets.

References