They might be giants: luminosity class, planet frequency, and planet-metallicity relation of the coolest Kepler target stars

Andrew W. Mann, Eric Gaidos, Sébastien Lépine, Eric Hilton

Introduction

The NASA Kepler mission (Borucki et al., 2010) has ushered exoplanet science into a new phase of analysis based on the statistics of large samples. Among the more elementary statistics derived from Kepler results are the planet occurrence around stars (Howard et al., 2011, henceforth H11), the distribution of planet size (or mass Wolfgang & Laughlin, 2011; Gaidos et al., 2012), correlations between the presence of planets and the properties of the host stars (e.g. Schlaufman & Laughlin, 2011, henceforth SL11), and the characteristics of multi-planet systems (Fabrycky et al., 2012). These findings yield important constraints on models of planet formation and evolution, and are best established for solar-type stars (late F through early K spectral types) because they constitute the vast majority of Kepler targets.

The results of Kepler were first preceded by the findings of radial velocity surveys of solar-type stars. More than 15% of dwarf stars have close-in (∼\sim0.25 AU) planets with orbital periods less than 50 days (Howard et al., 2010, 2011) and this fraction increases with orbital period (Mayor et al., 2011). The same authors find that planet occurrence is inversely related to planet mass or radius, with “super-Earths” outnumbering Jupiter-size planets by more than an order of magnitude. Around solar-type stars, the presence of giant planets is strongly correlated with super-solar metallicity (Gonzalez, 1997; Santos et al., 2004; Fischer & Valenti, 2005; Johnson et al., 2010), but this correlation does not appear to hold for smaller planets (Sousa et al., 2008; Bouchy et al., 2009; Mayor et al., 2011). As with results from Kepler, these findings are primarily for solar-type stars because many nearby representatives are bright enough for ground-based Doppler radial velocity observations.

Very cool (late K and early M type) dwarf stars have become popular targets of planet searches (e.g. Charbonneau et al., 2009; Vogt et al., 2010; Bean et al., 2010; Mann et al., 2011; Apps et al., 2010; Fischer et al., 2012). Planets around cool stars are easier to detect because of the stars’ smaller masses and radii. Furthermore, because these stars are less luminous, close-in and thus detectable planets can still orbit within the “habitable zone,” where an Earth-like planet would avoid the “snowball” or runaway greenhouse climate states (Gaidos et al., 2007). However, the statistics of planets around these stars are poorly established. These stars are underrepresented in magnitude-limited Doppler surveys as well as the Kepler target list. Only 2% of Kepler target stars are classified as possible M types (cooler than 4000 K), whereas >>70% of all stars within 20 pc are M dwarfs (Henry et al., 1994; Chabrier, 2003; Reid et al., 2004).

Nevertheless, Kepler data has been used to draw two important conclusions about late-type exoplanet hosts. First, H11 found that the frequency of stars with planets on close-in (P<50<50d) orbits rises with decreasing effective temperature through early K-type and that an even higher fraction of M dwarf stars may host such planets. Second, SL11 claimed that late K dwarf stars, but not solar-type stars, hosting super-Earth to Neptune sized candidate transiting planets are more metal rich than stars for which transits have not been detected. These findings offer potential tests of theories of planet formation (Fischer & Valenti, 2005; Kennedy & Kenyon, 2008; Cumming et al., 2008)

Kepler targets are selected from the Kepler input catalog (KIC) based on the ability of the mission to find transiting planets, especially in the habitable zone; ideally, the target catalog should consist exclusively of dwarf stars for which the signal of a transiting planet is largest, and exclude sub-giant and giant stars. Brown et al. (2011) used D51 (Mg Ib line) photometry and Sloan gg-D51 color to exclude giants, however this is also sensitive to temperature and metallicity and is not available for all targets. The KIC includes Sloan (grizgriz) and 2MASS (JHKJHK) magnitudes; stellar parameters are estimated by forward modeling of the photometric data with the synthetic spectra of Castelli & Kurucz (2004), and effective temperature TeffT_{eff}, gravity log gg, and metallicity [M/H] as free parameters. Stellar mass and distance are then estimated using luminosity, TeffT_{eff}, and log gg from the stellar evolutionary models of Girardi et al. (2000). The combination of stellar mass and log gg then yields a stellar radius.

Brown et al. (2011) state that KIC radius estimates have average errors of 35% and are not reliable for stars cooler than 4000 K. H11 point out that, because of the difficulty in constraining log gg, the radii of some stars, particularly sub-giants, may be underestimated by a factor of 2 or more in the KIC. Gaidos et al. (2012) found that consistency between the Kepler candidate planet catalog and the M2K Doppler survey could be achieved if the former was incomplete compared to estimates based on KIC radii. They further point out that Kepler planet candidates were conspicuously sparse among late K stars with colors that are shared by both dwarfs and giant stars. Finally, Muirhead et al. (2012) (henceforth M11) showe that KIC estimates for the radii of many Kepler M dwarfs hosting planets are smaller than KIC values by as much as a factor of two. This discrepancy is not to be confused with the 5-10% radius difference between radii of the most refined models and measurements by interferometry and observations of eclipsing binaries (e.g. López-Santiago et al., 2010; Kraus et al., 2011).

Reliable stellar parameters are a prerequisite for robust statistical analysis of planets, especially transiting planets. These are needed not only for stars for which planet candidates have been detected (referred to as Kepler Objects of Interest or KOIs), but also for the target sample as a whole. The radius of a planet producing a given transit depth is proportional to the radius of its host star. Likewise, the transit signal produced by a planet of a given radius - and hence its detectability around a star in the survey - also depends on stellar radius. If some target stars are actually larger or even giant stars, then planets are less likely to be detected in that sample, which means that the most likely occurrence rate of those planets is higher. For M dwarf stars in general, and particularly for the coolest Kepler target stars, parameters such as radius are uncertain or even very unreliable (e.g. Johnson et al., 2012, M11).

Brown et al. (2011) metallicities are reliable to 0.4 dex for solar-type stars, but are essentially useless for stars with Teff<4000{}_{eff}<4000. Instead, SL11 use Sloan g−rg-r colors for a given J−HJ-H range (a proxy for spectral type) as an indicator of the amount of Fe line blanketing at blue wavelengths, and hence metallicity. They construct mean g−rg-r vs. J−HJ-H loci for KOIs and Kepler stars without identified transits. They find a significant difference between the g−rg-r colors of the two populations for stars with J−H≈0.62J-H\approx 0.62, corresponding to late K-type stars. Based on stellar models, SL11 argue that the late-type KOIs are ≃0.2\simeq 0.2 dex more metal rich than Kepler targets with no detected transit. However, K giants are significantly bluer than dwarfs in g−rg-r, for the same J−HJ-H (Yanny et al., 2009). Thus, contamination of the Kepler target sample by giants would shift the locus of target stars to bluer g−rg-r, but would not affect the KOI locus, as planets are less detectable, or completely undetectable around giant stars. Realizing this, SL11 constructed and analyzed artificial mixed data sets to estimate that a 10-30% contamination by giants would also produce the observed offset.

M11 use the equivalent widths of atomic lines in the KK (2.2 μ\mum) band (Rojas-Ayala et al., 2012) and their measurements of late-type KOIs’ metallicities are consistent with, or slightly metal-poor (median [M/H] = -0.10) compared to the solar neighborhood (M/H ≃−0.05\simeq-0.05 Johnson & Apps, 2009). SL11 and M11 are consistent with each other if the Kepler target list itself is biased toward metal-poor M dwarfs, or if the offset found by SL11 is due to high giant contamination in Kepler late-type target stars.

Moderate resolution spectra are nearly always sufficient to distinguish K and M giants from their dwarf cousins. In addition Ciardi et al. (2011) showed that some giant stars can be identified based on JHKJHK photometry alone. In this paper, we combine moderate resolution spectra of a sample of Kepler targets with KIC photometry to refine the planet occurrence rate for late-type stars calculated by H11, and determine if the giant fraction is high enough to explain the color offset observed by SL11. In Section 2 we present spectroscopy of a representative sample of late-type Kepler target stars. In Section 3 we use both spectroscopy and photometry to derive luminosity classes and calculate the giant fraction for late-type Kepler target stars. In Section 4 we use this information, plus radii based on stellar evolutionary models, to refine the planet occurrence around these stars. In Section 5 we calculate and compare the mean g−rg-r colors (as metallicity proxies) of KOIs and a bona fide dwarf sample, and show how and why our results differ from those of SL11.

Sample, Observations, and Reduction

Because derived KIC parameters may not always be reliable, we instead select our sample using photometry. A sample of stars with V−J>2.5V-J>2.5 will include >98%>98\% of all M dwarfs, as well as most of the K7 dwarfs in the sample (Lépine & Gaidos, 2011, henceforth LG11). Although 2MASS JJ magnitudes are available for almost the entire sample, VV magnitudes are not. Kepler magnitudes (KPK_{P}), however, are available for all target stars. For M0 stars, KP−V≃−0.43K_{P}-V\simeq-0.43keplergo.arc.nasa.gov/CalibrationZeropoint.shtml so we conservatively select stars with Kp−J>2Kp-J>2 observed in Quarters 0-2 by Kepler and retrieved from the Multimission Archive (STScI). We remove stars with a contaminating star within 1 arc second.

Bright Kepler target stars were selected in a fundamentally different way from dim stars (see Figure 1 and Batalha et al., 2010). We separately analyzed dim (KP>14K_{P}>14) and bright (KP<14K_{P}<14) stars. Bessell & Brett (1988) showed that giant stars tend to have more extreme J−HJ-H colors than their dwarf counterparts. However, we wanted to investigate how misidentified giant stars in the KIC are distributed with J−HJ-H color. Thus we further subdivided our sample into four J−HJ-H color bins: J−H≤0.70J-H\leq 0.70, 0.70<J−H≤0.760.70<J-H\leq 0.76, 0.76<J−H≤0.820.76<J-H\leq 0.82, and 0.82<J−H0.82<J-H for the bright stars and J−H≤0.62J-H\leq 0.62, 0.62<J−H≤0.650.62<J-H\leq 0.65, 0.65<J−H≤0.680.65<J-H\leq 0.68, and J−H>0.68J-H>0.68 for the dim stars. Color bins were designed such that each contains a similar number of stars. We observed a sample of stars within each bin, selected randomly with respect to J−HJ-H. We observed more bright stars because they are more observationally accessible, although we observed targets spanning all Kepler magnitudes to detect trends with KPK_{P}. In total we observed 382 stars covering 6.5<6.5< KPK_{P}<16<16, 0.40<J−H<1.000.40<J-H<1.00 and KIC effective temperatures 3200<3200< TeffT_{eff}<5050<5050 K. The distribution of observed targets is shown in J−HJ-H and KIC TeffT_{eff} space in Figure 2. A list of observed targets is given in Table 1.

Observations were obtained between June 16 and Aug 28 (2011) with the SuperNova Integral Field Spectrograph (SNIFS, Lantz et al., 2004) at the University of Hawaii 2.2m telescope on Mauna Kea and the Boller and Chivens CCD Spectrograph (CCDS) or the Mark III spectrograph (MkIII) at the MDM Observatory 1.3m McGraw-Hill telescope on Kitt Peak. SNIFS is an optical integral field spectrograph with R≃1300R\simeq 1300 that splits the signal with a dichroic mirror into blue (3000−52003000-5200 Å) and red (5000−95005000-9500 Å) channels. SNIFS images were resampled with microlens arrays, dispersed with grisms, and focused onto blue- and red-sensitive CCDs. Processing of SNIFS data was performed with the SNIFS pipeline, described in detail by Aldering et al. (2006) and Pereira et al. (2010). SNIFS processing included dark, bias, and flat-field corrections, assembling the data into red and blue 3D data cubes, and cleaning them for cosmic rays and bad pixels. After sky subtraction, the spectra are extracted with a PSF model, and wavelengths were calibrated with arc lamp exposures taken at the same telescope pointing as the science data.

The CCDS and MkIII spectrographs cover 5700−93005700-9300Å and 4400−83004400-8300Å with R≃1150R\simeq 1150 and ≃2300\simeq 2300, respectively. Standard reduction of data taken with the CCDS and MkIII was performed with IRAF, following the practice of overscan subtraction, division by flat field, and extraction of the spectra. Spectra were wavelength-calibrated against NeArXe comparison arcs. All observations (including SNIFS) were flux-calibrated and telluric lines were removed based on observations of the NOAO primary spectrophotometric standards Feige 66, Feige 110, and BD+284211. All spectra had a median S/N of >30>30 in the 6000-7000Å range, and the median S/N of all spectra in this range was 50.

Our spectroscopic set only covers KP−J>2.0K_{P}-J>2.0, but we also consider a separate ‘photometric sample’ that includes stars with 0.56<J−H<0.66⋃KP−J>20.56<J-H<0.66\bigcup K_{P}-J>2. This is done so we can ensure coverage of the sample of late K stars used by SL11 (see Section 5). The KIC includes JHKJHK photometry from 2MASS (Skrutskie et al., 2006) and visible-wavelength photometry through SDSS grizgriz and D51D51 filters. We add photometry from the Wide-field Infrared Survey Explorer (WISE, Wright et al., 2010), which includes 3.4μ\mum, 4.6μ\mum, 12μ\mum, and 22μ\mum bands.

Luminosity Class

We determine luminosity class by comparing the spectral indices or colors of Kepler target stars to those of stars drawn from ‘training sets’ of known giants or dwarfs. We first discuss how we construct our training sets. We then explain our choice of indices and color-color relations, based on previous work on giant/dwarf discrimination and derived empirically from examination of the differences between the dwarf and giant training set. We use the colors and spectroscopic indices of stars in the training sets to construct a likelihood estimator, such that we can calculate the likelihood that a given star is a giant (or dwarf). That calculation is explained in Section 3.4.

We construct an uncontaminated set of dwarf stars from a sample of high proper motion-selected late dK and dM stars (LG11). The brightest (J<9J<9) northern stars in the LG11 catalog have visible-wavelength spectra (Lépine et al. in prep), obtained with the same instruments and reduced in the same way as was done for Kepler targets observed for this paper. Although the sample from Lépine et al. (in prep) includes more than 1500 spectra, we construct our dwarf sample only from the 620 targets with spectra from SNIFS/UH2.2m, which includes the Ca II triplet feature at 8484−86628484-8662Å.

LG11 use J−HJ-H, and H−KH-K colors, combined with proper motion from SUPERBLINK (Lépine & Shara, 2005) and (for some targets) parallax information from Hipparcos (van Leeuwen & Fantino, 2005; van Leeuwen, 2007) to remove giant stars. Based on those stars in LG11 with parallaxes, we estimate that fewer than 0.5% of the resulting sample will be giants. However, because of strict cuts in J−HJ-H and H−KH-K, the LG11 sample is incomplete and biased against dwarfs with much redder or bluer colors. LG11 also use a color cut of V−J>2.7V-J>2.7 to select mostly M dwarfs. This excludes some mid- to late-K stars which will be included in our (KP−J>2⋃0.58<J−H<0.66K_{P}-J>2\bigcup 0.58<J-H<0.66) color cut for the photometric sample (see Section 2). We therefore add 60 late K and early M dwarfs included in the Hipparcos catalog that have UH2.2m spectra but lie outside the cuts imposed by LG11. These stars are confirmed to be dwarfs by their Hipparcos parallaxes. We also add 150 M dwarfs with spectra from SDSS, including 50 dwarf from West et al. (2011), with r−Jr-J and J−HJ-H colors consistent with our targets of interest. We verify that these targets are dwarfs using a cut with reduced proper motion, where the reduced proper motion in the SDSS gg band is:

and μ\mu is the proper motion in arcsec yr-1. This quantity is similar to the absolute magnitude, such that giant stars will have much lower reduced proper motions than dwarfs of the same color. We only select SDSS stars with Hg>2.2(g−r)+7.0H_{g}>2.2(g-r)+7.0, and μ>15\mu>15 arcsec yr-1, which we determine empirically from our UH2.2m targets with SDSS photometry.

Our sample of >>300 giant spectra is constructed from multiple catalogs, specifically Fluks et al. (1994), Danks & Dennefeld (1994), Allen & Strom (1995), Serote Roos et al. (1996), Montes et al. (1999), and Lançon & Wood (2000), as well as 80 bright stars we observed with UH2.2/SNIFS that are confirmed to be giants by Hipparcos. Many spectra have significantly higher resolution than our own observations. We convolve these data with a gaussian to match the resolution of our own sample to remove any resolution-dependency in our results. To include sufficient SDSS photometry, we supplement our giant training set by including 200 giant stars with spectra from SDSS all with r<16r<16 and proper motions consistent with zero. We require these SDSS spectra to have spectroscopic indices consistent with the rest of the giant training set. Because we select only SDSS stars with indices consistent with indices from spectra from the rest of the training stars, SDSS giant stars have no effect on our spectroscopic determination of luminosity class. Rather, these SDSS stars are added only for their photometry.

SDSS, 2MASS and WISE colors are available for much of our giant and dwarf training set; however, most lack D51D51 photometry, which covers the gravity-sensitive Mg Ib line at 5200Å. Instead, we synthesize equivalent g−D51g-D51 colors from the spectra of our training set. We obtain the zero point for the synthesized colors of those stars in our sample which have both spectra and gg and D51D51 magnitudes.

2. Spectroscopic Determination of Luminosity Class

Our determination of luminosity class uses six different gravity-sensitive molecular or atomic indices (Table 2 and Figure 3). Molecular and atomic indices are ratios of the average flux levels in a specified wavelength region to that of a pseudo-continuum region. Indices are useful for M dwarfs where the continuum is poorly defined. The values of most indices are a function of both gravity and temperature of the star. To remove this degeneracy we compare measured indices to the TiO5 spectral index. TiO5, as defined by Reid et al. (1995), is sensitive to spectral type and metallicity (Woolf & Wallerstein, 2006; Lépine et al., 2007) but it has minimal gravity dependance (Jao et al., 2008) (see Figure 3).

We show spectra of giant and dwarf stars with similar effective temperatures in Figure 3, with the location of each feature labeled. As can be seen, most atomic lines are weaker in giants than in dwarfs. Indeed the Na I doublet (8172-8197Å) and K I (7669-7705Å) lines are quite shallow in giants while relatively deep in dwarfs (Torres-Dodgen & Weaver, 1993; Schiavon et al., 1997; Reid & Hawley, 2005). Molecular lines provide additional luminosity-dependent spectral signatures. Metal hydride bands, such as the CaH bands defined by Reid et al. (1995) and Lépine et al. (2007) have been used for luminosity classification, although they are less useful for stars earlier than K7. The calcium triplet (8484−86628484-8662Å) is a useful indicator of gravity (e.g. Cenarro et al., 2001b; Kraus & Hillenbrand, 2009), especially for M stars which emit comparatively more at red wavelengths. Giant and dwarf training sets overlaid on Kepler target star indices are shown in Figure 4.

3. Photometric Determination of Luminosity Class

We can use the available photometry to determine the luminosity class of a much larger sample of Kepler stars lacking spectra. Brown et al. (2011) primarily use g−D51g-D51 vs. g−rg-r and J−KJ-K vs. g−ig-i colors to separate Kepler late-type giants from dwarfs. Both giants and the coolest dwarfs in the sample have relatively weak Mg Ib lines, creating overlap between the dwarf and giant training sets at red g−rg-r. A similar effect happens with J−KJ-K. Near-infrared photometry (JHKJHK) has long been used to separate giants and dwarfs at redder colors (Bessell & Brett, 1988), in part due to strong CO and weak Na I and Ca I absorption in giant stars. But for K and early M stars with J−H<0.7J-H<0.7 and H−K<0.2H-K<0.2, the giant and dwarf sequences overlap, creating a sizable region of ambiguity. At mid-infrared wavelengths, most giant stars have warm dust emission, leading to significantly redder colors in the WISE bandpasses. Other relations can be derived from an examination of our giant and dwarf training sets. z−Kz-K vs. g−Jg-J follows a similar distribution to that of J−KJ-K vs. g−ig-i, but the giant and dwarf samples bifurcate at g−J≃3.0g-J\simeq 3.0, which makes this color useful for isolating the reddest giants. Giant and dwarf training sets overlaid on Kepler target star colors are shown in Figure 5.

4. Application of training sets to the Kepler sample

After each spectral index or color is measured or calculated for Kepler targets and both training sets, we identify stars as giants or dwarfs following the same technique as Gilbert et al. (2006). We begin by using the spectral index or color measurements of the training stars to produce a two-dimensional probability distribution function (PDF) for each index (or color). The PDFs are constructed by treating the strength of each index or color (henceforth SS) as a Gaussian distributed variable with respect to XX. For spectroscopic determination of luminosity class, XX is a parameter that primarily relates to the spectral type (although it may have some gravity dependence), while SS is a parameter that primarily relates to log gg. For the spectroscopic determination of luminosity class, XX is the TiO5 band and SS is one of our six gravity-sensitive indices (Na I, Ca II, Ba II/Fe I/Mn I/Ti I, K I, or CaH). For photometric determination of luminosity class, XX is defined as g−Jg-J, g−ig-i, J−HJ-H, g−rg-r, 3.4μ3.4\mum - 22μ22\mum, or J−3.4μJ-3.4\mum and SS is z−Kz-K, J−KJ-K, H−KH-K, g−D51g-D51, 4.6μ4.6\mum - 12μ12\mum, or K−4.6μK-4.6\mum, respectively. Values of SS are binned according to their corresponding XX value. Bins in XX are designed to contain an equal number of stars (20-25) in each bin, and because of this are not equally spaced in XX. The mean (S‾\overline{S}) and standard deviation (σS\sigma_{S}) of the distribution is computed in each bin. The two-dimensional PDF takes the form:

where CC is a normalization such that the entire PDF integrates to 1. PDFs for both giant and dwarf training sets overlaid on Kepler target star indices or colors are shown in Figures 4 and 5 for the spectroscopic and photometric sets, respectively.

The likelihood that star ii is a dwarf for a given index jj is:

where wjw_{j} is a weighting factor. Weights are calculated by determining the efficiency of a given feature at separating giants from dwarfs as a function of XX. We take a random subsample (half the total sample) from each training set, and add Poisson noise to the spectra/colors consistent with our observations or given photometric errors. We then apply Equations 2 - 4 to the subsamples using wj(X)=1w_{j}(X)=1 for all X,jX,j. Values of wjw_{j} are then set based on the fraction of dwarfs/giants correctly identified within a training set. wj(X)=1w_{j}(X)=1 if the feature/color identifies 100% of the targets within a given XX bin correctly and wj(x)=0w_{j}(x)=0 if the feature/color identifies 50% or less (i.e. no better than guessing) of the targets correctly. Weights are linearly interpolated (based on the fraction of stars correctly identified) between these two values.

Repeating the calculation of LiL_{i} using wj=1w_{j}=1 for all jj does not change the classification of any stars with spectra (i.e. our results from spectra are essentially independent of our choice of weighting scheme). However, this is not the case for luminosity classes determined from color-color relations. The reason for this is the significant overlap between the PDFs of the color metrics for giant and dwarf training sets (e.g. 2.3<g−J<2.82.3<g-J<2.8 and 1.6<z−K<1.91.6<z-K<1.9, see Figure 5). In overlapping regions, indices or colors will give similar probabilities for a star being a giant or a dwarf, making the metric less useful in giant/dwarf discrimination. This problem is solved by our weighting scheme, as regions where giant and dwarf training sets overlap tend to have lower weights. We show a plot of the weights for the color-color relations in Figure 6. Weighting factors are set to 0 if any of the relevant indices/colors for a given star are missing or lie outside the range of our training sets.

We identify all Kepler target stars with spectra as a giant or a dwarf with better than 99% (Li>2.0L_{i}>2.0 or Li<−2.0L_{i}<-2.0) confidence. The full list of determined luminosity classes for stars with spectra is given in Table 1. For the photometric sample, ≃97%\simeq 97\% stars are placed into unambiguous giant or dwarf categories (⟨Li⟩>1.5\langle L_{i}\rangle>1.5 for dwarfs or ⟨Li⟩<−1.5\langle L_{i}\rangle<-1.5 for giants). However, ≃3%\simeq 3\% of the sample are more ambiguous, most of which lack photometry in several bands.

Since giant/dwarf assignments based on spectroscopy are very accurate, only binomial errors are considered for the spectroscopic sample. For uncertainty estimates from the photometric sample, we re-apply our likelihood calculations using 1000 different subsets of our training sets, adding random (Poisson) noise to the photometry, and then recalculating the giant fraction in each case. The variation in giant fraction is added in quadrature with binomial errors. This does not consider systematic errors (e.g. systematic photometric errors, discrepancies between training sets and Kepler target stars, etc).

5. Giant Star Fraction

We find that, for the coolest Kepler stars (KP−J>2K_{P}-J>2), giant stars dominate the bright (KP<14K_{P}<14) Kepler target stars but are relatively rare among dim (KP>14K_{P}>14) targets. The fraction of giants is 96±1%96\pm 1\% for bright stars, 7±3%7\pm 3\% for dim stars, and 52±3%52\pm 3\% for the combined set (based on our spectroscopy). Photometric assignments (considering KP−J>2K_{P}-J>2) give consistent giant fractions: 97±2%97\pm 2\% for bright stars, 11±3%11\pm 3\% for dim stars, and 55±3%55\pm 3\% for all stars with KP−J>2K_{P}-J>2. The fractions in each brightness bin decrease somewhat when we apply a KIC log g>4.0g>4.0 cut. The giant fraction becomes 74±8%74\pm 8\% for bright stars and 3±2%3\pm 2\% for dim stars. The fraction of giants for all stars significantly decreases to 10±2%10\pm 2\%, due mainly to the large number of stars lacking any log gg classification, most of which are giants and all of which are removed by this cut.

Planet occurrence

Following the work of H11, we calculate the planet occurrence, ff, which is defined as the total number of planets, within a given range (in orbital period and radius) and considering all orbital inclinations, per star within a given range (in TeffT_{eff}, log gg, and KPK_{P}). Planet occurrence will be somewhat higher than the fraction of stars with planets due to the presence of multi-planet systems, but if the rate of planet multiplicity is low, then these two quantities will be nearly identical.

We first calculate the planet occurrence following the nonparametric method of Gaidos et al. (2012). The total planet occurrence, ff, is the sum of individual planet occurrences (fif_{i}) over all ii planets that fall within a given range in orbital period and radius. The most probable occurrence of the iith Kepler detected planet in the population of jj Kepler target stars is:

where di,j=1d_{i,j}=1 if the S/N of a planet transit around the jjth star is sufficient to detect the transit, and 0 otherwise, pi,jp_{i,j} is the geometric probability of a transit, and jj is summed over all target stars that fall within a given range in TeffT_{eff}, log gg, and KPK_{P}. We consider a planet detected (di,j=1d_{i,j}=1) if:

where δ\delta is the transit depth, NN is the number of transits that occur over the observation interval, τ\tau is the transit duration in minutes, and σCDPP\sigma_{CDPP} is the 30 minute combined differential photometric precision (CDPP) of Kepler. We use Quarter 1-2 30 minute CDPP values from Kepler. Our detection threshold S/N=7S/N=7 matches what is used by Borucki et al. (2011) and Batalha et al. (2012).

For small planets on nearly circular orbits,

where PP is the orbital period in days and M∗M_{*} and R∗R_{*} are the star’s mass and radii in solar units. Values for M∗M_{*} and R∗R_{*} are computed by interpolating a grid of stellar radii/masses from the Dartmouth Stellar Evolution Database (DSEP Dotter et al., 2008) at estimated values of TeffT_{eff}, [Fe/H], and age. We use DSEP because radii and masses derived from their isochrones are in good agreement (<0.03<0.03 RMS deviation in radius) with current observations from interferometry (Dotter et al., 2008; Feiden et al., 2011).

For exoplanet hosts we use the metallicities given in M11, but for field stars metallicities are drawn from a random gaussian distribution of metallicities with [Fe/H]‾=−0.07\overline{[Fe/H]}=-0.07 and σ[Fe/H]=0.20\sigma_{[Fe/H]}=0.20. This distribution is designed to be consistent with the distribution of M dwarfs in the solar neighborhood (Johnson & Apps, 2009; Casagrande et al., 2011). Ages are assigned randomly assuming a constant star formation rate (excluding ages <100<100 Myr). However, since M dwarfs do not change significantly while on the main sequence, our results are not changed when we fix all ages to 5 Gyr. The resulting stellar radii from the DSEP grid are used in conjunction with values of Rp/R∗R_{p}/R_{*} from Borucki et al. (2011) to compute planetary radii.

Estimates of TeffT_{eff} are inferred from our optical spectra. We compare our visible spectra to a grid of models of K- and M-dwarf spectra generated by the BT-SETTL version of PHOENIX (Allard et al., 2010). Details of the comparison, sub-grid interpolation, and error calculations are described in Lépine et al. (in prep). The grid of models spans TeffT_{eff} of 3000-5000 K in steps of 100 K, log gg values of 0.0-5.0 in steps of 0.5 dex, and metallicities of [M/H] = -1.5, -1, -0.5, 0, +0.3, and +0.5. α\alpha/Fe is taken to be solar. We report the TeffT_{eff} of the best-fit interpolated model, and the standard deviation of TeffT_{eff} among the set of interpolated models that are nearby in parameter space in Table 1.

Our calculated values of TeffT_{eff} are shown in Figure 7 vs. the temperature given in the KIC (Brown et al., 2011). BT-SETTL temperatures are systematically lower than KIC temperatures by 110−35+15110^{+15}_{-35} K for the dwarf stars, and 150−35+10150^{+10}_{-35} K for the giant stars. Errors are calculated by bootstrap resampling. This is consistent with other determinations using the atmospheric models of Allard et al. (2010), including other determinations on Kepler KOI stars (M11). Our calculated temperatures are tightly correlated with KIC temperatures. When KIC temperatures are corrected for our observed offset, the standard deviation of the difference in calculated temperatures (σKIC−Phoenix\sigma_{KIC-Phoenix}) is 90 K, suggesting that the KIC temperatures for low-mass stars are more precise but are less accurate than suggested by Brown et al. (2011). For field stars with visible-wavelength spectra, we adopt our calculated TeffT_{eff} values, and for stars with exoplanet candidates we use the TeffT_{eff} from M11. For the remaining stars we adjust the KIC effective temperatures of Kepler stars downward randomly by 110−35+15110^{+15}_{-35} K to keep the temperatures consistent with those of the KOI stars and those with spectra in our sample. This offset is randomized to account for errors in the systematic difference between temperatures calculated from our spectra and those listed in the KIC.

Following H11, we compute the planet occurrence with 2R\earth<RP<32\earth2R_{\earth}<R_{P}<32_{\earth} and P<50P<50 days around stars with 3400<Teff<41003400<T_{eff}<4100 using Equations 5 - 6. Again following H11, we exclude stars with KP>15K_{P}>15 where the accuracy of the planet candidate parameters are more questionable and the false positive rate is higher (Morton & Johnson, 2011; Borucki et al., 2011). We calculate the standard deviation of the frequency using a Monte Carlo analysis. Stellar parameters are perturbed randomly (see above) accounting for errors from M11 on KOI metallicity and TeffT_{eff}, and random errors from derived from TeffT_{eff} fits (see Figure 7) to our spectra. Other stars are given a random error of 90 K. We perturb transit parameters RP/R∗R_{P}/R_{*} and period according to errors given by Borucki et al. (2011). Planetary radii are recalculated from perturbed values of RP/R∗R_{P}/R_{*} and R∗R_{*}.

We remove planets from the KOI sample using the false positive probabilities from Morton & Johnson (2011) (e.g., a planet candidate with a 5% false positive probability is removed in 5% of the simulations). We remove giant stars from the sample using the calculated photometric likelihoods (Section 3.3) for each star, such that a star with a 10% likelihood of being a giant star will be removed from the sample in 10% of the Monte Carlo (MC) simulations. This also applies to stars with detected planet candidates, causing the planet to be removed, i.e. we consider the planet detection to be a false positive if the star is a giant. The number of KOIs and target stars simulated varies somewhat for each Monte Carlo run, but there are typically ≃14\simeq 14 KOIs around ≃1300\simeq 1300 stars in a given simulation.

We find that there are 0.37±0.080.37\pm 0.08 planets (with 2R\earth<RP<32R\earth2R_{\earth}<R_{P}<32R_{\earth} and P<50P<50 days) per star in the temperature range 3400<Teff<41003400<T_{eff}<4100. For comparison we run an additional Monte Carlo simulation but only remove giant stars with KIC log g>4.0g>4.0 as in H11. This test yields a planet occurrence of 0.26±0.050.26\pm 0.05, slightly lower than when giant stars are properly removed. To test how our results depend on our choice of stellar radii model (DSEP) we also run two simulations using the Yonsei-Yale (Demarque et al., 2004) isochrones: one with giant stars removed as explained above and another removing just giants with KIC log g>4.0g>4.0. The runs using Yonsei-Yale are included because their models are commonly used to derive radii for Kepler targets (e.g. Batalha et al., 2012). However, radii and masses derived from DSEP are a far better match to observations of late-type stars (Dotter et al., 2008; Feiden et al., 2011), and planet occurrence calculated using the DSEP models should be considered more reliable. The resulting Monte Carlo distributions are shown in Figure 8.

2. Parametric Likelihood estimation

We also perform a parametric maximum likelihood estimation of the fraction of stars with planets with radii 2R⊕<R<32R⊕2R_{\oplus}<R<32R_{\oplus} and orbital period P<50P<50 d (see H11 for a similar analysis). For discrete, binomial (detection or non-detection) events, the likelihood is expressed as:

where the first product is of detections, the second is of non-detections, and ρi\rho_{i} is the probability that a planet with properties in the appropriate ranges orbits the iith star and is detected by Kepler to transit. For this formulation, we have assumed that ρ≪1\rho\ll 1. We adopt the specific power-law form dN=CRi−αP−βdln⁡R⋅dln⁡PdN=CR_{i}^{-\alpha}P^{-\beta}d\ln R\cdot d\ln P for the intrinsic distribution of planets. If both α\alpha and β\beta are >0>0 then the normalization factor CC is given by:

where ff is the total planet occurrence. We do not model multi-planet systems; that level of analysis is not justified given the large uncertainties in our parameters.

Following the usual procedure, we maximize the logarithm of LL:

where Dj(Rj,Pj)D_{j}(R_{j},P_{j}) is the probability of detecting the jjth planet around its host star, including the geometric factor (note Dj(Rj,Pj)=djpjD_{j}(R_{j},P_{j})=d_{j}p_{j}, see Equation 5 and 7), and

We then substitute Equation 9 for CC. Ignoring terms that do not depend on α\alpha, and thus do not affect its maximum likelihood value, we find the following quantity must be maximized:

The simultaneous solution for the planet occurrence is found by maximizing the terms that depend on ff and is simply

where NpN_{p} is the number of detected planets. Equation 15 immediately suggests a reduction in the last terms of Equations 13 and 14 to NpN_{p}, which is independent of α\alpha and β\beta and can be ignored.

Because there are too few systems in our sample to get a robust estimate of β\beta, we fix β=0\beta=0 with a cut-off at P1=1P_{1}=1 d, consistent with the findings of previous analyses (Cumming et al., 2008; Wolfgang & Laughlin, 2011, H11). Equation 15 becomes:

Artificial Monte Carlo data sets suggest that ff is robustly recovered, but that recovered values of α\alpha are biased downwards. Using the cool KOIs defined here, stellar parameters derived as explained above, and Monte Carlo data sets generated by sampling with replacement, we find that f=0.34±0.08f=0.34\pm 0.08, consistent with our nonparametric calculation. As before, we repeat our Monte Carlo simulation but only removing giant stars with KIC log g>4g>4, and another run using the Yonsei-Yale evolutionary tracks (Demarque et al., 2004) instead of those of DSEP. The resulting Monte Carlo distributions are shown in Figure 8.

Planet-host Metallicities

SL11 use g−rg-r vs. J−HJ-H colors to conclude that late-type (J−H≃0.62J-H\simeq 0.62) exoplanet hosts are redder and more metal-rich than stars without transiting planets. Because giant stars have bluer g−rg-r colors at a given J−HJ-H color (Bessell & Brett, 1988; Gilbert et al., 2006), a significant number of giant star interlopers in their sample will cause field stars to appear metal poor. Giant stars have stellar radii 10-100 times larger than dwarfs, significantly reducing the depth in a light curve for a given transiting planet, making it much less likely that they will appear as KOIs (with the exception of false positives).

We can test their findings by creating a “pure” dwarf sample, and comparing its color distribution to that of the KOI sample. Our KP−J>2K_{P}-J>2 spectroscopic sample is systematically redder in J−HJ-H than the 0.56<J−H<0.660.56<J-H<0.66 bin used in SL11, preventing us from making a direct comparison. Instead, we construct samples of giants and dwarfs in the J−H≃0.62J-H\simeq 0.62 bin based on our photometric determination of luminosity class. For both the dwarf and giant samples, we select Kepler target stars with photometry in all bands used in our photometric assignment of luminosity class (J,H,K,D51,g,r,J,H,K,D51,g,r, and all four WISE bands). We then select stars with a >90%>90\% likelihood of being dwarfs based on our analysis in Section 3.3. The resulting dwarf sample is ≃2500\simeq 2500 stars. This sample may still contain giants. We add Poisson noise to the photometry of both the training sets and the Kepler 0.56<J−H<0.660.56<J-H<0.66 target star sample, and take random subsamples of both training sets. We then reapply these subsamples to the modified photometry of the Kepler sample. We repeat this process 1000 times. By analyzing the number of giant stars in each of these new samples we find that our dwarf sample is <1%<1\% giant stars at 95% confidence, ignoring possible systematic errors.

We use this dwarf sample, following the method of SL11, to compare the g−rg-r colors at a given J−HJ-H (a proxy of effective temperature) of the exoplanet host stars with our dwarf sample. Figure 9 shows g−rg-r colors as a function of J−HJ-H colors for the dwarf, giant, planet-host, and KIC log g>4.0g>4.0 sample. We find no significant difference in color between the KOI stars and our dwarf sample. Unlike the KIC log g>4.0g>4.0 sample, the locus of our photometrically selected dwarf sample is consistent with the locus of the KOI sample at J−H≃0.62J-H\simeq 0.62. For stars with KP−J>2.0K_{P}-J>2.0 we find an offset in g−rg-r color of only −0.05±0.03-0.05\pm 0.03 between the spectroscopically confirmed dwarfs and late-type KOI stars hosting Earth-to-Neptune sized planets. When we use our photometric sample of dwarfs in the J−H≃0.62J-H\simeq 0.62 bin we find an offset of 0.01±0.020.01\pm 0.02 and we can rule out the offset of 0.08 seen by SL11 with >99.7%>99.7\% certainty. Our photometric selection may remove some metal-poor dwarfs. However, even when we include stars ≥60%\geq 60\% likelihood of being dwarfs, which will necessarily increase the number of interloping giants, the offset is still only 0.03±0.020.03\pm 0.02 (consistent with zero offset).

In spite of the low giant fraction for dim Kepler target stars, it is not sufficient to simply repeat the SL11 analysis exclusively for stars with KP>14K_{P}>14. Since SL11 only examine stars with KIC log g>4.0g>4.0, it is far more important to investigate the g−rg-r distribution of misidentified giants in the 0.56<J−H<0.660.56<J-H<0.66 color range (i.e. giant stars that were assigned log g>4g>4 in the KIC). In fact the fraction of misidentified dim giant stars in their J−H≃0.62J-H\simeq 0.62 bin is higher (12%12\%), than it is for the KP−J>2K_{P}-J>2 star sample. We show why this is the case in Figure 10, which shows the distribution of giants, dwarfs, and misidentified giants in J−HJ-H vs. g−rg-r space. Misidentified giants are more concentrated at 0.58<J−H<0.630.58<J-H<0.63. Further, the misidentified giants in this J−HJ-H range are much more blue than the dwarfs in the same range. Thus by selecting a color bin centered on J−H=0.62J-H=0.62, SL11 are over-selecting giant stars, even after applying a KIC log g>4g>4 cut (≃15%\simeq 15\% of this sample are giant stars). This concentration of misidentified giants is the most likely explanation for the color offset seen by SL11, and also explains why the same g−rg-r offset is not seen at redder J−HJ-H colors (see Figure 9).

Discussion

We use visible-wavelength spectra to determine the properties of a subset of late-type Kepler target stars. We separate giants from dwarfs by comparing our spectra to those of stars with known luminosity class, and determine effective temperatures by comparing with PHOENIX model spectra. We extend our results to a larger collection of Kepler stars using photometry from the KIC, 2MASS, and WISE catalogs. We apply our luminosity class determinations to refine estimates of the planet occurrence around stars with 3400<Teff<41003400<T_{eff}<4100, and compare the colors – and hence metallicities of stars with and without detected Earth and Neptune sized planets. We draw four major conclusions:

1. Among stars redder than KP−J=2K_{P}-J=2 (≃\simeq K5 and later), bright (KP<14K_{P}<14) stars are predominantly (96±1%96\pm 1\%) giants, while dim stars (KP>14K_{P}>14) are predominantly (93±3%93\pm 3\%) dwarfs. These fractions improve somewhat when we consider stars with KIC determined log g>4.0g>4.0 (74±8%74\pm 8\% and 97±2%97\pm 2\% respectively). Overall, 52±3%52\pm 3\% of Kepler stars with KP−J>2K_{P}-J>2 are giants. However, only 10±2%10\pm 2\% of said stars with KIC log g>4.0g>4.0 are giants, a consequence of the large number of late-type stars lacking any temperature or log gg values in the KIC.

2. KIC effective temperatures, based on the models of Castelli & Kurucz (2004) and grizgriz and JHKJHK photometry, are systematically higher by 110−35+15110^{+15}_{-35} K compared to those derived from our own spectra and PHOENIX BT-SETTL atmosphere models (Allard et al., 2010).

3. Adopting the temperature scale from BT-SETTL and radii/masses from the Dartmouth Stellar Evolution Database (Dotter et al., 2008) and removing stars we identify as giants based on nonparametric and parametric Monte Carlo calculations we find a planet occurrence rate of ≃0.36±0.08\simeq 0.36\pm 0.08 for planets with radii 2R\earth<RP<32R\earth2R_{\earth}<R_{P}<32R_{\earth} and periods 1<P<501<P<50 days per star in the temperature range 3400<Teff<41003400<T_{eff}<4100. Using the KIC determined luminosity classes leads to a somewhat lower planet occurrence of 0.26±0.050.26\pm 0.05.

4. The g−rg-r colors of exoplanet host stars at J−H≃0.62J-H\simeq 0.62 are consistent with an unbiased sample of Kepler dwarf stars, ruling out any large difference between hosts of Earth-to-Neptune sized planets and those without any detected planets.

Surprisingly, there are hundreds of stars in our photometric sample that could have been easily identified as giants with KIC photometry, but were assigned log g>4g>4. The KIC primarily uses g−D51g-D51 vs g−rg-r colors to identify giants, and many late-type stars with KIC log g>4.0g>4.0 have g−D51g-D51 vs g−rg-r colors consistent with giants (and inconsistent with dwarfs).

Our calculated giant fraction is consistent with other independent measurements. Gaidos et al. (2012) compare radial velocity data from M2K (Apps et al., 2010; Fischer et al., 2012) to Kepler results and note that the completeness of the coolest Kepler target stars may be quite low (≃50%\simeq 50\%), much of which could be explained by an underestimate of the frequency of giant stars. Additionally, Ciardi et al. (2011) find that bright Kepler M stars are“predominantly giants, regardless of the KIC classification” based on JHKJHK photometry alone. Our giant fraction is also consistent with the current understanding of Galactic structure: based on a simulation from TRILEGAL (Girardi et al., 2005), ≃92\simeq 92 of stars near the center of the Kepler field with KP<14K_{P}<14 and KP−J>2.0K_{P}-J>2.0 are giants.

Interestingly, we find two KOIs with colors consistent with giant stars. KOI 667 and KOI 977 both fall within our giant training set in multiple color relations, and outside our dwarf training set. M11 identify KOI 977 as a giant, and they also note that KOI 667 consisted of 5 objects within 6\arcsec6\arcsec which may be contaminating 2MASS or WISE photometry. One of these objects could be an eclipsing binary, diluted by the other stars. KOI 667 also has a relatively high (10%) false positive probability based on Galactic structure models (Morton & Johnson, 2011).

Our values of TeffT_{eff} are consistent with results reported elsewhere also using BT-SETTL, including observations of the late-type KOIs with near-infrared spectra M11. These authors find a similar systematic offset of 123−32+24123^{+24}_{-32} K between their temperatures and KIC assigned temperatures. KIC temperatures are based on the models of Castelli & Kurucz (2004) and the evolutionary tracks of Girardi et al. (2000), which, although reliable for solar-mass stars, are untrustworthy for stars with Teff<3750T_{eff}<3750 K (Brown et al., 2011).

Our planet occurrence estimate is slightly higher than that of H11, who, using results from Kepler, find a planet occurrence rate of 0.30±0.080.30\pm 0.08 for stars with 3600<Teff<41003600<T_{eff}<4100. The difference is primarily due to reliance on luminosity class determinations by Brown et al. (2011), which we find to be inaccurate. However, the difference is within 1σ1\sigma. For both our work and that of H11, errors are dominated by the low number of late-type stars (and therefore planets around them) in the Kepler field and very high random (∼35%\sim 35\%) errors in stellar radii.

In addition to random errors (e.g. stellar radii and Rp/R∗R_{p}/R_{*}) that are included in our Monte Carlo simulation, there may be large systematic uncertainties in atmosphere models and evolutionary tracks, which can change the resulting frequency. When we use the Yonsei-Yale isochrones, it decreases our planet occurrence by ≃0.08\simeq 0.08. Interestingly, this difference is similar in size to the random errors in our Monte Carlo analysis (≃0.08\simeq 0.08), and the difference between proper giant removal and using KIC log g>4.0g>4.0 (≃0.10\simeq 0.10). This suggests that giant star removal, improved stellar characterization of the dwarf stars, and use of reliable stellar models of late-type stars are of roughly equal importance in characterizing the planet occurrence around very cool stars.

The lack of a strong correlation between host-star metallicity and the presence of Earth-to-Neptune sized planets is consistent with what is found for solar-type stars, e.g. Mayor et al. (2011). This also matches the findings of M11, who determine that among the late-type Kepler exoplanet hosts in our sample the median [M/H] is −0.11±0.02-0.11\pm 0.02. This distribution is consistent with stars in the solar neighborhood (−0.05-0.05 to −0.15-0.15, Johnson & Apps, 2009; Schlaufman & Laughlin, 2010; Casagrande et al., 2011). A metallicity difference could only be present if Kepler target stars are significantly more metal poor than stars in the solar neighborhood. As explained in Gaidos et al. (2012), Kepler late K and M stars are <250<250 pc from the Sun, and ≲60\lesssim 60 pc above the galactic plane. Most of the stars will be in the thin disk, and have metallicities similar to that of the solar neighborhood.

Our analysis of the g−rg-r colors of planet hosts contradicts the results of SL11, who find a 4σ4\sigma difference between g−rg-r colors of late-type exoplanet hosts and stars with no exoplanets present. Their result is most likely an artifact of the large number of stars which were misclassified as dwarfs in the KIC. SL11 state that their result can be reproduced if their sample of KIC log g>4g>4 stars is between 10% and 30% giants, which they calculate by adding stars with KIC log g<4g<4 stars (test giant stars) into their control sample, and measuring the resulting g−rg-r color offset. We find that the giant fraction is above 10%10\% for this color range. Further, if the KIC log g>4g>4 sample that SL11 used was significantly contaminated with giants, the sample will have bluer colors than a true dwarf sample. Adding test giants (to measure the resulting color change) to an already giant-star contaminated sample will create smaller changes in the overall color of a sample than if the sample had contained only dwarf stars. Thus more test giant stars will be required to produce a given color offset, creating an artificially high estimate for the level of giant contamination required to produce the observed color difference.

Although the g−rg-r colors of exoplanet hosts in our sample are consistent with our dwarf sample, we cannot rule out small offsets (≲0.05\lesssim 0.05) in g−rg-r color. It is possible that any metallicity effect is sufficiently small that it is diluted to non-detection by the large number of undetected exoplanets in the dwarf sample. As Kepler continues to discover planets of smaller radii and at larger orbital periods, the answer may become more clear.

References