Prevalence of Earth-size planets orbiting Sun-like stars

Erik A. Petigura, Andrew W. Howard, Geoffrey W. Marcy

S1 The Best42k Stellar Sample

We restrict our planet search to Sun-like stars with well-determined photometric properties and low photometric noise. We select stars having revised Kepler Input Catalog (KIC) parameters. Effective temperatures are based on the Pinsonneault et al. (?) revisions to the KIC effective temperatures. Surface gravities are based on fits to Yonsei-Yale stellar evolution models (?) assuming [Fe/H] = −0.2-0.2. Further details regarding isochrone fitting can be found in Batalha et al. (?); Burke et al., submitted; and Rowe et al., in prep. These revised stellar parameters are tabulated on the Exoplanet Archive with the prov_prim flag set to “Pinsonneault.” Out of the 188,329 stars observed at some point during Q1–Q15, we selected stars that:

Have revised KIC stellar properties. (155,046 stars),

TeffT_{\rm eff} = 4100–6100 K (63,915 stars), and

Figure S1 shows the position of the 155,046 stars with revised stellar properties along with the “solar subset” corresponding to G and K dwarfs. Figure S2 shows the distribution of brightness and noise level of the Best42k stellar sample.

S2 Planet Search Photometric Pipeline

We search for planet candidates in the Best42k stellar sample using the TERRA\tt TERRA pipeline described in detail in Petigura & Marcy (2012) and in Petigura, Marcy, and Howard (2013; P13, hereafter) (?, ?). We review the major components of TERRA\tt TERRA below, noting the changes since P13.

TERRA\tt TERRA begins by conditioning the photometry in the time-domain. TERRA\tt TERRA first searches for single cadence outliers, mostly due to cosmic rays. TERRA\tt TERRA also searches for abrupt drops in the raw photometry known as Sudden Pixel Sensitivity Drops (SPSDs) discussed by Stumpe et al. (?). SPSDs are particularly challenging since they mimic transit ingress, and aggressive attempts to remove them run the risk of removing real transits. TERRA\tt TERRA removes the largest SPSDs, but they remain a source of non-astrophysical false positives that we remove during manual triage (Section S3.2).

TERRA\tt TERRA also removes trends longer than ∼\sim10 days. In P13, this high-pass filtering was implemented by fitting a spline to the raw photometry with the knots of the spline separated by 10 days. But in this work we employ high-pass filtering using Gaussian Process regression (?), which gives finer control over the timescales removed. We adopt a squared exponential kernel with a 5-day correlation length. After this high-pass filter, TERRA\tt TERRA identifies systematic noise modes via principle components analysis on large number of stars.

S2.2 Grid-based transit search

We search for periodic box-shaped dimmings by evaluating the signal-to-noise ratio (SNR) of a putative transit over a finely-spaced grid of period, PP; epoch, t0t_{0}; and transit duration, ΔT\Delta T. In P13, we searched over a period range, PP = 5–50 days, and over transit durations ranging from 1.5–8.8 hr. But in this work, we extend our search in orbital period to PP = 0.5–400 days. Since we search over nearly three decades in orbital period, and because transit duration is proportional to P1/3P^{1/3}, we let the range of trial transit durations vary with period. We break our period range into 10 equal logarithmic intervals. Then, using photometrically determined parameters for each star, namely M⋆M_{\star} and R⋆R_{\star}, we compute an approximate, expected transit duration (ΔTcirc\Delta T_{\text{circ}}) for the simple case of circular orbits with impact parameter, b=1b=1. However, we actually search over ΔT\Delta T = 0.5–1.5 ΔTcirc\Delta T_{\text{circ}} to account for a range of impact parameters and orbital eccentricities and for mis-characterized M⋆M_{\star} and R⋆R_{\star}. As an example, Table S1 shows our trial ΔT\Delta T for a star with solar mass and radius.

S3 Data Validation

If TERRA\tt TERRA detects a (PP, t0t_{0}, ΔT\Delta T) with SNR > 12, we flag the light curve for additional scrutiny. While the grid-based component of TERRA\tt TERRA is well-matched to exoplanet transits, there are other phenomena that can produce SNR > 12 events and contaminate our planet sample. We distinguish between two classes of contaminates: “astrophysical false positives” such as diluted eclipsing binaries (EBs), and “non-astrophysical false positives” such as noise that can mimic a transit. We establish a series of quality control measures called “Data Validation” (DV), designed to remove formally strong dimmings (i.e. SNR > 12) found by the blind photometric pipeline that are not consistent with an astrophysical transit. DV consists of two steps:

Machine triage: Select potential transits by automated cuts.

Manual triage: Manually remove light curves that are inconsistent with a Keplerian transit.

Manual triage is accomplished by inspection of DV summary plots which contain numerous useful diagnostics necessary to warrant planet status. The diagnostics permit a multi-facted evaluation of the integrity (as a potential planet candidate) of a given dimming identified by the photometric pipeline. Figure S3 shows an sample DV report, this for KIC-5709725 that passed examination.

The product of the DV quality control is a list of “eKOIs,” for which most instrumental events identified preliminarily and erroneously by the photometric pipeline have been rejected. The resulting planet candidates are analogous to the KOIs Kepler Project. Astrophysically plausible causes (i.e. transiting planets and background eclipsing binaries) are retained among our eKOIs. We address astrophysical false positives in Section S5.

Prior to the identification of final eKOIs, we carry out machine triage to identify a set of “Threshold Crossing Events” (TCEs) that can be classified by a human in a reasonable amount of time. TCE status requires a SNR > 12; however, we find that 16227 light curves (out of the 42000 target stars) meet this criterion. Outliers and correlated noise are responsible for the majority of SNR > 12 events. We show set of diagnostic plots for such an outlier in Figure S4. Here, an uncorrected sudden pixel sensitivity dropoff at tt = 365.3 days, raises the noise floor to SNR∼\sim15 for P≲100P\lesssim 100 days. Its contribution to SNR is averaged down for shorter periods.

We flag such outliers by comparing the most significant period, PmaxP_{\text{max}}, to nearby periods. We call the ratio of the maximum SNR to the median of the next tallest five peaks between [Pmax/1.4,Pmax×1.4[P_{\text{max}}/1.4,P_{\text{max}}\times 1.4] the s2n_on_grass statistic. We require s2n_on_grass > 1.2 for TCE status. After that cut, 3438 TCEs remain. We also require PmaxP_{\text{max}} > 5 d, which leaves 2184 TCEs.

S3.2 Manual Triage

The sample of 2184 TCEs has a significant degree of contamination from non-astrophysical false positives. In P13, we relied on aggressive automatic cuts that removed nearly all of the non-astrophysical false positives (final sample was ∼90%\sim 90\% pure). However, by comparing our sample to that of Batalha et al. (?), we found that these automatic cuts were removing a handful of compelling planet candidates.

In this work, we aim for higher completeness and rely more heavily on visual inspection of light curves. We assess whether a TCE is due to a string of three or more transits or instead caused by outlier(s) such as SPSDs. Figure S5 shows an example of a light curve that passed machine triage, but was removed manually. During manual triage, we do not attempt to distinguish between planets and astrophysical false positives. The end product is a list of 836 eKOIs, which are analogous to KOIs produced by the Kepler Project in that they are highly likely to be astrophysical in origin but false positives have not been ruled out.

S4 KOIs That Fail Data Validation

As a cross-check of our DV quality control methods, we performed the same inspection on 235 KOIs that had been identified by the Kepler Project and which appear currently in the online Exoplanet Archive (?). These KOIs have periods longer than 50 days, representative of long period transiting planets that enjoy a reduced number transits (compared to short-period planets) during the 4-year lifetime of the Kepler mission. We found four KOIs, 2311.01, 2474.01, 364.01, and 2224.01, that are not consistent with an astrophysical transit. We show the raw light curves around the published ephemerides in Figure S6. All four have RP ≤ 2.04 R⊕R_{P}~{}\leq~{}2.04~{}R_{\oplus}, and three have P ≥ 173P~{}\geq~{}173 days. Due to the small number of KOIs near the habitable zone, inclusion of these KOIs would bias occurrence measurements upward by a large amount. Vetting all 3000 KOIs in the Exoplanet Archive requires a substaintial effort, and is beyond the scope of this paper. However, these four KOIs are a reminder than detailed, expert, and visual vetting of DV diagnostics existing KOIs is useful.

S5 Removal of Astrophysical False Positives

We take great care to cleanse our sample of false positives (FPs). Some transits are so deep (δF≳10%\delta F\gtrsim 10\%) that they can only be caused by an EB. However, if an EB is close enough to a Kepler target star, the dimming of the EB can be diluted to the point where it resembles a planetary transit. For each eKOI, we assess four indicators of EB status. Here, we list the indicators along with the number of eKOIs removed from our planet sample due to each cut:

Radius too large (115). We consider any transit where the best fit planet radius is larger than 20 R⊕R_{\oplus} to be stellar. Planets are generally smaller than 1.5 RJ = 16 R⊕1.5~{}R_{\text{J}}~{}=~{}16~{}R_{\oplus} especially for PP > 5 days, where planets are less inflated. Our cut at 20 R⊕R_{\oplus} allows some margin of safety to account for mis-characterized stellar radii.

Secondary eclipse (44). The expected equilibrium temperature for a planet with PP > 5 days is too small to produce a detectable secondary eclipse. Therefore, the presence of a secondary eclipse indicates the eclipsing body is stellar. We search for secondary eclipses by masking out the primary transit and searching for additional transits at the same period. If an eKOI, such as KIC-8879427 shown in Figure S7, has a secondary eclipse, we designate it an EB.

Variable depth transits (27). Since Kepler photometric apertures are typically two or three pixels (8 or 12 arcsec) on a side, light from neighboring stars can contribute to the overall photometry. A faint EB, when diluted with the target star’s light, can produce a dimming that looks like a planetary transit. If the angular separation between the two stars is large enough, the EB will contribute a different amount of light at each Kepler orientation. For eKOIs like KIC-2166206 shown in Figure S8, the contribution of a nearby EB results in a season-dependent transit depths. Since the target apertures are defined to include nearly all (≳90%\gtrsim 90\%) of the light from the target star, variations between quarters produce a negligible effect on transits associated with the target star, i.e. fractional changes of ≲1%\lesssim 1\%.

Centroid offset (31). Kepler project DV reports exist for nearly all (609/650) of the eKOIs that survive the previous cuts and are available on the Exoplanet Archive. We inspect the transit astronomy diagnostics (?) for significant motion of the transit photocenter in and out of transit. eKOIs with significant motion are designated false positives.

We remove a small number of eKOIs (11) with V-shaped transits. Since planets are so much smaller than their host stars, ingress/egress durations are short compared to the duration of the transit, i.e. planetary transits are box-shaped. Stellar eclipses tend to be V-shaped. Limb-darkening, the 30-minute integration time, and the possibility of grazing incidence blur this distinction. We assessed transit shape visually rather than using more detailed approaches based on light curve fitting and models of Galactic structure (?, ?). Only 1.3% of eKOIs are removed in this way and are a small effect compared to other uncertainties in our occurrence measurements.

We also remove five eKOIs with large TTVs. Since TERRA\tt TERRA’s light curve fitting assumes constant period, fits are biased toward smaller planet radii in the presence of transit timing variations ≳ΔT\gtrsim\Delta T. If the resulting error is ≳25%\gtrsim 25\%, we remove that eKOI. While these eKOIs are likely planets, our constant period model results in a significant bias in derived planet radii. Given the small number of eKOIs with such large TTVs, our decision to remove them has small effect on our statistical results based off of hundreds of planets.

We compute planet occurrence from the 603 eKOIs that survive the above cuts. We show the distribution of TERRA\tt TERRA candidates and FPs on the PP–RPR_{P} plane in Figure S9. All 836 eKOIs are listed in Table S1. For each eKOI, Table S1 lists KIC identifier, transit ephemeris, FP designation, Mandel-Agol fit parameters, adopted host star parameters, and size. We also crossed checked our eKOIs against the catalog Kepler team KOIs accessed from the NASA Exoplanet Archive (?) on 13 September 2013. If the Kepler Project KOI number exists for an eKOI, we include it in Table S1.

S6 Planet Radius Refinement

We fit the phase folded transit photometry of each eKOI with a Mandel-Agol model (?). That model has three free parameters: RP/R⋆R_{P}/R_{\star}, the planet to star radius radio, τ\tau, the time for the planet travel a distance R⋆R_{\star} during during transit; and bb, the impact parameter. Following P13, we account for the covariance among the three parameters using an MCMC exploration of the parameter posteriors. The error on RP/R⋆R_{P}/R_{\star} in Table S1 incorporates the covariance with τ\tau and bb.

Because photometry alone only provides the RP/R⋆R_{P}/R_{\star}, knowledge of the planet population depends heavily on our characterization of their stellar hosts. We obtained spectra of 274274 eKOIs with HIRES on the Keck I telescope using the standard configuration of the California Planet Survey (Marcy et al. 2008). These spectra have resolution of ∼\sim50,000 and SNR of ∼\sim45/pixel at 5500 Å. We obtained spectra for all 62 eKOIs with PP > 100 days.

We determine stellar parameters using a routine called SpecMatch\tt SpecMatch (Petigura et al., in prep). SpecMatch\tt SpecMatch compares a target stellar spectrum to a library of ∼800\sim 800 spectra from stars that span the HR diagram (TeffT_{\rm eff} = 3500–7500 K; log⁡g\log g = 2.0–5.0). Parameters for the library stars are determined from LTE spectral modeling. Once the target spectrum and library spectrum are placed on the same wavelength scale, we compute χ2\chi^{2}, the sum of the squares of the pixel-by-pixel differences in normalized intensity. The weighted mean of the ten spectra with the lowest χ2\chi^{2} values is taken as the final value for the effective temperature, stellar surface gravity, and metallicity. We estimate SpecMatch\tt SpecMatch-derived stellar radii are uncertain to 10% RMS, based on tests of stars having known radii from high resolution spectroscopy and asteroseismology.

S7 Completeness

When measuring planet occurrence, understanding the number of missed planets is as important as the planet catalog itself. We measure TERRA\tt TERRA’s planet finding efficiency as a function of PP and RPR_{P} using the injection/recovery framework developed for P13. We briefly review the key aspects of our pipeline completeness study; for more detail, please see P13. We generate 40,000 synthetic light curves according to the following steps:

Select a star randomly from the Best42k sample,

draw (PP,RPR_{P}) randomly from log-uniform distributions over 5–400 d and 0.5–16 R⊕R_{\oplus},

draw impact parameter and orbital phase randomly from uniform distributions over 0–1,

inject the model into the “simple aperture photometry” of a random Best42k star.

We process the synthetic photometry with the calibration, grid-based search, and DV components of TERRA\tt TERRA. We consider a synthetic light curve successfully recovered if the injected (PP, t0t_{0}) agree with the recovered (PP,t0t_{0}) to 0.1 days. Figure S10 shows the distribution of recovered simulations as a function of injected planet size and orbital period.

Pipeline completeness is determined in small bins in (PP,RPR_{P})-space by dividing the number of successfully recovered transits by the total number of injected transits on a bin-by-bin basis. This ratio is TERRA\tt TERRA’s recovery rate of putative planets within the Best42k sample. Pipeline completeness is higher among a more rarefied sample of low noise stars. However, a smaller sample of stars yields fewer planets.

We show survey completeness for a dense grid of PP and RPR_{P} cells in Figure S11. Completeness falls toward smaller RPR_{P} and longer PP. Above 2 R⊕R_{\oplus}, completeness is greater than 50% even for the longest periods searched (except for the RPR_{P} = 2–2.8 R⊕R_{\oplus}, PP = 283–400 days bin). Completeness falls precipitously toward smaller planet sizes; very few simulated planets smaller than Earth are recovered. Compared to a 1 R⊕R_{\oplus} planet, a 2 R⊕R_{\oplus} planet produces a transit with 4 times the SNR and is much easier to detect. For planets larger than 2 R⊕R_{\oplus}, we note a gradual drop in completeness toward longer periods, that steepens at ∼\sim300 days. Above ∼\sim300 days, the probability that a two or more transits land in data gaps becomes appreciable, and the completeness falls off more rapidly.

Measuring completeness by injection and recovery captures the vagaries in planet search pipeline. Real and synthetic transits are treated the same way, up until the manual triage section. Recall from Section S3.2 that 836 of 2184 TCEs pass machine triage. We perform no such manual inspection of TCEs from the injection and recovery simulations. A potential concern is that a planet may pass machine triage, but is accidentally thrown out in manual triage. Such a planet would be missing from our planet catalog, but not properly accounted in the occurrence measurement by lower completeness. However, because our SNR > 12 threshold for TCE status is high, distinguishing non-astrophysical false positives and eKOIs is easy. Therefore, we consider it unlikely that they are cut during the manual triage stage, and do not expect the lack of manual vetting of the injected TCEs to bias our completeness measurements.

S8 Planet Occurrence

Here, we expand on the key planet occurrence results presented in the main text. We describe our method for extrapolation into the RPR_{P} = 1–2 R⊕R_{\oplus}, PP = 200–400 day domain. We give additional details regarding our measurement of the prevalence of Earth-size planets in the HZ. We also discuss two minor corrections to our occurrence measurements due to planets in multiplanet systems and false positives (FPs).

In the main text, we reported 5.7−2.2+1.7%5.7^{+1.7}_{-2.2}\% occurrence of planets with RPR_{P} = 1–2 R⊕R_{\oplus} and PP = 200–400 days based on extrapolation from shorter periods. The use of such extrapolation is supported by uniform planet occurrence per log⁡P\log P interval. Cumulative Planet Occurrence (CPO) is helpful to understand the detailed shape of the planet period distribution. If planet occurrence is constant per log⁡P\log P interval, CPO is a linear function in log⁡P\log P. The slope of the CPO conveys planet occurrence: the higher the planet occurrence, the steeper the slope of the CPO.

Figure S12 shows CPO for RPR_{P} = 2–4 R⊕R_{\oplus} planets. Planet occurrence increases with period from 5 days up to ∼10\sim 10 days, and is consistent with uniform for larger periods. This change in the planet period distribution was noted in previous work (?, ?, ?). We fit a line to the CPO from 50–200 days and extrapolate into the 200–400 day range. The extrapolation predicts 6.4−1.2+0.5%6.4^{+0.5}_{-1.2}\% occurrence, which agrees with our measured value of 5.0±2.1%5.0\pm 2.1\% to 1 σ\sigma. We estimate errors on our extrapolation by fitting subsets of the CPO that span half the original period range. We fit 100 subsections ranging from PP = 50–100 days up to PP = 100–200 days.

We also compare occurrence in the PP = 50–100 day, RPR_{P} = 1–2 R⊕R_{\oplus} domain based on extrapolation to our measured value. Figure S13 shows the CPO for RPR_{P} = 1–2 R⊕R_{\oplus} planets. We fit the CPO from PP = 12.5–50 days. This fit predicts an occurrence of 6.5−1.7+0.9%6.5^{+0.9}_{-1.7}\% in the 50–100 day range, in good agreement with our measured value of 5.8±1.6%5.8\pm 1.6\%. The uniformity in the occurrence of small planets as a function of period, lends support to the same kind of modest extrapolation into the RPR_{P} = 1–2 R⊕R_{\oplus}, PP = 200–400 day domain.

S8.2 Planet Occurrence in the Habitable Zone

We consider a planet to reside in the habitable zone if it receives a similar amount of light flux, FPF_{P}, from its host star as does the Earth. As described in the main text, we consider the most recent theoretical work on habitability of planets following the seminal work by Kasting (?, ?, ?, ?, ?).

We adopt an inner edge of the HZ at 0.5 AU for a Sun-like star where a planet would receive four times the light flux that Earth does. This inner edge is slightly more conservative than that found by Zsom et al. (?). The outer edge of the HZ less well understood. Kasting found the outer edge to be at 1.7 AU (?); Pierrehumbert and Gaidos (?) found it could extend to 10 AU for planets with thick H2 atmospheres. Here, we adopt an intermediate value of 2 AU for solar analogs where the stellar flux is 1/4 that incident on the Earth. This outer edge is consistent with the presence of liquid water on Mars in its past. Mars might still have liquid water today, if it were more massive. Thus following the theory of planetary habitability, we adopt a habitable zone for stars in general based on stellar flux between 4x and 1/4 the solar flux falling on the Earth: FPF_{P} = 0.25–4 F⊕F_{\oplus}.

The stellar light flux hitting a planet, FPF_{P}, depends linearly on stellar luminosity, L⋆L_{\star}, and inversely as the square of the distance between the planet and the star. Stellar luminosity, L⋆L_{\star}, is given by:

where σ=5.670×10−8\sigma=5.670\times 10^{-8} W m-2 K-4 is the Stefan-Boltzmann constant. In our study, the stellar radii and temperatures, TeffT_{\rm eff}, are computed two ways. We obtained high SNR spectra with high spectral resolution using the Keck Observatory HIRES spectrometer for all of the 62 stars that host planets with periods over 100 days, approaching the HZ. For those 62 stars, we performed a SpecMatch analysis (?) to determine TeffT_{\rm eff} and the surface gravity, log⁡g\log g, and metalicity, [Fe/H]. These stellar values were matched to stellar evolution models (Yonsei-Yale) to yield the radii and masses of the stars. The resulting values of stellar radii are uncertain by 10%, as determined by calibrations with nearby stars having parallaxes and hence having more accurately determined stellar radii. The values of TeffT_{\rm eff} are accurate to within 2%. Thus, summing the fractional errors in quadrature, the resulting stellar luminosities for the 62 stars (having PP > 100 days) are measured but carry uncertainties of 25%. For those stars without Keck spectra, we adopted photometric stellar radius and mass, for which the stellar radii are in error by 35% and the TeffT_{\rm eff} values are uncertain by 4%, giving errors in luminosity of 80%. We estimated the star-planet separation (aa) using PP, M⋆M_{\star}, and Kepler’s third law. The stellar light flux falling on a planet is now easily calculated from FPF_{P} ∝\propto L⋆L_{\star} /a2/a^{2}. In what follows, we quote the flux falling on a planet relative to that falling on the Earth.

We find 10 planets having radii 1–2 R⊕R_{\oplus} that fall within the stellar incident flux domain of the habitable zone, 0.25–4 F⊕F_{\oplus}. As a reference, we plot their phase folded light curves in Figure S15 along with the KIC identifier, period, radius, and stellar light flux. To compute the prevalence of such planets within the HZ, we apply the usual geometric correction for orbital tilts too large to cause transits, augmenting the counting of each transiting planet by a/R⋆a/R_{\star} total planets. We compute FPF_{P} for each synthetic planet in our completeness measurement study. Figure S16 shows stellar flux level and radii of the 10 habitable zone planets, having size 1–2 R⊕R_{\oplus}, along with the synthetic HZ planets from our completeness study. Because the number of synthetic trials is small for FPF_{P} < 1 F⊕F_{\oplus}, we compute occurrence using the 8 planets with FPF_{P} = 1–4 F⊕F_{\oplus}. We find 11±4%11\pm 4\% of Sun-like stars have a RPR_{P} = 1–2 R⊕R_{\oplus} planet that receives FPF_{P} = 1–4 F⊕F_{\oplus} light energy from their host star.

We account for the entire HZ (extending out to 0.25 F⊕F_{\oplus}) by extrapolating occurrence in FPF_{P}, assuming constant planet occurrence per log⁡P\log P interval. Figure S14 shows the CPO as a function of FPF_{P}. Planet occurrence is constant from ∼100 F⊕\sim 100~{}F_{\oplus} down to ∼4 F⊕\sim 4~{}F_{\oplus}, beyond which, small number fluctuations are significant. Assuming the occurrence of planets is constant in log⁡FP\log F_{P} implies that the same number of 1–2 R⊕R_{\oplus} planets have incident fluxes of 1–0.25 F⊕F_{\oplus} as have fluxes of 1–4 F⊕F_{\oplus} where we computed directly the occurrence of planets to be 11%. Thus, 22±8%22\pm 8\% of Sun-like stars have a RPR_{P} = 1–2 R⊕R_{\oplus} planet within our adopted habitable zone with fluxes of 0.25-4.0 F⊕F_{\oplus}.

S8.3 Occurrence Including Planets in Multi-planet Systems

For systems harboring more than one planet, TERRA\tt TERRA only detects the planet with highest SNR, i.e. the most significant planet. The actual rate of planet occurrence is higher than we report when the missed planets in these multi-transiting systems is included. (Note that this correction only applies to multi-transiting system and not all multi-planet systems.) We estimate the size of this effect using the Q12 sample of KOIs from the Kepler project, which includes stars with multiple planets. We selected the 1190 “candidates” with well-determined periods (σ(P)<0.1\sigma(P)<0.1 days) that orbit stars in the Best42k. In order to make a fair comparison between our planet sample and the Q12 sample, we computed the SNR of each of the 1190 candidates using TERRA\tt TERRA. We excluded 82 KOIs with SNR < 12, i.e. candidates that would have been deemed sub-significant by TERRA\tt TERRA.

For each planet in a multi-transiting system, we rank order each candidate by its “Relative SNR” defined as:

Figure S17 shows the distribution of Q12 candidates in the Best42k as points on the PP–RPR_{P} plane. We highlight points corresponding to the most significant planet. We assess the boost in planet counts due to multi-transiting systems for different domains in PP and RPR_{P}. For PP > 50 days and RPR_{P} < 4 R⊕R_{\oplus}, this multi-boost factor ranges from 21 to 28%, neglecting bins with fewer than 8 detected planets that suffer from small number fluctuations. Had we included additional planets, our occurrence measurements would rise by ∼25%\sim 25\%, which is comparable to or slightly smaller than the fractional occurrence error for small planets in long-period orbits.

S8.4 Correction due to False Positives

As discussed earlier, the sample of eKOIs is polluted by astrophysical false positives. Like the Kepler team, we do our best to identify and remove transits that are clearly due to eclipsing binaries, but cannot remove all eclipsing binary configurations. Thus, our sample, as well as those produced by the Kepler team, still contain a false positive component.

Fressin et al. (2013) addressed the contamination of the February 2012 Kepler Project sample of KOIs (?) by FPs that were not removed by the Kepler Project vetting process. FPs include background eclipsing binaries, physically associated eclipsing binaries (hierarchical triples), and physically associated stars, which themselves have a transiting planet. We consider the last scenario to be a FP because even though the transiting object is a planet, the radius is at least 1.4 times larger (1.4 corresponds stars of equal brightness). Fressin et al. (2013) added FPs to the Batalha et al (2012) sample of KOIs according to models of galactic structure, stellar binarity, and assumptions about the distributions of planets. Simulated FPs that would exhibit a detectable secondary eclipse or a significant centroid offset were removed, assuming the Kepler Project vetting process catches these FPs.

Fressin et al. (2013) found an overall FP rate of 8.8±1.9%8.8\pm 1.9\% for 1.25–2.0 R⊕R_{\oplus} planets and 12.3±3.0%12.3\pm 3.0\% for 0.8–1.25 R⊕R_{\oplus} planets. Again, note that this fractional occurrence correction is small compared to our reported errors for small planets in long-period orbits. Stars with bound companions with transiting planets are the dominant fraction of FPs for small planets (76% for 1.25–2.0 R⊕R_{\oplus} planets and 66% for 0.8–1.25 R⊕R_{\oplus} planets). FPs of this type are very difficult to identify. A Sun-like star with VV = 14.7 mag (typical for our sample) is 1 kpc away. The binary star separation distribution peaks at 50 AU (?) or 0.05 arcsec assuming a face-on orbit. Detecting companions separated by 0.05 arcsec is near the limits of current ground-based AO. Even if a companion was detected, we still wouldn’t know which star harbored the transiting planet.

We adopt a 10% FP rate for planets having PP = 50–400 days and RPR_{P} = 1–2 R⊕R_{\oplus}. Adopting a false positive rate that is constant with period is justified because the occurrence of Neptune-sized planets is approximately constant with period, as shown in the main text. In the context of the occurrence of Earth-size planets with PP = 200–400 days (5.7−2.2+1.7%5.7^{+1.7}_{-2.2}\%) and Earth-size planets in the HZ (22±8%22\pm 8\%), FPs contribute 10% fractional uncertainty and are secondary compared to statistical uncertainty.

References