A Southern Photometric Quasar Catalog from the Dark Energy Survey Data Release 2
Qian Yang, Yue Shen
I Introduction
Active galactic nuclei (AGNs) and their high-luminosity counterparts, quasars, are accreting supermassive black holes (SMBHs) at the center of massive galaxies. Understanding the evolution of the quasar population across cosmic time is crucial to understanding the physics of accretion and the coevolution of SMBHs and their host galaxies.
Large quasar surveys provide the necessary samples for measuring the abundance of quasars as functions of redshift and luminosity. In addition, these surveys enable a broad range of quasar science, such as quasar lens searches and their constraints on cosmology and the evolution of massive galaxies (Oguri et al. 2006; Oguri et al. 2012, e.g.,), finding projected quasar pairs (Hennawi et al. 2006a; Prochaska et al. 2013, e.g.,) and binary quasars (Hennawi et al. 2006b; Hennawi et al. 2010, e.g.,), and measuring quasar clustering (Martini & Weinberg 2001; Shen et al. 2007, e.g.,). Large quasar surveys also provide an opportunity to identify rare objects, such as extreme variability quasars (Rumbaugh et al. 2018, e.g.,), and study their fueling mechanisms (MacLeod et al. 2016; Yang et al. 2018, e.g.,). Finally, large samples of quasars are often used to define the celestial reference frame (Gaia Collaboration et al. 2018a, e.g.,).
Distant quasars have been discovered beyond redshift 7 (Mortlock et al. 2011; Bañados et al. 2018; Yang et al. 2020; Wang et al. 2021), where massive SMBHs had formed within less than one billion years after the Big Bang. A large sample of quasars over a broad range of redshifts enables the study of the evolution of SMBHs as well as the intergalactic medium. For example, the Ly forest in quasar spectra can be used to measure baryon acoustic oscillations as a probe for cosmology (Dawson et al. 2013).
While the Sloan Digital Sky Survey (York et al. 2000) has provided large samples of quasars in the northern hemisphere, there is a lack of large spectroscopically confirmed quasar samples in the southern hemisphere. There are over 750k quasars in the SDSS DR16 quasar catalog (Lyke et al. 2020). In contrast, there are less than 24k spectroscopically confirmed quasars in the southern hemisphere in the Million Quasars Catalog (Flesch 2019, v6.5) with . A large sample of quasars in the southern hemisphere will be important for quasar-related studies in the next few decades, given increasing investments of ground-based facilities covering the southern sky, in particular, the Vera C. Rubin Observatory Legacy Survey of Space and Time (Ivezić et al. 2019, LSST,).
The Dark Energy Survey (Abbott et al. 2018, DES;) is a wide-area visible and near-infrared imaging survey covering 5000 deg2 of the high Galactic latitude sky, mostly in the southern hemisphere. The multi-band deep DES photometry enables the photometric selection of a large quasar sample in the southern hemisphere. In this work, we perform a systematic selection of quasar candidates using public photometric data from DES over the 5000 deg2 wide-survey footprint, combined with publicly available near-infrared (NIR) and mid-infrared (MIR) photometric data. With these data, we classify quasars, galaxies, and stars in all DES detected photometric sources with probabilities and estimated photo-zs (for galaxy and quasar candidates).
The structure of this paper is as follows. In §II, we describe the imaging surveys and the training samples. We describe the selection methods in §III. We present the quasar catalog in §IV, and discuss the selection completeness and efficiency in §V. We conclude in §VI. Throughout this paper, we adopt a flat CDM cosmology with parameters , , and km s-1 Mpc-1. The Milky Way extinctions of extragalactic objects in DES bands are corrected using the dust reddening map of Schlegel et al. 1998. In this work, the term “quasar” is used to broadly refer to an unobscured broad-line AGN regardless of its luminosity. We also only consider quasar targets where the continuum emission is dominated by the quasar rather than the host galaxy.
II Data and Samples
We use the second public data release (DR2) of the DES (Abbott et al. 2021), including data from the DES wide-area survey covering 5000 deg2 of the southern Galactic cap in five broad photometric bands (). We use the DES DR2 coadded photometric catalog, including 691 million distinct astronomical objects, the vast majority of which are non-transient and non-moving objects. For the DES coadded photometry, we use the IMAFLAGS_ISO flag to remove unreliable detections, which is set if any pixel is masked in all of the contributing exposures for a give band. This flag is mainly set for saturated objects and objects with missing data (Abbott et al. 2018). The median coadded catalog point-source depth at in the bands are 25.0, 24.5, 23.7, 22.6, and 21.3, respectively (PSF mag). We use both PSF and AUTO photometry in DES depending on the fitting template class (see below).
For the NIR data, we make use of all public NIR imaging in the DES area, including the VISTA Hemisphere Survey (McMahon et al. 2013, VHS,), the VISTA Kilo-Degree Infrared Galaxy Survey (Edge et al. 2013, VIKING,), the UKIDSS Large Area Surveys (Lawrence et al. 2007, ULAS,), and the Ukirt Hemisphere Survey (Dye et al. 2018, UHS,). For these NIR surveys, we use the aperture-corrected magnitude in a 2 diameter circle. For areas not covered by these NIR surveys, we use the shallower all-sky Two Micron All Sky Survey (Skrutskie et al. 2006, 2MASS,) data https://doi.org/10.26131/IRSA2. Figure 1 shows the sky coverages of different NIR imaging surveys. The VHS survey covers most of the DES area in the southern sky. We summarize the depths of these NIR surveys in Table 1. We ignore the slight filter differences among different NIR surveys, which result in minor magnitude differences (normally less than 0.05 mag).
In the MIR, we use the unblurred coadds of the Wide-field Infrared Survey Explorer (Wright et al. 2010, WISE,) imaging data (Lang 2014; Meisner et al. 2019, unWISE,). We use the unWISE photometry from coadds of the WISE and NEOWISE (through the sixth year NEOWISE data release in 2020). The unWISE catalog has advantages over the AllWISE catalog https://doi.org/10.26131/IRSA1 since it is based on significantly deeper imaging and features improved photometric modeling in crowded regions (Schlafly et al. 2019). The 5 depth in AB magnitude in unWISE and bands are 21.7 and 20.9, respectively.
Gaia DR2 (Gaia Collaboration et al. 2018b) contains celestial positions for 1.7 billion sources, and parallaxes and proper motions for 1.3 billion sources. We use Gaia astrometry information to help rule out stars with detected proper motion or parallax.
II.2 Training Samples
We consider three object classes, quasars, stars, and galaxies, for which we build empirical color templates from training samples. Stars and most quasars are point-like objects, and galaxies are mostly extended sources. Each object is fit to three classes of color templates (quasar, star and galaxy). When fitting with quasar and star templates, we default to use DES PSF photometry for the object. When fitting with the galaxy template, we default to use the AUTO photometry in DES. At the faint end, for some objects without DES PSF photometry in some bands, we use the DES AUTO photometry for all three classes.
We then use spectroscopically confirmed quasars, stars and galaxies to build our color templates. For quasars, we start from the SDSS DR16 quasar catalog (Lyke et al. 2020), but remove unreliable high-redshift quasars misclassified by the SDSS pipeline. Specifically, we removed quasars that were only classified as “QSO” by the pipeline but not confirmed by visual inspection (most of these are pipeline misclassifications of low-redshift quasars or non-quasars). Next, we supplement spectroscopically confirmed quasars from the Million Quasars Catalog, v6.5 (Flesch 2019). We added sources with type as “Q” and “A”, which are broad-line quasars and broad-line Seyferts respectively. This supplementary sample is necessary because it includes confirmed high-redshift quasars at ; and quasars from the 2dF QSO Redshift Survey (Croom et al. 2004, 2QZ,), the 2dF-SDSS LRG and QSO survey (Croom et al. 2009, 2SLAQ,), the Australian Dark Energy Survey (Yuan et al. 2015, OzDES,), and the Large Sky Area Multi-object Fiber Spectroscopic Telescope (LAMOST) quasar catalog (Yao et al. 2019). The redshifts of the majority of the spectroscopically confirmed quasars are lower than 3.5 (99% of SDSS quasars). The number of spectroscopically confirmed quasars decreases rapidly with redshift, specifically from 9178 at to 634 at , and to 37 at . So at the high redshift end, using only these confirmed quasars may lead to strong biases from individual quasars. To improve the color coverage of quasars, we add simulated quasars (McGreer et al. 2013) at high redshift () The DECam Y band extends to 10700 . Therefore, at , the Ly emission of quasars starts to drop out of DES band.. The simulated quasar models include a broken power-law continuum, UV/optical emission lines, pseudo-continuum from FeII emission, and redshift-dependent Ly forest absorption due to neutral hydrogen. The numbers per redshift bin of simulated quasars are close to those of SDSS quasars at . We simulated a large number of quasars to ensure a sufficient statistical sample to avoid the impact of random fluctuations.
We consider contamination from stars and galaxies in our quasar selection. We use spectroscopic galaxies and stars from the SDSS DR16 (Ahumada et al. 2020). SDSS galaxies are representative of the low-redshift galaxy population but are limited to given the nature of optical SDSS surveys. However, the lack of representation of galaxies in the training sample does not affect our quasar selection, since these high- galaxies are typically much fainter in the observed-frame optical than our quasar targets. We supplement the sample with stars from the fifth data release of the LAMOST survey (Luo et al. 2015). We restrict to high-galactic-latitude stars in LAMOST with , as the DES footprint is all at . Compared with SDSS, the supplemental LAMOST stars are mainly at the bright end (). The star training sample is representative of different types of stars, from white dwarf to late-type stars. For example, more than half of the 68k white dwarfs from the Montreal White Dwarf Database http://www.montrealwhitedwarfdatabase.org are in our star training sample, and the other half are mainly out of the SDSS sky coverage. Among the 10k brown dwarfs compiled by Best et al. 2018 from the DwarfArchives http://DwarfArchives.org., 83% of them are in our star training sample.
We summarize the numbers of different classes of objects from different catalogs in Table 2. We cross-matched the sources with the imaging surveys described in §II.1 with a search radius of 2. The numbers of sources detected by different imaging surveys are also included in Table 2. Since most SDSS sources are in the northern sky and not covered by DES, we convolve the SDSS spectra with the DES filter curves to generate synthetic DES photometry in the training samples. Because the DES band spans from 9400 to 10700, we do not use spectra taken by the SDSS-I/II spectrographs (only up to 9200), and use spectra taken by the SDSS BOSS spectrographs (up to 10400 ) whenever applicable.
III Target Selection Algorithms
The photo- problem is a regression problem, relying on the description of the probability distribution of redshift for a specific class of objects. Quasar target selection is a classification problem, depending on the probability estimates for different classes of objects, such as quasars, stars, and galaxies.
We briefly describe the prior, likelihood, and posterior probabilities in our Bayesian analysis. In Bayes’ theorem, the posterior probability of model parameters given data can be written as
where is the likelihood and is the prior probability of model parameters . Here is the normalizing constant (also called evidence), and is usually ignored in the inference.
In our photo- regression problem, is the redshift, is the multi-dimensional relative fluxes (i.e., flux ratio wrt the flux in a reference band). The prior distribution, , is the predicted number distribution of an object class (i.e., quasar or galaxy), as a function of redshift . We describe the prior distribution in §III.2 and likelihood distribution in §III.3.
III.2 Prior Distribution
The number densities of quasars and galaxies depend on redshift and luminosity, which can be estimated using the observed luminosity functions of quasars and galaxies. The number density of stars depends on stellar type and luminosity, as well as locations in the sky. All these number densities refer to the absolute sky densities of the three classes of objects.
We implement the optical quasar luminosity function (QLF) from Palanque-Delabrouille et al. 2016, which is based on quasars over a wide redshift range of and magnitudes as faint as 22.5 mag in band. We extrapolate this QLF to the faint end and the high redshift end. We also tested the bolometric QLFs from Hopkins et al. 2007 and Shen et al. 2020. However, there are additional issues of utilizing these bolometric QLFs due to uncertainties in bolometric corrections and k-corrections, which are redshift and luminosity dependent. The QLF from Palanque-Delabrouille et al. 2016 works well for quasar photo- estimation over broad ranges of redshift and optical magnitude (Yang et al. 2017), as desired here.
We implement the galaxy luminosity function (GLF) from Montero-Dorta & Prada 2009 based on SDSS data. There are different subclasses of galaxies, such as late-type and early-type galaxies. Our galaxy training sample is not rigorously labelled with different sub-types, lacking information such as star formation rate or morphology. So we simply treat all galaxies as a single class in this work. Most of these SDSS galaxies are at , with a small fraction of them at higher redshifts. Since there are very few spectroscopically confirmed galaxies at (0.02%) in the training sample, we restrict to . Galaxies at are too faint in the observed-frame optical to contaminate our quasar selection.
We estimate the number density of stars for typical high Galactic-latitude fields In principle, our algorithm can implement a stellar number density as a function of Galactic latitude to improve the quasar selection in regions with higher stellar densities. For simplicity, in this work we only consider the typical stellar density at high Galactic latitudes, and leave this additional implementation to future work., from a Milky Way synthetic simulation with the Besançon model (Robin et al. 2003). Yang et al. 2017 showed that the simulated star number distribution is close to the observed distribution. We performed such a simulation of stars with DES filters in a 100 deg2 region with a central position at R.A.h and decl., which is close to the median central position of the DES survey. The number density of stars also depends on stellar types. Our star training sample is not well classified into different stellar types. Instead, we use color, , as an alternative parameter for different stellar types. We describe what colors are used specifically for stars in §III.3.
Using the QLF, GLF, and star simulations described above, we derive the quasar number (per deg2) prior distribution as a function of redshift, ; the galaxy number prior distribution as a function of redshift, ; and the star number prior distribution as a function of color, , in a set of magnitude bins. Our algorithm can be improved with better QLF and GLF for a wider range of redshifts and magnitudes. Our galaxy photo- can be further improved with galaxy training samples and GLFs labelled with different sub-types.
III.3 Likelihood Function
The key problem of our target selection/classification is to describe the likelihood of a series of attributes, , for a given redshift and magnitude. Specifically, in our algorithm represents the multi-dimensional relative fluxes.
The colors of quasars change as a function of redshift, due to the shift of quasar emission lines moving in and out of different filters. Quasar colors also change as a function of magnitude due to multiple reasons: (1) the colors of quasar at the faint end or at low redshift are more affected by their host galaxy light; (2) quasars are usually bluer when brighter; (3) the equivalent widths of quasar emission lines are often anti-correlated with the continuum emission (Baldwin 1997, i.e., the Baldwin effect;).
The colors of quasars at similar redshift and magnitude are usually similar. To fit the color distribution in multi-dimensional space, we can, for example, (1) fit the color in each color dimension with a Gaussian distribution, such as using the method (Richards et al. 2001, e.g.,); (2) fit the colors in multi-dimensional space with multivariate Gaussian distribution, such as using the multivariate method (Weinstein et al. 2004, e.g.,); (3) fit the colors in multi-dimensional space with mixture of multiple multivariate Gaussian distributions, such as the XDQSOz technique (Bovy et al. 2011; Bovy et al. 2012); (4) fit the colors in multi-dimensional space with Machine Learning techniques (Yèche et al. 2010; Shu et al. 2019, e.g.,); or (5) fit the colors in multi-dimensional space with more flexible parametric distribution, such as the multivariate Skew-t distribution (Yang et al. 2017).
Skewt-QSO is an algorithm for quasar selection and photo- estimation (Yang et al. 2017). The color distribution of quasars shows skewed and tail features mainly due to intrinsic dust reddening. Skewt-QSO describes the color distribution of quasars in a specific redshift and magnitude range by multivariate Skew-t distributions. Yang et al. 2017 demonstrated that the Skew-t distribution better describes the color distribution of quasars than Gaussian or Skew-normal distributions. Skewt-QSO also achieves better photo- accuracy compared to other quasar photo- algorithms, such as the empirical color–redshift relation (Richards et al. 2001; Weinstein et al. 2004, e.g.,) and the XDQSOz algorithm (Bovy et al. 2012). Here we briefly describe the Skewt-QSO algorithm (see more details in Yang et al. 2017).
The probability density function (PDF) of a multivariate Skew-t distribution, denoted by , can be expressed as (Lachos et al. 2014),
where is the -dimensional variate (relative fluxes), is the mean vector, is the covariance matrix, is the degree of freedom, is the shape parameter, and is the Mahalanobis distance . and denote the PDF and cumulative distribution function (CDF) of the Student-t distribution,
where is the gamma function. When and , the Skew-t distribution becomes the normal distribution, .
As redshift increases, the Ly emission begins to drop out and the Lyman- forest begins to move into blue bands. We use the relative fluxes instead of colors because at the faint end even negative flux (e.g., non-detection) is useful. Yang et al. 2017 used band as the reference band. However, the Ly emission of quasars begins to drop out of band. Using a fixed reference band for relative fluxes will lead to large uncertainties for high-redshift quasars. Here we use a flexible reference band to compute relative fluxes. We choose the reference band as the band with the maximum signal-to-noise ratio (S/N) in DES photometry.
The likelihood function of the multivariate attribute, , for a given can be described by the multivariate Skew-t distribution. Here, is the multi-dimensional relative fluxes. For quasars and galaxies, is the redshift, ; for stars, is the color, .
To model the colors of quasars (construct the likelihood functions), we divide the quasar training sample described in §II.2 in redshift bins of and magnitude bins of . This bin size is large enough to enclose enough quasars in one bin and small enough for quasar photo- estimation (Yang et al. 2017). To model the colors of galaxies, we divide the galaxy training sample in redshift bins of and magnitude bins of . We divide the star training sample into color bins, where the color (wrt the reference band) can be treated as a parameter similar to redshift for quasars/galaxies. We use different colors (specifically , , , , or ) when the reference band is different (, , , , or ). We divide the star training sample into color bins of and magnitude bins of . Thus we obtain a series of Skew-t parameters as a function of redshift (or color) and magnitude for quasars, galaxies, and stars, respectively. Using the multivariate Skew-t distributions with these parameters, we obtain the likelihood functions to describe quasars, galaxies, and stars in the multi-dimensional color space as a function of redshift (or color), in different magnitude bins.
For an object, at a given magnitude, with the multivariate attribute (multi-dimensional relative fluxes) and their uncertainties, we using Equation (2) to estimate the likelihood in each quasar redshift bin, , the likelihood in each galaxy redshift bin, , and the likelihood in each star color bin, .
III.4 Joint Posterior Probability
For an object with available photometric data in multiple bands (thus we know its magnitude and multi-dimensional relative fluxes), we obtain the joint posterior probability (Equation 1) by combining the prior probability described in §III.2 and the likelihood function described in §III.3 for quasar, galaxy, and star classes, respectively.
For the quasar class, we obtain the joint posterior probability at each redshift. The quasar class PDF is obtained as
We identify peaks in the PDF automatically using the function in R package https://cran.r-project.org/web/packages/pracma/index.html. We obtain the quasar photo-, denoted as photoz-QSO, from the primary peak with the highest integrated probability within a redshift range , where and denote the locations of zero probability on both sides of the peak as identified by pracma. A parameter describing the probability that the true redshift is located within the primary peak, , can be computed as
which is used to quantify the uncertainty of photoz-QSO.
Similar to the quasar class, the PDF of the galaxy class is
The identified photo- of galaxies is denoted as photoz-Galaxy, and the probability that the true redshift is located within is
The total probabilities of quasar, galaxy, and star are
Therefore, the normalized probability of an object being a quasar is expressed as
The quasar candidate selection flowchart is shown in Figure 2. For a given object with relative fluxes and magnitudes, we calculate the posterior probability of the object being a quasar, a galaxy, or a star combining their likelihood and prior probabilities. We compare these probabilities, and classify the candidate as a quasar, a galaxy, or a star when , , or is the maximum probability, respectively. By construction, these three probabilities are normalized to have a unity sum, i.e., . We also obtain photoz-QSO for quasar candidates and photoz-Galaxy for galaxy candidates.
III.5 Proper Motions and Parallaxes
It has been shown that proper motion and parallax detections from Gaia can help reduce stellar contamination in photometric quasar selection (Shu et al. 2019; Calderone et al. 2019; Wolf et al. 2020, e.g., ). We define the parallax significance, , as
where parallax_error is the measurement uncertainty of parallax.
Following Hambly et al. 2008, we define the proper motion significance, , as
where pmra (pmra_error) is the proper motion (measurement error) in right ascension, and pmdec (pmdec_error) is the proper motion (measurement error) in declination This PMSIG definition neglects correlated errors in pmra and pmdec.. We use PLXSIG and PMSIG as additional criteria in our quasar target selection to exclude stars.
IV Results
Table 3 summarizes the photo- regression and classification results for spectroscopically confirmed objects (quasars, galaxies, and stars) for different photometric band combinations. In total, we used 15 most frequent photometric band combinations. In general, the photo- regression and classification results are better when more bands are used, as expected.
The difference between the photo- () and the spectroscopic redshift () is quantified as . The photo- accuracy is the fraction of objects in a test sample with smaller than 0.1. A larger represents a higher photo- accuracy. In addition, a smaller standard deviation of measured for the test sample, , would indicate that the photo- result is better overall. When using bands from DES, combined with all available IR bands ( in NIR and in MIR), the photo- accuracy is as good as 92.2% for quasars and 98.1% for galaxies for our spectroscopic training samples. The standard deviation of , , is 0.147 and 0.035 for quasars and galaxies, respectively.
As shown in Table 3, with fewer bands, decreases and increases for both quasars and galaxies, as expected. When only using DES bands, for quasars and 90.0% for galaxies; and 0.085 for quasars and galaxies, respectively. For comparison, and for quasars when only using SDSS bands (Yang et al. 2017). The photoz-QSO accuracy using DES photometry is slightly worse than that using SDSS photometry because there is no band in DES, which is useful for quasar photo- calculation at low redshift. As shown in Yang et al. 2017, is 72.8% and is 0.31 using the XDQSOz method (Bovy et al. 2012). Using our algorithm and the DES photometric data, the photo- accuracy is comparable or slightly better than the XDQSOz algorithm using SDSS photometry.
Figure 3 shows the photo- versus the spectroscopic redshift with the fewest bands (DES only, upper panels) and most bands (DES_YJHK_W1W2, bottom panels) for quasars (left panels) and galaxies (right panels). The color-map shows the number density. For quasars, using DES data alone, there is an apparent degeneracy between and , since the DES colors of quasars at both redshifts are similar as the C III] (Mg II) line shifts into the band at (). This degeneracy is resolved with the inclusion of NIR and MIR data. For galaxies, there is some degeneracy at and this problem is also alleviated by including NIR and MIR data.
Our algorithm not only calculates quasar and galaxy photo-, but also classifies quasars, galaxies, and stars based on the maximum probability. In Table 3, we show the fraction of objects classified as quasars, galaxies and stars in the spectroscopic training samples. We used the , , and parameters for the classification. As illustrated in Figure 2, a target is classified as a quasar when its is higher than and (thus the normalized probability in Equation 10 is larger than ). We successfully classify 94.7% of quasars, 99.3% of galaxies, and 96.3% of stars when all bands are available. Fewer quasars are mis-classified as stars when including MIR photometry since stars usually radiate thermal emission and are faint in the MIR. At the faint end, without NIR and/or MIR photometry, more quasars are mis-classified as galaxies due to contamination from their host galaxies.
Figure 4 shows the distribution of -band magnitude for the 102321 spectroscopically confirmed quasars in DES footprint along with our photometric classifications. In this figure, we use all available bands, and 83%, 14%, and 2% of them are classified as quasars, galaxies, and stars, respectively. Quasars at the faint end, especially at , might be misclassified as galaxies.
Stars and galaxies mis-classified as quasars will decrease the purity of the selected quasar candidate sample. Using our benchmark sample of spectroscopically confirmed galaxies and stars, only a small fraction (0.1%-0.3%) of stars are mis-classified as quasars, and a small fraction (0.2%-0.5%) of galaxies are mis-classified as quasars (see Table 3). These contamination rates are based on the most loose quasar selection criteria. Using a higher cut, the contamination from stars and galaxies can be further reduced. Of course, the absolute contamination fraction depends on the densities of stars and galaxies in the targeting field. In §IV.2, we show the full selection criteria, as well as the completeness and efficiency (purity) for our quasar selection for typical high Galactic-latitude fields.
IV.2 Quasar Candidates
We now perform quasar target selection over the 5000 deg2 DES wide-field area. Table 4 summarizes the steps to select quasar candidates. We use the following criteria to optimize the quasar selection:
The maximum S/N in five DES bands is greater than 5, .
At least two DES bands have a S/N greater than 3, .
We request a baseline quality criterion of IMAFLAGS_ISO = 0 in all DES bands.
The Gaia proper motion significance, , and parallax significance, , are smaller than 5.
The Skewt-QSO probability of quasars, , is larger than those of stars, , and galaxies, , i.e., and ; thus by construction.
In total, there are 691,483,608 sources in the DES DR2 coadded photometric catalog. Among these sources, there are 1.47, 645.88, and 44.13 million sources classified as quasars, galaxies, and stars, respectively, using the Skewt-QSO probabilities only (i.e., criterion 5). Using criteria (1)–(5) above, we photometrically classify 1,352,947 as quasar candidates, 334,484,173 as galaxy candidates and 36,950,258 as star candidates (criterion 4 were not applied to star candidates). Figure 5 shows the -band magnitude and redshift distributions of the 1.35 million quasar candidates. In both panels, the y axes are in logarithmic scale. The left panel shows that the targets are highly complete at with log-linear number increasing from bright to faint magnitude. In the right panel, the redshift distribution peaks around 1.5. The QLF studies show that the number density of luminous quasars peaks between redshifts 2 and 3 (Richards et al. 2006, e.g., ). For quasars with the same absolute magnitude, the apparent magnitude becomes fainter as redshift increases, and in the fainter regime the selection completeness decreases, so the redshift peak moves towards lower redshift. The blue histogram in the right panel shows the quasar candidates with . With higher completeness at , the redshift distribution between 0.5 and 2.2 becomes flatter and the peak moves slightly to higher redshift. There are few quasar candidates with as the Lyman- emission line drops out of band at .
Among the set of selection criteria, the first criterion excludes 43.4% of DES sources. The second, third, and forth criteria further exclude 4.8% of DES sources. The most crucial criterion is the fifth criterion from the Skewt-QSO probability, excluding 51.6% of DES sources. Quasars are normally point-like sources, but low-redshift quasars and faint quasars can be extended sources. Therefore, we did not perform any morphological cuts based on DES imaging.
For higher selection efficiency (purity), we can adopt higher thresholds. We tested the completeness and efficiency (purity) of quasar selection in the Stripe 82 (S82) region of SDSS, where the spectroscopic completeness of photometric objects is relatively high. Specifically, we use the S82 region with or and . Since the completeness and efficiency vary with magnitude and decrease dramatically at the faint end, here we use quasars brighter than , which is appropriate for current spectroscopic quasar surveys. Following Yang et al. 2017, the efficiency (purity) is calculated based on the quasar number estimated from the QLF as
where is the number of quasars per deg2 calculated from the QLF (Palanque-Delabrouille et al. 2016). Figure 6 shows the completeness and efficiency for different values of using the S82 spectroscopically confirmed quasar sample. As we adopt a higher threshold, the completeness decreases and the efficiency of the selection increases. With our fiducial criteria (1-5), there are very few spectroscopically confirmed quasars (0.4%) and quasar candidates (2%) with . The completeness and efficiency are both high (85%) when using a threshold of 0.7. Therefore, for a high-completeness selection, we recommend to use our fiducial quasar catalog, selected using criteria (1)-(5). For a higher efficiency selection while maintaining a high completeness (), we recommend to add one more criterion of , which results in 0.95 million quasar candidates.
Figure 7 shows the completeness and efficiency as function of -band magnitude. The completeness using one more criterion of is lower than that only using criteria (1)-(5), while the efficiency behaves in the opposite sense. The completeness falls below 80% at . The drop of efficiency (purity) at the bright end is mainly due to enhanced contamination from mis-classified low-redshift bright galaxies. Our algorithm can select some weak-line or narrow-line AGNs. For example, among 5741 narrow-line AGNs in the Million Quasars Catalog (type=“K” or “N”, i.e., narrow-line quasars or narrow-line Seyferts) within the DES footprint, our algorithm selects 492 of them. Therefore, the completeness will increase and the contamination rate will decrease if we include narrow-line objects in our quasar selection. On the other hand, the measurements of the QLF are generally difficult at the bright end given the rapid decrease in the spatial density of quasars toward high luminosities. Therefore our estimated efficiency at the bright end is highly impacted by the quality of the QLF measurement.
Of course, the efficiency of quasar selection also depends on the field stellar density. In sky regions with high stellar densities, the purity would decrease, as more stars would be mis-classified as quasars (even if the fraction of stars mis-classified as quasars is as low as 0.1%).
We also tested applying the most crucial criterion from the Skewt-QSO probability first, resulting in 1.47 million quasar candidates (2.14% of all the source in DES DR2). The other criteria further rule out 0.017% (118,054) sources, demonstrating that is the most useful parameter to rule out contamination. The Gaia astrometry criteria rule out 5110 additional sources. In the bright regime where the Gaia detects proper motion, 4375 out of 536956 sources are rejected at . This confirms that our Skewt-QSO probability criterion selects very few stellar contaminants with large parallaxes/proper motions. Of course, our photometric quasar sample may still contain many faint stars without reliable Gaia DR2 astrometry.
Using probability distributions of parallax/proper motion as a prior probability or machine learning approaches as in Shu et al. 2019 will make better use of Gaia astrometric information. However, as shown in Table 4, the Skewt-QSO color selection has already ruled out the majority of stars, and using the additional parallax/proper motion cuts of PMSIG 5 and PLXSIG 5 only rules out 0.1% additional sources of the 1.35 million quasar candidates after the Skewt-QSO criteria. In that sense, more refined parallax/proper motion cuts are unnecessary, since the primary selection of our quasar candidates is the Skewt-QSO color selection.
In Table 4, we also list the numbers of spectroscopically confirmed quasars in the DES-DR2 source catalog that pass our selection criteria. Criteria (1)-(5) recover 83.1% (90.7%) of all () spectroscopically confirmed quasars. Using , the completeness is 78.0% (87.6%) for all () quasars.
We provide probabilities of quasars, galaxies, and stars for the entire DES DR2 coadded photometric catalog, which contains a total of 691,483,608 sources. The format of our final catalogs is described in Table 5 for the million quasar candidates and in Table 8 for the full DES-DR2 source catalog. These catalogs can be downloaded from URL http://quasar.astro.illinois.edu/paper_data/DES_QSO/.
V Discussion
Using spectroscopically confirmed quasars, we further quantify our quasar selection completeness as a function of both magnitude and redshift. As shown in Yang et al. 2017, the photometric redshift accuracy and the classification success rate of the Skewt-QSO algorithm are high even when using different training and testing samples.
Figure 8 shows the completeness as a function of -band magnitude (left panels) and redshift (right panels) using spectroscopically confirmed quasars. The black diamonds/lines represent the selection results using our fiducial criteria (1)-(5). The blue circles/lines represent the results using one more criterion of . The completeness is higher than 80% at for both selections. The completeness decreases rapidly at , which is the consequence of decreasing photometric accuracy and the lack of infrared detection at the faint end. The right panel in Figure 8 shows the completeness as a function of redshift for quasars. The overall completeness is for quasars over . At the low redshift end (), the selection completeness decreases as redshift decreases, which is due to enhanced contamination from bright host galaxies of quasars at low redshift that causes the quasar not to be selected based on color. At high redshift (), the completeness decreases with redshift. At the completeness estimation suffers from the small spectroscopic sample size, so we use a larger redshift bin of 0.4 at , comparing to a bin of 0.1 at , to avoid large fluctuations. In addition, these completeness estimates are based on spectroscopically confirmed quasars, thus suffering from their own selection effects and incompleteness. So the total completeness might be even lower than our estimation from the spectroscopic sample, especially at the faint end where the original spectroscopic sample suffers the most from incompleteness in selection.
We use simulated quasars to remedy for the small sample size of real quasars at high redshift. The simulation procedure is described in §II.2. We quantify the selection completeness using a sample of 0.8 million simulated quasars, spanning a wide range of redshifts and magnitudes. Figure 9 shows the 2D completeness as a function of redshift (x-axis) and -band magnitude (y-axis) using and , color-coded by selection completeness. The solid cyan line shows the location where the completeness is 80%. Figure 9 confirms that at low redshift (), the completeness decreases with decreasing redshift due to increasing host galaxy contamination. At , the completeness does not decrease with redshift, indicating that the trend observed in Figure 8 is mainly due to the small sample statistics at high redshift. At , the completeness is roughly constant and starts to decrease with increasing magnitude around . At certain redshifts, for example, 1.8, 3.0, 4.8, and 5.5, the 80% selection completeness is achieved at shallower magnitude limit due to contamination from different types of stars, with decreasing effective temperatures.
V.2 Comparison to Gaia QSOC Redshifts
Gaia DR3 release includes redshift estimates for extragalactic sources using low resolution BP/RP spectra (Gaia Collaboration et al. 2022). In Gaia DR3, the Quasi Stellar Object Classifier (QSOC) systematically publish redshift predictions for 1,834,118 sources, with a very low threshold on the Discrete Source Classifier (DSC) quasar probability of and a warning flag of redshift estimation of (Delchambre et al. 2022). There are 47451 spectroscopically confirmed quasars both in our 1.35 million quasar candidate catalog and in the 1.8 million Gaia QSOC redshift sample. Figure 10 shows the comparison between photo-s and spectroscopic redshifts for our algorithm (left panel) and the Gaia QSOC redshifts (right panel) for these 47451 quasars. The photo- accuracy is 93.4% and 61.1% from our algorithm and Gaia QSOC, respectively. Using photo-s from our algorithm, the vast majority is along the 1:1 line. Gaia QSOC redshifts from the low-resolution spectra have smaller scatter along the 1:1 line, indicating smaller redshift uncertainties than our photo-zs (as expected); but there are additional stripes that represent misidentified emission lines in Gaia low-resolution spectra. In particular, 4% of the Gaia QSOC redshifts are incorrectly predicted at , grossly over-predicting the abundance of high-redshift quasars. Using a more stringent cut of with empty warning flags, described by Delchambre et al. 2022, the Gaia QSOC redshift accuracy increases to 95.0%, but the sample is downsized to only 20%. In comparison, using a higher-quality cut in our algorithm of , i.e., the integrated probability of the identified photo- peak is higher than 0.5, our photo- accuracy increases to 94.7% while the sample is only slightly downsized to 95%.
VI Summary
We perform quasar target selection in the southern hemisphere over the deg2 DES-wide survey area. We utilize public DES DR2 optical photometry and available NIR photometric data, from various surveys including VHS, VIKING, UHS, ULAS, and 2MASS. In the MIR, we use the all-sky unWISE photometric data. Our algorithm can efficiently classify sources into the categories quasars, galaxies, and stars, as well as derive photo- for quasar and galaxy candidates.
Our algorithm can successfully classify 94.7% of quasars, 99.3% of galaxies, and 96.3% of stars when all bands are available, benchmarked on spectroscopically confirmed samples. The classification and photo- success rate decrease when fewer bands are available. The quasar (galaxy) photo- accuracy , the fraction of objects with smaller than 0.1, is as high as 92.2% (98.1%) when all bands are available, and decreases to 72.2% (90.0%) when only using 5-band photometry from DES.
We select 1.4 million quasar candidates over the DES wide survey footprint, and provide all classification probabilities to customarily select quasar samples with different completeness and efficiency (purity). Selection completeness and efficiency are anti-correlated. We recommend to use our fiducial criteria (§IV.2) for the most complete quasar sample. We recommend to use one more criterion of for a higher-purity selection and simultaneous high completeness ().
We provide our quasar, galaxy, and star probabilities for all billion sources in the DES DR2 coadd photometric catalog. This catalog will be useful for a broad range of extragalactic and galactic sciences in the southern hemisphere.
Appendix A Quasar Targets for the Black Hole Mapper in SDSS-V
We perform quasar target selection for the Black Hole Mapper (BHM) program in SDSS-V (Kollmeier et al. 2017), in particular, targets in the reverberation mapping (RM) fields. We use the same algorithm but utilize additional optical photometric data. Table 6 summarizes the photometric data we use for the seven initial BHM-RM fields. We use DES photometric data in the XMM-LSS, CDFS, EDFS, and ELAIS-S1 fields. We use the Pan-STARRS1 data (Chambers et al. 2016) for the COSMOS and SDSS-RM fields in the northern sky. For the S-CVZ field we use optical photometric data from Gaia DR2 or the NOAO Source Catalog (Nidever et al. 2018, NSC; ). In addition, we make use of available NIR data in these fields. We use unWISE and photometric data in all fields. Since there are very few high-redshift (-band dropouts at ) quasars in these small fields, we use -band (or Gaia -band) as the reference band for quasar target selection in these BHM-RM fields. The final quasar target catalog for BHM-RM fields is presented in Table 7. We required the Skewt-QSO probability criteria (i.e., and in non-S-CVZ fields, and and in the S-CVZ field). The SDSS-V BHM-RM quasar targets (v0.5) were selected from this catalog with further criteria on , PLXSIG, PMSIG, and magnitude limits on -band (or Gaia -band) magnitude. Specifically, we used a criterion of for SDSS-V BHM-RM quasar targets (v0.5). This criterion is explained in more detail in Yang et al. 2017.
Appendix B A Catalog for All DES DR2 Sources
We publicly release our quasar, galaxy, and star probabilities for all (0.69 billion) photometric sources in the DES DR2 coadded source catalog. We assign photoz and probability parameters as those of quasars when , and as those of galaxies when . The last three columns in Table 8 are the probabilities that only use likelihood, i.e. no prior probabilities were used in Eq. (8). These likelihood parameters are useful for redshift ranges where the luminosity functions may not be well measured, for example, for high-redshift quasars at .