Equivalence of distance-based and RKHS-based statistics in hypothesis testing
Dino Sejdinovic, Bharath Sriperumbudur, Arthur Gretton, Kenji Fukumizu
Introduction
The problem of testing statistical hypotheses in high dimensional spaces is particularly challenging, and has been a recent focus of considerable work in both the statistics and the machine learning communities. On the statistical side, two-sample testing in Euclidean spaces (of whether two independent samples are from the same distribution, or from different distributions) can be accomplished using a so-called energy distance as a statistic [Székely and Rizzo (2004, 2005), Baringhaus and Franz (2004)]. Such tests are consistent against all alternatives as long as the random variables have finite first moments. A related dependence measure between vectors of high dimension is the distance covariance [Székely, Rizzo and Bakirov (2007), Székely and Rizzo (2009)], and the resulting test is again consistent for variables with bounded first moment. The distance covariance has had a major impact in the statistics community, with Székely and Rizzo (2009) being accompanied by an editorial introduction and discussion. A particular advantage of energy distance-based statistics is their compact representation in terms of certain expectations of pairwise Euclidean distances, which leads to straightforward empirical estimates. As a follow-up work, Lyons (2013) generalized the notion of distance covariance to metric spaces of negative type (of which Euclidean spaces are a special case).
On the machine learning side, two-sample tests have been formulated based on embeddings of probability distributions into reproducing kernel Hilbert spaces [Gretton et al. (2007, 2012a)], using as the test statistic the difference between these embeddings: this statistic is called the maximum mean discrepancy (MMD). This distance measure was also applied to the problem of testing for independence, with the associated test statistic being the Hilbert–Schmidt independence criterion (HSIC) [Gretton et al. (2005, 2008), Smola et al. (2007), Zhang et al. (2011)]. Both tests are shown to be consistent against all alternatives when a characteristic RKHS is used [Fukumizu et al. (2009), Sriperumbudur et al. (2010)].
Despite their striking similarity, the link between energy distance-based tests and kernel-based tests has been an open question. In the discussion of [Székely and Rizzo (2009), Gretton, Fukumizu and Sriperumbudur (2009), page 1289] first explored this link in the context of independence testing, and found that interpreting the distance-based independence statistic as a kernel statistic is not straightforward, since Bochner’s theorem does not apply to the choice of weight function used in the definition of the distance covariance (we briefly review this argument in Section 5.3). Székely and Rizzo (2009), Rejoinder, page 1303, confirmed that the link between RKHS-based dependence measures and the distance covariance remained to be established, because the weight function is not integrable. Our contribution resolves this question, and shows that RKHS-based dependence measures are precisely the formal extensions of the distance covariance, where the problem of nonintegrability of weight functions is circumvented by using translation-variant kernels, that is, distance-induced kernels, introduced in Section 4.1.
In the case of two-sample testing, we demonstrate that energy distances are in fact maximum mean discrepancies arising from the same family of distance-induced kernels. A number of interesting consequences arise from this insight: first, as the energy distance (and distance covariance) derives from a particular choice of a kernel, we can consider analogous quantities arising from other kernels, and yielding more sensitive tests. Second, in relation to Lyons (2013), we obtain a new family of characteristic kernels arising from general semimetric spaces of negative type, which are quite unlike the characteristic kernels defined via Bochner’s theorem [Sriperumbudur et al. (2010)]. Third, results from [Gretton et al. (2009), Zhang et al. (2011)] may be applied to obtain consistent two-sample and independence tests for the energy distance, without using bootstrap, which perform much better than the upper bound proposed by Székely, Rizzo and Bakirov (2007) as an alternative to the bootstrap.
In addition to the energy distance and maximum mean discrepancy, there are other well-known discrepancy measures between two probability distributions, such as the Kullback–Leibler divergence, Hellinger distance and total variation distance, which belong to the class of -divergences. Another popular family of distance measures on probabilities is the integral probability metric [Müller (1997)], examples of which include the Wasserstein distance, Dudley metric and Fortet–Mourier metric. Sriperumbudur et al. (2012) showed that MMD is an integral probability metric and so is energy distance, owing to the equality (between energy distance and MMD) that we establish in this paper. On the other hand, Sriperumbudur et al. (2012) also showed that MMD (and therefore the energy distance) is not an -divergence, by establishing the total variation distance as the only discrepancy measure that is both an IPM and -divergence.
The equivalence established in this paper has two major implications for practitioners using the energy distance or distance covariance as test statistics. First, it shows that these quantities are members of a much broader class of statistics, and that by choosing an alternative semimetric/kernel to define a statistic from this larger family, one may obtain a more sensitive test than by using distances alone. Second, it shows that the principles of energy distance and distance covariance are readily generalized to random variables that take values in general topological spaces. Indeed, kernel tests are readily applied to structured and non-Euclidean domains, such as text strings, graphs and groups [Fukumizu et al. (2009)].
The structure of the paper is as follows: in Section 2, we introduce semimetrics of negative type, and extend the notions of energy distance and distance covariance to semimetric spaces of negative type. In Section 3, we provide the necessary definitions from RKHS theory and give a review of the maximum mean discrepancy (MMD) and the Hilbert–Schmidt independence criterion (HSIC), the RKHS-based statistics used for two-sample and independence testing, respectively. In Section 4, the correspondence between positive definite kernels and semimetrics of negative type is developed, and it is applied in Section 5 to show the equivalence between a (generalized) energy distance and MMD (Section 5.1), as well as between a (generalized) distance covariance and HSIC (Section 5.2). We give conditions for these quantities to distinguish between probability measures in Section 6, thus obtaining a new family of characteristic kernels. Empirical estimates of these quantities and associated two-sample and independence tests are described in Section 7. Finally, in Section 8, we investigate the performance of the test statistics on a variety of testing problems.
This paper extends the conference publication [Sejdinovic et al. (2012)], and gives a detailed technical discussion and proofs which were omitted in that work.
Distance-based approach
This section reviews the distance-based approach to two-sample and independence testing, in its general form. The generalized energy distance and distance covariance are defined.
We will work with the notion of a semimetric of negative type on a nonempty set , where the “distance” function need not satisfy the triangle inequality. Note that this notion of semimetric is different to that which arises from the seminorm (also called the pseudonorm), where the distance between two distinct points can be zero.
Let be a nonempty set and let be a function such that , {longlist}[1.]
if and only if , and
Then is said to be a semimetric space and is called a semimetric on .
Note that in the terminology of Berg, Christensen and Ressel (1984), satisfying (1) is said to be a negative definite function. The following proposition is derived from Berg, Christensen and Ressel (1984), Corollary 2.10, page 78, and Proposition 3.2, page 82.
If satisfies (1), then so does , for .
is a semimetric of negative type if and only if there exists a Hilbert space and an injective map , such that
2 Energy distance
Unless stated otherwise, we will assume that is any topological space on which Borel measures can be defined. We will denote by the set of all finite signed Borel measures on , and by the set of all Borel probability measures on .
Following Lyons (2013), the notion can be generalized to a metric space of negative type, which we further extend to semimetrics. Before we proceed, we need to first introduce a moment condition w.r.t. a semimetric .
For , we say that has a finite -moment with respect to a semimetric of negative type if there exists , such that . We denote
We are now ready to introduce a general energy distance .
Let be a semimetric space of negative type, and let . The energy distance between and , w.r.t. is
If is a general semimetric, however, a different line of reasoning is needed, and we will come back to this condition in Remark 21, where its sufficiency will become clear using the link between positive definite kernels and negative-type semimetrics established in Section 4.
Note that the energy distance can equivalently be represented in the integral form,
whereby the negative type of implies the nonnegativity of , as discussed by Lyons [(2013), page 10].
3 Distance covariance
Let and be semimetric spaces of negative type, and let and , having joint distribution . The generalized distance covariance of and is
As with the energy distance, the moment conditions ensure that the expectations are finite (which can be seen using the Cauchy–Schwarz inequality). Equivalently, the generalized distance covariance can be represented in integral form,
where is viewed as a function on . Furthermore, Lyons (2013), Theorem 3.20, shows that distance covariance in a metric space characterizes independence [i.e., if and only if and are independent] if the metrics and satisfy an additional property, termed strong negative type. The discussion of this property is relegated to Section 6.
Kernel-based approach
In this section, we introduce concepts and notation required to understand reproducing kernel Hilbert spaces (Section 3.1), and distribution embeddings into RKHS. We then introduce the maximum mean discrepancy (MMD) and Hilbert–Schmidt independence criterion (HSIC).
We begin with the definition of a reproducing kernel Hilbert space (RKHS).
, and
.
If has a reproducing kernel, it is said to be a reproducing kernel Hilbert space (RKHS).
Let be a kernel on , and . The kernel embedding of into the RKHS is such that for all .
Alternatively, the kernel embedding can be defined by the Bochner integral . If a measurable kernel is a bounded function, exists for all . On the other hand, if is not bounded, there will always exist , for which diverges. The kernels we will consider in this paper will be continuous, and hence measurable, but unbounded, so kernel embeddings will not be defined for some finite signed measures. Thus, we need to restrict our attention to a particular class of measures for which kernel embeddings exist (this will be later shown to reflect the condition that random variables considered in distance covariance tests must have finite moments). Let be a measurable kernel on , and denote, for ,
Note that the kernel embedding is well defined , by the Riesz representation theorem.
2 Maximum mean discrepancy
As we have seen, kernel embeddings of Borel probability measures in do exist, and we can introduce the notion of distance between Borel probability measures in this set using the Hilbert space distance between their embeddings.
Let be a kernel on , and let . The maximum mean discrepancy (MMD) between and is given by Gretton et al. (2012a), Lemma 4,
The following alternative representation of the squared MMD [from Gretton et al. (2012a), Lemma 6] will be useful
See Gretton et al. (2012a) and Sriperumbudur et al. (2012) for details.
3 Hilbert–Schmidt independence criterion (HSIC)
The MMD can be employed to measure statistical dependence between random variables [Gretton et al. (2005, 2008), Smola et al. (2007), Gretton and Györfi (2010), Zhang et al. (2011)]. Let and be two nonempty topological spaces and let and be kernels on and , with respective RKHSs and . Then, by applying Steinwart and Christmann [(2008), Lemma 4.6, page 114],
is a kernel on the product space with RKHS isometrically isomorphic to the tensor product .
Let and be random variables on and , respectively, having joint distribution . Furthermore, let be a kernel on , given in (14). The Hilbert–Schmidt independence criterion (HSIC) of and is the MMD between the joint distribution and the product of its marginals .
Following Smola et al. (2007), Section 2.3, we can expand HSIC as
It can be shown that this quantity is equal to the squared Hilbert–Schmidt norm of the covariance operator between RKHSs [Gretton et al. (2005)]. We claim that is well defined as long as and . Indeed, this is a sufficient condition for to exist, since it implies that , which can be seen from the Cauchy–Schwarz inequality,
Furthermore, the embedding of the product of marginals also exists, as it can be identified with the tensor product , where exists since , and exists since .
Correspondence between kernels and semimetrics
In this section, we develop the correspondence of semimetrics of negative type (Section 2.1) to the RKHS theory, that is, to symmetric positive definite kernels. This correspondence will be key to proving the equivalence between the energy distance and MMD, and the equivalence between distance covariance and HSIC in Section 5.
Semimetrics of negative type and symmetric positive definite kernels are closely related, as summarized in the following lemma, adapted from Berg, Christensen and Ressel (1984), Lemma 2.1, page 74.
As a consequence, defined above is a valid kernel on whenever is a semimetric of negative type. For convenience, we will work with such kernels scaled by .
Let be a semimetric of negative type on and let . The kernel
is said to be the distance-induced kernel induced by and centred at .
Indeed, if were given by (16), it would suffice to take , since . By varying the point at the center , we obtain a family
of distance kernels induced by . The following proposition follows readily from the definition of and shows that one can always express (2) from Proposition 3 in terms of the canonical feature map for the RKHS .
Let be a semimetric space of negative type, and . Then: {longlist}[1.]
.
is nondegenerate, that is, the Aronszajn map is injective.
Note that while Lyons [(2013), page 9] also uses the results in Proposition 3 to characterize metrics of negative type using embeddings to general Hilbert spaces, the relation with the theory of reproducing kernel Hilbert spaces is not exploited in his work.
2 Semimetrics generated by kernels
We now further develop the link between semimetrics of negative type and kernels. We start with a simple corollary of Proposition 3.
Let be any nondegenerate kernel on . Then,
defines a valid semimetric of negative type on .
Whenever the kernel and semimetric satisfy (18), we will say that generates . If two kernels generate the same semimetric, we will say that they are equivalent kernels.
The relationship between positive definite kernels and semimetrics of negative type is illustrated in Figure 1.
The requirement that kernels be characteristic (as introduced below Definition 10) is clearly important in hypothesis testing. A second family of kernels, widely used in the machine learning literature, are the universal kernels: universality can be used to guarantee consistency of learning algorithms [Steinwart and Christmann (2008)]. While these two notions are closely related, and in some cases coincide [Sriperumbudur, Fukumizu and Lanckriet (2011)], one can easily construct nonuniversal characteristic kernels as a consequence of Proposition 18. See Appendix B for details.
3 Existence of kernel embedding through a semimetric
In Section 3.1, we have seen that a sufficient condition for the kernel embedding of to exist is that . We will now interpret this condition in terms of the semimetric generated by , by relating to the space of measures with finite -moment w.r.t. .
Let . Suppose . Then we have
where we have used that is a convex function of . From the above it is clear that for .
which implies , thereby satisfying the result for . Suppose the result holds for , that is, for . Let for . Then we have
We are now able to show that is sufficient for the existence of , that is, to show validity of Definition 5 for general semimetrics of negative type . Namely, we let be any kernel that generates , whereby . Thus,
where the first term is finite as , the second term is finite as , and the third term is finite by noticing that and .
Proposition 20 gives a natural interpretation of conditions on probability measures in terms of moments w.r.t. . Namely, the kernel embedding , where kernel generates the semimetric , exists for every with finite half-moment w.r.t. , and thus the MMD, between and is well defined whenever both and have finite half-moments w.r.t. . Furthermore, HSIC between random variables and is well defined whenever their marginals and have finite first moments w.r.t. semimetric and generated by kernels and on their respective domains and .
Main results
In this section, we establish the equivalence between the distance-based approach and the RKHS-based approach to two-sample and independence testing from Sections 2 and 3, respectively.
We show that for every , the energy distance is related to the MMD associated to a kernel that generates .
Let be a semimetric space of negative type and let be any kernel that generates . Then
In particular, equivalent kernels have the same maximum mean discrepancy.
Since generates , we can write . Denote . Then
where we used the fact that . This result may be compared with that of Lyons [(2013), page 11, equation (3.9)] for embeddings into general Hilbert spaces, where we have provided the link to RKHS-based statistics (and MMD in particular). Theorem 22 shows that all kernels that generate the same semimetric on give rise to the same metric on (possibly a subset of) , whence is merely an extension of the metric induced by on point masses, since
In other words, whenever kernel generates , is an isometry between and , endowed with the MMD metric ; and the Aronszajn map is an isometric embedding of a metric space into . These isometries are depicted in Figure 2. For simplicity, we show the case of a bounded kernel, where kernel embeddings are well defined for all , in which case and endowed with the Hilbert-space metric inherited from are also isometric (note that this implies that the subsets of RKHSs corresponding to equivalent kernels are also isometric).
2 Equivalence between HSIC and distance covariance
We now show that distance covariance is an instance of the Hilbert–Schmidt independence criterion.
Let and be semimetric spaces of negative type, and let and , having joint distribution . Let and be any two kernels on and that generate and , respectively, and denote
Then, .
where we used that , and that when does not depend on one or more of its arguments, since also has zero marginal measures. Convergence of integrals of the form is ensured by the moment conditions on the marginals. We remark that a similar result to Theorem 24 is given by Lyons [(2013), Proposition 3.16], but without making use of the link with kernel embeddings. Theorem 24 is a more general statement, in the sense that we allow to be a semimetric of negative type, rather than metric. In addition, the kernel interpretation leads to a significantly simpler proof: the result is an immediate application of the HSIC expansion in (3.3).
As in Remark 23, to ensure the existence of the distance covariance, we impose a stronger condition on the marginals: and , while and are sufficient for the existence of the Hilbert–Schmidt independence criterion.
By combining the Theorems 22 and 24, we can establish the direct relation between energy distance and distance covariance, as discussed in Remark 7.
As introduced by Székely, Rizzo and Bakirov (2007), the notion of distance covariance extends naturally to that of distance variance and of distance correlation (by analogy with the Pearson product-moment correlation coefficient),
The distance correlation can also be expressed in terms of associated kernels—see Appendix A for details.
3 Characteristic function interpretation
for a particular choice of weight function given by
for a finite nonnegative Borel measure . It follows [Gretton, Fukumizu and Sriperumbudur (2009)] that
which is in clear correspondence with (21). The weight function in (22) is not integrable, however, so we cannot find a continuous translation invariant kernel for which coincides with the distance covariance. Indeed, the kernel in (20) is not translation invariant.
Distinguishing probability distributions
Theorem 3.20 of Lyons (2013) shows that distance covariance in a metric space characterizes independence if the metrics satisfy an additional property, termed strong negative type. We review this notion and establish the interpretation of strong negative type in terms of RKHS kernel properties.
The semimetric space , where is generated by kernel , is said to have a strong negative type if ,
Since the quantity in (23) is, by equation (6), exactly , , the following is immediate:
Let kernel generate . Then has a strong negative type if and only if is characteristic to .
Thus, the problem of checking whether a semimetric is of strong negative type is equivalent to checking whether its associated kernel is characteristic to an appropriate space of Borel probability measures. This conclusion has some overlap with [Lyons (2013)]: in particular, Proposition 29 is stated in [Lyons (2013), Proposition 3.10], where the barycenter map is a kernel embedding in our terminology, although Lyons does not consider distribution embeddings in an RKHS.
From Lyons (2013), Theorem 3.25, every separable Hilbert space is of strong negative type, so a distance kernel induced by the (inner product) metric on is characteristic to the appropriate space of probability measures.
Consider the kernel in (20), and assume for simplicity that and are bounded, so that we can consider embeddings of all probability measures. It turns out that need not be characteristic—that is, it may not be able to distinguish between any two distributions on , even if and are characteristic. Namely, if is the distance kernel induced by and centred at , then for all . That means that for every two distinct , we have . Thus, given that and have strong negative type, the kernel in (20) characterizes independence, but not equality of probability measures on the product space. Informally speaking, distinguishing from is an easier problem than two-sample testing on the product space.
Empirical estimates and hypothesis tests
In this section, we outline the construction of tests based on the empirical counterparts of MMD/energy distance and HSIC/distance covariance.
So far, we have seen that the population expression of the MMD between and is well defined as long as and lie in the space , or, equivalently, have a finite half-moment w.r.t. semimetric generated by . However, this assumption will not suffice to establish a meaningful hypothesis test using empirical estimates of the MMD. We will require a stronger condition, that (which is the same condition under which the energy distance is well defined). Note that, under this condition we also have , as .
Given i.i.d. samples and , the empirical (biased) -statistic estimate of (3.2) is given by
Recall that if generates , this estimate involves only the pairwise -distances between the sample points.
2 Independence testing
In the case of independence testing, we are given i.i.d. samples , and the resulting -statistic estimate (HSIC) is [Gretton et al. (2005, 2008)]
Let be an i.i.d. sample from , with values in , s.t. and . Then
3 Test designs
Experiments
In this section, we assess the numerical performance of the distance-based and RKHS-based test statistics with some standard distance/kernel choices on a series of synthetic data examples.
In the two-sample experiments, we investigate three different kinds of synthetic data. In the first, we compare two multivariate Gaussians, where the means differ in one dimension only, and all variances are equal. In the second, we again compare two multivariate Gaussians, but this time with identical means in all dimensions, and variance that differs in a single dimension. In our third experiment, we use the benchmark data of Sriperumbudur et al. (2009): one distribution is a univariate Gaussian, and the second is a univariate Gaussian with a sinusoidal perturbation of increasing frequency (where higher frequencies correspond to harder problems). All tests use a distance kernel induced by the Euclidean distance. As shown on the left-hand plots in Figure 3, the spectral and bootstrap test designs appear indistinguishable, and significantly outperform the test designed using the quadratic form bound, which appears to be far too conservative for the data sets considered. The average Type I errors are listed in Table 1, and are close to the desired test size of for the spectral and bootstrap tests.
We also compare the performance to that of the Gaussian kernel, commonly used in machine learning, with the bandwidth set to the median distance between points in the aggregation of samples. We see that when the means differ, both tests perform similarly. When the variances differ, it is clear that the Gaussian kernel has a major advantage over the distance-induced kernel, although this advantage decreases with increasing dimension (where both perform poorly). In the case of a sinusoidal perturbation, the performance is again very similar.
In addition, following Example 15, we investigate performance of kernels obtained using the semimetric for . Results are presented in the right-hand plots of Figure 3. In the case of sinusoidal perturbation, we observe a dramatic improvement compared with the case and the Gaussian kernel: values (and smaller) offer virtually error-free performance even at high frequencies [note that yields the energy distance described in Székely and Rizzo (2004, 2005)]. Small improvements over a wider range are also observed in the cases of differing mean and variance.
We observe from the simulation results that distance-induced kernels with higher exponents are advantageous in cases where distributions differ in mean value along a single dimension (with noise in the remainder), whereas distance kernels with smaller exponents are more sensitive to differences in distributions at finer lengthscales (i.e., where the characteristic functions of the distributions differ at higher frequencies).
2 Independence experiments
To assess independence tests, we used an artificial benchmark proposed by Gretton et al. (2008): we generated univariate random variables from the Independent Component Analysis (ICA) benchmark densities of Bach and Jordan (2002); rotated them in the product space by an angle between and to introduce dependence; filled additional dimensions with independent Gaussian noise; and, finally, passed the resulting multivariate data through random and independent orthogonal transformations. The resulting random variables and were dependent but uncorrelated. The case (sample size) and (dimension) is plotted in Figure 4 (left). As observed by Gretton, Fukumizu and Sriperumbudur (2009), the Gaussian kernel using the median inter-point distance as bandwidth does better than the distance-induced kernel with . By varying , however, we are able to obtain a wide performance range: in particular, the values (and smaller) have an advantage over the Gaussian kernel on this dataset. As for the two-sample case, bootstrap and spectral tests have indistinguishable performance, and are significantly more sensitive than the quadratic form-based test, which failed to detect any dependence on this dataset.
Conclusion
We have established an equivalence between the generalized notions of energy distance and distance covariance, computed with respect to semimetrics of negative type, and distances between embeddings of probability measures into certain reproducing kernel Hilbert spaces. As a consequence, we can view energy distance and distance covariance as members of a much larger class of discrepancy/dependence measures, and we can choose among this larger class to design more powerful tests. For instance, Gretton et al. (2012b) recently proposed a strategy of selecting from a candidate kernels so as to asymptotically optimize the relative efficiency of a two-sample test. Moreover, kernel-based tests can be performed on the data that do not lie in a Euclidean space. This opens the door to new and powerful tools for exploratory data analysis whenever an appropriate domain-specific notion of distance (negative type semimetric) or similarity (kernel) can be defined. Finally, the family of kernels that arises from the energy distance/distance covariance can be employed in many additional kernel-based applications in statistics and machine learning, such as conditional dependence testing and estimating the chi-squared distance [Fukumizu et al. (2008)], Bayesian inference [Fukumizu, Song and Gretton (2011)] and mixture density estimation [Sriperumbudur (2011)].
Appendix A Distance correlation
As described by Székely, Rizzo and Bakirov (2007), the notion of distance covariance extends naturally to that of distance variance and of distance correlation (by analogy with the Pearson product-moment correlation coefficient),
Distance correlation also has a straightforward interpretation in terms of kernels,
Appendix B Link with universal kernels
We briefly remark on how our results on equivalent kernels relate to the notion of universal kernels on compact metric spaces in the sense of Steinwart and Christmann (2008), Definition 4.52:
A continuous kernel on a compact metric space is said to be universal if its RKHS is dense in the space of continuous functions on , endowed with the uniform norm.
The family of universal kernels includes the most popular choices in machine learning literature, including the Gaussian and the Laplacian kernel. The following characterization of universal kernels is due to Sriperumbudur, Fukumizu and Lanckriet (2011):
Let be a continuous kernel on a compact metric space . Then, is universal if and only if is a vector space monomorphism, that is,
Acknowledgments
D. Sejdinovic, B. Sriperumbudur and A. Gretton acknowledge support of the Gatsby Charitable Foundation. The work was carried out when B. Sriperumbudur was with Gatsby Unit, University College London. B. Sriperumbudur and A. Gretton contributed equally.