The local relaxation flow approach to universality of the local statistics for random matrices
Laszlo Erdos, Benjamin Schlein, Horng-Tzer Yau, Jun Yin
Introduction
A central question concerning random matrices is the universality conjecture which states that local statistics of eigenvalues of large square matrices are determined by the symmetry type of the ensembles but are otherwise independent of the details of the distributions. In particular they coincide with that of the corresponding Gaussian ensemble. The most commonly studied ensembles are
(i) hermitian, symmetric and quaternion self-dual matrices with identically distributed and centered entries that are independent (subject to the natural restriction of the symmetry);
(ii) sample covariance matrices of the form , where is an matrix with centered real or complex i.i.d. entries.
There are two types of universalities: the edge universality and the bulk universality concerning energy levels near the spectral edges and in the interior of the spectrum, respectively. Since the works of Sinai and Soshnikov , the edge universality is commonly approached via the fairly robust moment method ; very recently an alternative approach was given in .
The bulk universality is a subtler problem. In the simplest case of the hermitian Wigner ensemble, it states that, independent of the distribution of the entries, the local -point correlation functions of the eigenvalues (see (2.3) for the precise definition later), after appropriate rescaling and in the limit, are given by the determinant of the sine kernel
Similar statement is expected to hold for all other ensembles mentioned above but the explicit formulas are somewhat more complicated. Detailed formulas for the different Wigner ensembles can be found e.g., in . The various sample covariance ensembles have the same local statistics for their singular values as the local eigenvalue statistics of the corresponding Wigner ensembles.
For ensembles of hermitian, symmetric or quaternion self-dual matrices that remain invariant under the transformations for any unitary, orthogonal or symplectic matrix , respectively, the joint probability density function of all the eigenvalues can be explicitly computed. These ensembles are typically given by the probability density
where is a real function with sufficient growth at infinity and is the flat Lebesgue measure on the corresponding symmetry class of matrices. The eigenvalues are strongly correlated and they are distributed according to a Gibbs measure with a long range logarithmic interaction potential. The joint probability density of the eigenvalues of with distribution (1.2) can be computed explicitly:
where for hermitian, symmetric and symplectic ensembles, respectively, and const. is a normalization factor. The formula (1.3) defines a joint probability density of real random variables for any even when there is no underlying matrix ensemble. This ensemble is called the invariant -ensemble. Quadratic corresponds to the Gaussian ensembles; we note that these are the only ensembles that are simultaneously invariant and have i.i.d. matrix entries. These are called the Gaussian Orthogonal, Unitary and Symplectic Ensembles (GOE, GUE, GSE for short) in case of , respectively. Somewhat different choices of lead to two other classical ensembles, the Laguerre and the Jacobi ensembles, that also have matrix interpretation for (e.g., the Laguerre ensemble corresponds to the Gaussian sample covariance matrices which are also called Wishart matrices), see for more details. The local statistics can be obtained via a detailed analysis of orthogonal polynomials on the real line with respect to the weight function . This approach was originally applied to classical ensembles by Dyson , Mehta and Gaudin and Mehta that lead to classical orthogonal polynomials. Later general methods using orthogonal polynomials were developed to tackle a very general class of invariant ensembles by Deift et.al., see and references therein, and also by Bleher and Its and Pastur and Schcherbina .
Many natural matrix ensembles are typically not unitarily invariant; the most prominent examples are the Wigner matrices or the sample covariance matrices mentioned in (i) and (ii). For these ensembles, apart from the identically distributed Gaussian case, no explicit formula is available for the joint eigenvalue distribution. Thus the basic algebraic connection between eigenvalue ensembles and orthogonal polynomials is missing and completely new methods needed to be developed.
The bulk universality for hermitian Wigner ensembles has been established recently in , by Tao and Vu in and in . These works rely on the Wigner matrices with Gaussian divisible distribution, i.e., ensembles of the form
where is a Wigner matrix, is an independent standard GUE matrix and is a positive constant. Johansson (see also Ben Arous and Péché and the recent paper ) proved the bulk universality for the eigenvalues of such matrices by an asymptotic analysis on an explicit formula for the correlation functions adapted from Brézin-Hikami . Unfortunately, the similar formula for symmetric or quaternion self-dual Wigner matrices, as well as for real sample covariance matrices, is not very explicit and the technique of cannot be extended to prove universality. Complex sample covariance matrices can however be handled with an analogous formula and universality without any Gaussian component is a work in progress .
A key observation of Dyson is that if the matrix is embedded into a stochastic matrix flow, i.e. one considers where the matrix elements of are independent standard Brownian motions with variance , then the evolution of the eigenvalues is given by a system of coupled stochastic differential equations (SDE), commonly called the Dyson Brownian motion (DBM) . If we replace the Brownian motions by the Ornstein-Uhlenbeck processes to keep the variance constant, then the resulting dynamics on the eigenvalues, which we still call DBM, has the GUE eigenvalue distribution as the invariant measure. Similar stochastic processes can be constructed for symmetric, quaternion self-dual and sample covariance type matrices, and, in fact, on the level of eigenvalue SDE they can be extended to other values of (see (5.5) and (5.8) for the precise formulas).
The result of can be interpreted as stating that the local statistics of GUE is reached via DBM for time of order one. In fact, by analyzing the dynamics of DBM with ideas from the hydrodynamical limit, we have extended Johansson’s result to . The key observation of is that the local statistics of eigenvalues depend exclusively on the approach to local equilibrium which in general is faster than reaching the global equilibrium. Unfortunately, the identification of local equilibria in still uses explicit representations of correlation functions by orthogonal polynomials (following e.g. ), and the extension to other ensembles is not a simple task.
In we introduced an approach based on a new stochastic flow, the local relaxation flow, which locally behaves like DBM, but has a faster decay to equilibrium. This method completely circumvented explicit formulas and it resulted in proving universality for symmetric Wigner matrices (the method applies to hermitian and quaternion self-dual Wigner matrices as well). As an input of this method, we needed a fairly detailed control on the local density of eigenvalues that could be obtained from our previous works on Wigner matrices .
In this paper we will prove a general theorem which states that as long as the eigenvalues are at most distance near their classical location on average, the local statistics is universal and in particular it coincides with the Gaussian case for which explicit formulas have been computed. To introduce this flow, denote by the location of the -th eigenvalue that will be defined in (2.12). We first define the pseudo equilibrium measure by
where is the probability measure for the eigenvalue distribution of the corresponding Gaussian ensemble. In case of Wigner matrices, is the measure for the general ensemble ( and for GUE):
In this setting, it is natural to view eigenvalues as random points and their equilibrium measure as Gibbs measure with a Hamiltonian . We will freely use the terminology of statistical mechanics. Note that the additional term in confines the -th point near its classical location, but the probability w.r.t. the equilibrium measure of the event that near its classical location will be shown to be very close to . Furthermore, we will prove that the local statistics of the measures and are identical in the limit and this justifies the term pseudo equilibrium measure.
The local relaxation flow is defined to be the reversible flow (or the gradient flow) generated by the pseudo-equilibrium measure. The main advantage of the local relaxation flow is that it has a faster decay to global equilibrium (Theorem 4.2) compared with the DBM. The idea behind this construction can be related to the treatment of metastability in statistical physics. Imagine that we have a double well potential and we wish to treat the dynamics of a particle in one of the two wells. Up to a certain time, say , the particle will be confined in the well where the particle initially located. However, the potential of this particle, given by the double well, is not convex. A naive idea is to regain the convexity before the time is to modify the potential to be a single well! Now as long as we can prove that the particle was confined in the initial well up to , there is no difference between these two dynamics. But the modified dynamics, being w.r.t. a convex potential, can be estimated much more precisely and this estimate can be carried over to the original dynamics up to the time .
In our case, the convexity of the equilibrium measure is rather weak and in fact, it comes from the quadratic confining potential of (1.6). So the potential is convex, just not “convex enough”. There is no sharp transition like jumping from one metastable state to another as in the double well case. Instead, there are two time scales: in short time the local equilibrium is formed, on longer time, it approaches the global equilibrium. The approach to the local equilibrium is governed by a strong intrinsic convexity in certain directions due to the interactions (see (2.10) later for a precise formula). To reveal this additional convexity, in our previous paper we introduced a pseudo equilibrium measure where we replaced the long range part of the interaction by a mean-field potential term using the classical locations of far away particles. This potential term inherited the intrinsic convexity of the interaction and it could be directly used to enhance the decay to the local statistics. One technical difficulty with this approach was that we needed to handle the singular behavior of the logarithmic interaction potential. In this paper we show that the pseudo equilibrium measure can be defined by adding a Gaussian term. This simple modification turns out to be sufficient and is also model-independent. Since the Gaussian modification is regular, we no longer need to deal with singularities. The price to pay is that we need a slightly stronger local semicircle law which will be treated in Section 8.
The method of local relaxation flow itself proves universality for Wigner matrices with a small Gaussian component (typically of variance with some ). In other words, we can prove universality for a Wigner ensemble whose single entry distribution (the distribution of its matrix elements) is given by , where is the generator of the Ornstein-Uhlenbeck process and is any initial distribution (We remark that in our approach of decay to equilibrium, the Brownian motion in the construction of DBM is always replaced by the Ornstein-Uhlenbeck process). To obtain universality for Wigner matrices without any Gaussian component, it remains to prove that for a given Wigner matrix ensemble with a single entry distribution we can find and such that the eigenvalue distributions of the ensembles given by and are very close to each other. By the method of reverse heat flow introduced in , we choose to be an approximation of . Although the Ornstein-Uhlenbeck evolution cannot be reversed, we can approximately reverse it provided that is sufficient smooth and the time is short. This enables us to compare local statistics of Wigner ensembles with and without small Gaussian components assuming that the single entry distribution is sufficiently smooth (see Section 6).
As an application, we will use this method to prove the bulk universality of sample covariance ensembles. The necessary apriori control on the location of eigenvalues will be obtained by a local semicircle law. In addition to sample covariance ensembles, we will outline the modifications needed for proving the bulk universality of symplectic ensembles.
Universality for the local relaxation flow
We consider time dependent permutational symmetric probability measures with density , , with respect to the measure . The dynamics is characterized by the forward equation
with a given permutation symmetric initial data . Here the generator is defined via the Dirichlet form as
Formally, we have . In Appendix A we will show that under general conditions on the generator can be defined as a self-adjoint operator on an appropriate domain and the dynamics is well defined for any initial data. Strictly speaking, we will consider a sequence of Hamiltonians and corresponding dynamics and parametrized by , but the -dependence will be omitted. All results will concern the limit.
With a slight abuse of notations, we will sometimes also use to denote the density of the measure with respect to the Lebesgue measure. The correlation functions of the equilibrium measure are denoted by
We now list our main assumptions on the initial distribution and on its evolution . We first define the subdomain
of ordered sets of points . In the application to the sample covariance matrices, we will use the subdomain
Assumption I. The Hamiltonian of the equilibrium measure has the form
The condition can be relaxed to , see remark after (4.11).
Alternatively, in order to discuss the case of the sample covariance matrices, we will also consider the following modification of Assumption I.
Assumption I’. The Hamiltonian of the equilibrium measure has the form
where and . The function satisfies the same conditions as in Assumption I.
It is easy to check that the condition (2.7) guarantees that the following bound holds for the normalization constant
with some exponent depending on .
It follows from Assumption I (or I’) that the Hessian matrix of satisfies the following bound:
This convexity bound is the key assumption; our method works for a broad class of general Hamiltonians as long as (2.10) holds. In particular, an arbitrary many-body potential function can be added to the Hamiltonians (2.6), (2.8), as long as is convex on the open sets and , respectively. The argument in the proof of the main Theorem 2.1 remains unchanged, but the technical details of the regularization of the singular dynamics (Appendix B) becomes more involved. We do not pursue this direction here since we do not need it for the application for Wigner and sample covariance matrices.
Let denote the location of the -th point under the limiting density, i.e., is defined by
We will call the classical location of the -th point. Note that may not be uniquely defined if the support of is not connected but in this case the next Assumption III will not be satisfied anyway.
Assumption III. There exists an such that
Under Assumption II, the typical spacing between neighboring points is of order away from the spectral edges, i.e., in the vicinity of any energy with . Assumption III guarantees that typically the random points remain in the vicinity of their classical location.
where is the exponent from Assumption III.
The main general theorem is the following:
for where is the exponent from Assumption III.
Suppose in addition to the Assumption I-IV, that there exists an such that, for any
for some constants and only depending on . Then for we have
This theorem shows that the local statistics of the points in the bulk with respect to the time evolved distribution coincides with the local statistics with respect to the equilibrium distribution as long as . In many applications, the local equilibrium statistics can be explicitly computed and in the limit it becomes independent of , in particular this is the case for the classical matrix ensembles (see next section). The restriction on the time will be removed by the reverse heat flow argument (see Section 6) for matrix ensembles.
Since the eigenvalues fluctuate at least on a scale , the best possible exponent in Assumption III is , but we will only be able to prove it for some for the ensembles considered in this paper. Similarly, the optimal exponent in (2.16) is . If we use these optimal estimates, , , and we choose , thus , then we can choose , i.e., we obtain the universality with essentially no averaging in . On the other hand, the error estimate is the strongest, of order , for an averaging on an energy window of size . These errors become weaker if time is reduced. These considerations are not important in this paper, but will be useful when good estimates on and can be obtained.
Convention: Throughout the paper the letters denote positive constants whose values may change from line to line and they are independent of the relevant parameters. Since we will always take the limit at the end, all estimates are understood for sufficiently large .
Universality for Matrix Ensembles
Now we specialize Theorem 2.1 to Wigner and sample covariance matrices with i.i.d. entries. In the next sections we give the precise definitions of these ensembles; formulas for the equilibrium measure and the dynamics will be deferred until Section 5.
In order to apply Theorem 2.1 to Wigner and sample covariance matrices, we need to check that Assumptions I-IV are satisfied for these ensembles. Assumptions I or I’ are satisfied by the definition of the Hamiltonian, the precise formulas are given in Section 5. Assumption II is satisfied since the density of eigenvalues is given by the Wigner semicircle law (3.7) for Wigner matrices . In case of the sample covariance matrices, the singular values of will play the role of ’s and their density is given by the Marchenko-Pastur law (3.14) after an obvious transformation (3.15) . In fact, in Section 8 we prove a local version of the Marchenko-Pastur law in analogy with our previous work on the local semicircle law for Wigner matrices . In Section 9 (Theorem 9.1) we will show that Assumption III is satisfied for these ensembles (more precisely, we will prove that Assumption III is satisfied for sample covariance matrices; the proof for Wigner matrices is analogous, and will not be given in details). Assumption IV will be proved in Lemma 8.1 for the sample covariance matrices, for Wigner matrices the proof was given, e.g., in Theorem 4.6 of . We remark that the assumption that the matrix entries are identically distributed, will only be used in checking Assumptions III and IV. Assumption II holds under much more general conditions on the matrix entries. Finally, the apriori estimate on the entropy follows from the smoothing property of the OU-flow (see Section 5).
To fix the notation, we assume that in the case of real symmetric matrices, the matrix elements of are given by
Finally, for the quaternion self-dual case we assume that is a by complex matrix that can be viewed as an matrix with elements consisting of blocks of the form
The dual of the quaternion is defined to be which corresponds to the hermitian conjugate of the matrix (3.3).
where has a law with zero expectation and variance . The spectrum of is doubly degenerate and we will neglect this degeneracy, i.e., we consider only real (typically distinct) eigenvalues, .
In particular, the typical spacing between neighboring eigenvalues is of order in the bulk of the spectrum.
We will often need to assume that the distributions and have Gaussian decay, i.e., there exists such that
In several statements we can relax this condition to assuming only subexponential decay, i.e., that there exists and such that
For some statements we will need to assume that the measures , satisfy the logarithmic Sobolev inequality, i.e., for any density with it holds that
and a similar bound holds for . We remark that (3.10) implies (3.8), see, e.g. .
2 Sample Covariance Matrix
The real sample covariance matrix ensemble consists of symmetric matrices of the form . Here is an real matrix with fixed and we assume that . The elements of are given by
Moreover, analogously to (3.6), the empirical density of eigenvalues converges weakly in probability to the Marchenko-Pastur law
Most of the analysis will be done for the singular values of that are denoted by . They are supported asymptotically in and therefore the typical spacing between neighboring singular values is of order . Their empirical density converges to
We remark that the assumption that is symmetric is used only at one technical step, namely when we refer to the large deviation result for the extreme eigenvalues of the sample covariance matrices in (see Lemma 9.2 below). The similar result for Wigner matrices has been proven without the symmetry condition, see Theorem 1.4 in .
3 Main Theorems
With the remarks at the beginning of Section 3, Theorem 2.1 applies directly to prove universality for Wigner and sample covariance ensembles with a small Gaussian component; we will not state these theorems separately. To remove the small time restriction from Theorem 2.1, we will apply the reverse heat flow argument. This will give our main result:
Here denotes the probability measure of the eigenvalues of the appropriate Gaussian ensemble, i.e. GUE, GOE, GSE for the case of hermitian, symmetric, and, respectively, quaternion self-dual Wigner matrices; and the ensembles of real or complex sample covariance matrices with Gaussian entries (Wishart ensemble) in case of the covariance matrices . These measures are given in (1.6), with , for Wigner matrices, and, expressed in terms of singular values, in (5.6), with , for sample covariance matrices.
Remark 1.: In the case of symmetric and hermitian Wigner matrices, the condition (3.16) can be removed by applying the Four-moment theorem of Tao and Vu (Theorem 15 of ) as in the proof of Corollary 2.4 of . Similar remark applies to the sample covariance ensembles and to the quaternion self-dual Wigner ensemble provided the corresponding Four-moment theorem is established.
We also remark that a manuscript by Ben-Arous and Péché with a similar statement is in preparation for complex sample covariance matrices that holds for a fixed , i.e., without averaging over the energy parameter in (3.17).
Remark 2.: After the first version of this manuscript was posted on the arxiv, the question that whether the four moment theorem for sample covariance matrices holds was settled in . In particular, gives an alternative proof of the universality of local statistics for the complex sample covariance ensemble when combined with the result of . For the real sample covariance ensemble the universality was established for distributions whose first four moments match the standard Gaussian variable. An important common ingredient to both our approach and that of is the local Marchenko-Pastur law, established in Proposition 8.1; a slightly different version suitable for the application to prove the four moment theorem is proved in .
The four moment theorem in compares the distributions of individual eigenvalues for two different ensembles. For our application to the correlation functions and gap distributions, an alternative approach is to use the recent Green function comparison theorem . This will also remove the smoothness and logarithmic Sobolev inequality restrictions in Theorem 3.1.
We now state our result concerning the eigenvalue gap distribution both for Wigner and sample covariance ensembles. For any and with we define the density of eigenvalue pairs with distance less than in the vicinity of by
Consider an Wigner or sample covariance matrix as in Theorem 3.1 such that the probability measure of the matrix elements satisfies the logarithmic Sobolev inequality (3.10) and, additionally, is symmetric in the sample covariance matrix case. Suppose that the initial density satisfies
where is the probability measure of the eigenvalues of the appropriate Gaussian ensemble, as in Theorem 3.1.
Theorem 3.2 shows that, in particular, the probability to find no eigenvalue in the interval is asymptotically the same as in the corresponding classical Gaussian ensemble. Theorems 3.1 and 3.2 will follow from Theorem 2.1 and the reverse heat flow argument that we present in Section 6. We remark that the additional condition on the symmetry of in the case of sample covariance matrices stems from using a result from on the lowest eigenvalue of these matrices, see Lemma 9.2.
Theorem 3.2 can be proven directly from Theorem 4.1 since the test functions of the form
determine the distribution of the random variable uniquely. Here we take to be the set
Local Relaxation Flow
Suppose that the Hamiltonian given in (2.6) satisfies the convexity bound (2.10) with . Let be the solution of the forward equation (2.1) with an initial density . Fix a positive , set and define
Then for any sufficiently small , there exist constants , depending only on and such that for any and for any , we have
where is the number of the elements in .
The proof of this theorem is similar but much simpler than that of Theorem 2.1 of . The estimate (4.3) improves slightly over the similar estimate in by a factor due to the improvement in (4). Theorem 2.1 will follow from the fact that in case , the assumption (2.13) guarantees that
with the choice and using . More precise error bound will be obtained by relating to . Therefore the local statistics of observables involving eigenvalue differences coincide in the limit. To complete the proof of Theorem 2.1, we will have to show that the convergence of the observables is sufficient to identify the correlation functions of the ’s in the sense prescribed in Theorem 2.1. The details will be given in Section 7.
Proof of Theorem 4.1. Without loss of generality we can assume in the sequel that . To see this, note that any can be approximated by a sequence of bounded functions in -norm with arbitrary precision and the dynamics is a contraction in (see Appendix A), thus and are arbitrarily close in . Since is bounded on the left hand side of (4.3), this is sufficient to pass to the limit .
Every constant in this proof depends on and , and we will not follow the precise dependence. We can assume that . Given , we define
Notice that the choice of depending on which is the main reason that appears in the denominator on the right hand side of (4.3). We now introduce the pseudo equilibrium measure, , defined by
where is chosen such that is a probability measure, in particular with
Note that the additional term confines the -th point near its classical location. We will prove that the probability w.r.t. the equilibrium measure of the event that near its classical location is very close to . Thus there is little difference between the two measures and and in fact, we will prove that their local statistics are identical in the limit . The main advantage of the pseudo equilibrium measure comes from the fact that it has a faster decay to global equilibrium as shown in Theorem 4.2.
Since the additional potential is uniformly convex with
Here we have used in the last estimate. If this assumption is replaced by
for some constant independent of , then there will be an extra term in (4.10). Assuming , we have , then this extra term can be controlled by the term and the same proof will go through. Since for the applications in this paper, the condition is satisfied, we will not use this remark here.
The in the first term comes from the additional convexity of the local interaction and it enhances the “local Dirichlet form dissipation”. In particular we have the uniform lower bound
This guarantees that the relaxation time to equilibrium for the dynamics is bounded above by . We recall the definition of the relative entropy of with respect to any probability measure
The first ingredient to prove Theorem 4.1 is the analysis of the local relaxation flow which satisfies the logarithmic Sobolev inequality and the following dissipation estimate. Its proof follows the standard argument in (used in this context in Section 5.1 of ). In Appendix B we will explain how to extend this argument onto the subdomain . Here we only remark that the key inputs are the convexity bounds (4.10, 4.12) on the Hessian of (4.10).
Suppose (4.10) holds. Consider the forward equation
with an initial condition and with the reversible measure . Assume that . Then we have the following estimates
with a universal constant . Thus the time to equilibrium is of order :
The estimate (4.15) on the second term in (4.10) plays a key role in the next theorem.
Proof. For simplicity, we will consider the case when , the general case easily follows by appropriately redefining the function . Let satisfy
with an initial condition . Thanks to the exponential decay of the entropy on time scale , see (4.17), and the entropy bound on the initial state , the difference between the local statistics w.r.t. and is subexponentially small in ,
giving the second term on the r.h.s. of (4.18). To compare with , by differentiation, we have
Here we used the definition of from (4.7) and note that the factor present in (4.7) cancels the factor from the argument of (4.2). From the Schwarz inequality and , the last term is bounded by
since is smooth and compactly supported. This proves Theorem 4.3.
As a comparison to Theorem 4.3, we state the following result which can be proved in a similar way.
Notice that by exploiting the local Dirichlet form dissipation coming from the second term on the r.h.s. of (4.14), we have gained the crucial factor in the estimate (4.18) compared with (4.20).
The final ingredient to prove Theorem 4.1 is the following entropy and Dirichlet form estimates.
Suppose that (2.10) holds and recall with . Let so that . Assume that with some fixed . Then the entropy and the Dirichlet form satisfy the estimates:
Proof. Recall that . The standard estimate on the entropy of with respect to the invariant measure is obtained by differentiating the entropy twice and using the logarithmic Sobolev inequality. The entropy and the Dirichlet form in (4.21) are, however, computed with respect to the measure . This yields the additional second term in the following identity that holds for any probability density :
where . In our application we set to be time independent, , hence we have
Since is invariant, the middle term on the right hand side vanishes, and from the Schwarz inequality
To obtain the first inequality in (4.21), we integrate (4.24) from to , using that and with some finite , depending on . This apriori bound follows from
where we used (2.9) and (4.6). The second inequality in (4.21) can be obtained from the first one by integrating (4.22) from to and using the monotonicity of the Dirichlet form in time.
Finally, we complete the proof of Theorem 4.1. Recall that and . Choose as density in Theorem 4.3. The condition can be guaranteed by the approximation argument from the beginning of the proof of Theorem 4.1. Then Theorem 4.5, Theorem 4.3 together with (4.25) and the fact that directly imply that
i.e., the local statistics of and can be compared. Clearly, equation (4.26) also holds for the special choice (for which ), i.e., local statistics of and can also be compared. This completes the proof of Theorem 4.1.
Equilibrium measure and Dyson Brownian motion
where is an arbitrary parameter, i.e., this corresponds to choosing in (2.6). With a slight abuse of notations we will use for both the measure and its density with respect to the Lebesgue measure. The specific value correspond to the GUE, GOE and GSE ensembles, respectively.
acting on . The measure is invariant and reversible with respect to the dynamics generated by . Define the Dirichlet form and entropy by
Let denote the probability measure on the set at the time with the given generator . Then satisfies the forward equation
with initial condition . This dynamics is the Dyson Brownian motion.
The Dyson Brownian motion is the corresponding system of stochastic differential equations for the vector that is given by
The treatment of the sample covariance ensembles is fully analogous, but the formulas change slightly. We use the convention in the sample covariance case that denotes the singular values of and are the eigenvalues of . Most of the formulas will be in terms of ’s; in particular we consider the joint distribution function of the singular values. The invariant measure for the singular values is given by (c.f. (5.1)):
where and when is a real matrix, when is a complex matrix. This formula can be obtained by direct calculation (see also Proposition 2.16 of or Fig. 1 of after appropriate rescaling). Define the generator (c.f. (5.2))
Finally, the stochastic differential equation is given by (c.f. (5.5))
In applications to Wigner matrices (), will be the joint probability density of the eigenvalues of the initial hermitian, symmetric or quaternion self-dual Wigner matrix . The limiting density is the Wigner semicircle law given in (3.7). The Dyson Brownian motion describes the eigenvalues of the matrix valued process
with . Here is a symmetric, hermitian or quaternion self-dual matrix-valued process whose offdiagonal elements are standard real, complex or quaternion Brownian motions with variance one and the diagonal elements of are real Brownian motions with variance and , in case , respectively. More precisely, let denote the density function of the distribution of one real component of the -th entry of , (there are two real components for the hermitian matrices and four for the quaternion matrices), then
Let denote the reversible measure for this process. The diagonal elements evolve according to an OU process with twice variance. For any , the solution to (5.9), , has the same distribution as
The generator of the induced stochastic process on the eigenvalues is given by (5.2). The equilibrium measure is the GUE, GOE or GSE eigenvalue distribution. Theorem 2.1 thus says in this case that the local eigenvalue statistics of a Wigner random matrix with a small Gaussian component coincides with the local statistics of the corresponding Gaussian ensemble. The entropy condition on in Theorem 2.1 can be easily obtained by
In the real or complex sample covariance case (), the matrix elements of evolve according to the OU process (5.10), i.e. has the same distribution as
where is an matrix whose elements are i.i.d real or complex Gaussian variables with mean 0 and variance .
Reverse heat flow
To remove the short time restriction from Theorem 2.1 in case of Wigner and sample covariance ensembles and to prove Theorems 3.1 and 3.2, we apply the reverse heat flow argument, presented first in and used also in Corollary 2.4 of .
For fixed or 4, recall the Ornstein-Uhlenbeck process from (5.10) with the reversible Gaussian measure . Let be a positive density with respect to , i.e., and we write . Suppose that for any fixed there are constants depending on such that
and the measure satisfies the subexponential decay condition. We will apply this for the initial distribution , so and differ by a Gaussian factor.
Suppose that satisfies the subexponential decay condition and (6.1) for some . Then there is a small constant depending on such that for there exists a probability density with mean zero and variance such that
for some depending on . Furthermore, can be chosen such that if the logarithmic Sobolev inequality (3.10) holds for the measure , then it holds for as well, with the logarithmic Sobolev constant changing by a factor of at most .
Furthermore, let , with some . Denote by . Then we also have
In order to prove Theorems 3.1 and 3.2, appropriate observables need to be chosen that depend on the matrix elements via the eigenvalues to express the quantities in (3.17) and (3.20). It is easy to see that may grow at most polynomially in . But we can always choose large enough to compensate for it with the choice allowed in Theorem 2.1. Here the verifications of the Assumptions I-IV of Theorem 2.1 were explained at the beginning of Section 3. This completes the proof of our main theorems.
Proof of Proposition 6.1. Define with some small positive depending on , where is a smooth cutoff function satisfying for and for . Set
By assumption (6.1), is positive and
for any if is small enough. To see this, take, e.g., and we have
where we have used , and the assumption (6.1).
Define and by definition, . Then
Since the Ornstein-Uhlenbeck is a contraction in , together with (6.1), we have
for sufficiently small . To estimate the last two terms, we also used that on the support of the measure decays subexponentially in .
Notice that may not be normalized as a probability density w.r.t. but this can be easily adjusted. To compute this normalization, take for example, and we have, by using ,
The last term is bounded by for any due to that has a subexponential decay and using the assumption (6.1) on .
We have proved that there is a constant , for any positive, such that is a probability density. Clearly,
and the same formulas hold if is replaced by since the OU flow preserves expectation and variance. Let be defined by
Then is a probability density w.r.t. with zero mean and variance . It is easy to check that the total variation norm of is smaller than any power of . Using again the contraction property of and (6.5), we get
Now we check the LSI constant for . Recall that was obtained from by translation and dilation. By definition of the LSI constant, the translation does not change it. The dilation changes the constant, but since our dilation constant is nearly one, the change of LSI constant is also nearly one. So we only have to compare the LSI constants between and . From (6.4) and that is nearly one, the LSI constant changes by a factor less than . This proves the claim on the LSI constant.
and this completes the proof of Proposition 6.1.
Proof of Theorem 2.1
We will set if . Our goal is to estimate the difference
Let be an -dependent parameter chosen at the end of the proof, in fact, will be chosen as with some small positive exponent , depending on . Let
and note that . We have the simple bound where
Note that is the same as but with replaced by the constant , i.e., is the equilibrium.
After performing the integration, we will eventually apply Theorem 4.1 to the function
For any and define sets of integers and by
where was defined in (2.12). Clearly . With these notations, we have
The error term , defined by (7.6) indirectly, comes from those indices, for which since unless , the constant depending on the support of . Thus
for any sufficiently large , assuming and using that is a bounded function. The additional factor comes from the integration. Taking the expectation with respect to the measure , we get
using Assumption III (2.13). We can also estimate
where the error term , defined by (7.9), comes from indices such that . It satisfies the same bound (7.8) as .
By the continuity of , the density of ’s is bounded by , thus and . Therefore, summing up the formula (7.5) for , we obtain from (7.6) and (7.9)
for each . A similar lower bound can be proved analogously and we obtain
Adding up (7.12) for all , we get
and the same estimate holds for the equilibrium, i.e., if we set in (7.13). We now subtract the these two formulas and apply (4.3) from Theorem 4.1 to each summand on the second term in (7.13). Choosing to minimize the two error terms involving , we conclude that
where we have used and that .
Together with (7.14), we have thus proved that
Choosing such that and then choose large enough so that the last term is smaller than, say, . We have thus proved that
for and this concludes (2.15).
For the proof of (2.17), we choose , and then by using (2.16) we can estimate directly as
for any , instead of (7.8). Therefore, the estimate on the right hand side of (7.12) and the subsequent estimates can be replaced by
provided . Choosing and following the same proof, we can improve the estimate (7.18) to
for . This proves (2.17) and we have completed the proof of Theorem 2.1.
Local Marchenko-Pastur law
In this section we establish that the empirical density of eigenvalues for sample covariance matrices is close to the Marchenko-Pastur law even on short scale. We do this by controlling the difference of the Stieltjes transform, establishing results analogous to Theorem 4.1. and Proposition 4.2 of . In this section, we focus on , in particular the lower spectral edge . The constants appearing in this subsection may depend on .
Before the detailed proof, we explain the main steps of the argument which is similar to the method we have successively developed in . The proof given here is somewhat complicated by fact that the matrix elements themselves are not independent but are generated as a quadratic expression of independent random variables. The first step, Lemma 8.1, is an apriori bound on the local density on short scales, , using resolvent expansion and a large deviation principle for quadratic forms. Expressing the resolvent of in terms the resolvents of its minors, we obtain a self-consistent equation (8.20) for the Stieltjes transform of the eigenvalues. This equation is very close to the defining quadratic equation of the Stieltjes transform of the Marchenko-Pastur law, see (8.7), with a perturbation term . This term can be estimated by large deviation arguments and using the a-priori bound on the local density. Then in Lemma 8.3 we investigate the stability of the self-consistent equation for . Although the perturbed equation has two solutions, only one of them can be close to . To select the correct solution, we use a continuity argument in the spectral parameter . For with a large imaginary part, say , the explicit formula (8.27) for the solution can be directly analyzed. For approaching to the real axis, we prove that the two unperturbed solutions remain far away from each other (8.35). Since the perturbed solutions are also continuous in the spectral parameter, for a sufficiently small perturbation they must remain in the vicinity of the correct solution of the unperturbed equation.
Let and . Consider the interval . Let denote the number of eigenvalues of in the interval . Suppose that , for some . Then there exist constants such that
for all large enough (independent of ).
We remark that the assumption on can be relaxed to . But we do not need this result here. For details, one can refer to Theorem 5.1 of .
Proof of Lemma 8.1. We observe, first of all, that
where we defined . It follows that
Denoting by the first column of and by the matrix consisting of the last column of , we have
Denote ’s the eigenvalues of the matrix . The ’s are also the eigenvalues of matrix and the other eigenvalues of are zeros. Then define as the normalized eigenvectors of associated with non-zero eigenvalues , i.e., the matrix elements of are given by
where we used the assumption that . Because the eigenvalues ’s are the eigenvalues of the matrix , they are interlaced with the eigenvalues of and . It follows from (8.2) that
where in the last step we used Lemma 4.7 from . The claim follows by the assumption that and that and are large enough.
Consider sample covariance matrices with an matrix with independent and identically distributed complex entries. Let . Recall and in (3.13) and define as
We will often drop the argument from the notation of for brevity. Then for any , satisfying , , the Stieltjes transform,
of the empirical eigenvalue distribution of satisfies
for any small enough (independent of and ) and . Here is the unique solution of
with positive imaginary part for all with .
Recall from (3.13). The function defined in (8.7) depends on and can be written as
where denotes the square root on complex plane whose branch cut is the negative real line. Explicit calculation shows is the Stieltjes transform of the Marchenko-Pastur density given in (3.14).
Using (8.6) and (3.14), we have the local Marchenko-Pastur law for the number of eigenvalues in a small interval:
Consider an interval within the bulk spectrum. Let be a suffiiciently small parameter. Suppose that and are chosen such that with a large constant and with given in (8.5). Then we have the convergence of the counting function, i.e.,
where denotes the number of eigenvalues of in the interval .
Proof of Corollary 8.2. The proof of (8.9) follows from the inequality (8.6) with a similar argument as the proof of Proposition 4.1 of .
We remark that, similarly to Theorem 3.1 in , the assumption on the lower bound can be relaxed to and obtain the local Marchenko-Pastur law on the shortest possible scale, at least away from the spectral edges.
Proof of Proposition 8.1. Let be the -th column of and let be the remaining matrix obtained from after removing the -th column . Let , be the non-zero eigenvalues and the eigenvectors of the matrix and we define . Then we have the formula
Note that the vector is independent of and . Therefore, we have
Define , then
So with and , we have
with some fixed large (using dyadic decomposition and (8.1), similarly to the argument in Lemma 4.2 of ) apart from an event of probability . Then with Proposition 4.5 of , we have
for sufficiently small . Since the eigenvalues of are interlaced with the eigenvalues of , we have
Combining (8.14), (8.15), and (8.16), we find that
for sufficiently small . With the assumption and , we obtain
On the other hand, with the definition of , for any , , such that, , , we have
Together with (8.17), we obtain, for , and sufficiently small ,
where is the line segment connecting points and . Then the Proposition 8.1 follows from the next lemma.
Assume is a positive semidefinite matrix with . For fixed , we recall the notation . Let and , . Denote the line segment connecting and . Suppose that for any , the Stieltjes transform satisfies the following self-consistent relation:
for some ’s. Then there exists depending only on , such that, whenever
with .
Proof of Lemma 8.3. We begin with a special case: . In this case if then . With the assumptions on and , it is easy to see that:
Insert it into (8.20), we obtain when ,
by . Explicit calculation shows
we note (see (8.7)). Then the following lemma implies, with , , , that (8.22) holds for if is small enough.
Let be the solutions of (8.26). Let and . For sufficiently small , depending on ,
Proof of Lemma 8.4. First, when is small enough, an easy calculation shows
Let and . Note that and therefore, by (8.31), . Hence, (8.30) follows from and from the inequality
which holds for any complex number and .
Now we prove (8.22) for the case . We first note that the two solutions of (8.20) are when . One can check that for , these two solutions are bounded by some constant :
Furthermore, can be bounded by for any as follows,
On the other hand, for any , we claim that if is close to or , then it should be really close to or , i.e., if
To see this, note that (8.37) together with (8.34) imply
Then with (8.20) and (8.21), we obtain that (8.25) holds for any . Using Lemma 8.4 again, we have (8.38).
We have seen that (8.22) and (8.37) (for small ) hold when . Because , are continuous functions of , with (8.38) and (8.37), we can see that when is small enough, (8.38) holds for every . This result shows that must be close to at least one of and and it is close to when .
Now we claim that if were close to , i.e.,
then is also close to , which implies that is always close to .
Again, with the continuity of and and (8.38), if is close to in the sense of (8.40), then there exists such that is close to both of and , i.e.,
Together with (8.40), we obtain is still close to the . It means that (8.22) holds for all ’s in our assumption, using the fact . This completes the proof of Lemma 8.3.
The following lemma shows that the expectation value of is close to .
Let , such that , , for some . Then we have
for large enough depending on .
Proof of Lemma 8.5. Using (8.34) and the estimate (8.6) from Proposition 8.1, we have
uniformly in within the range , .
From (8.10), (8.11), (8.13), (8.15) and (8.17), we obtain
Using (8.6), (8.34) and , we obtain that is bounded from below by a constant . Furthermore, for some ,
Combining this with (8.47), (8.48) and (8.49), we obtain:
Let , , and . Suppose for some , we have
when is sufficiently large (depending on ).
Proof of Lemma 8.6. We only prove the case of the real sample covariance matrix. The case of the complex sample covariance matrix can be treated similarly.
Let and be the eigenvalues and eigenvectors of . The derivative of with respect to the -th matrix element is given by
Using and , one can obtain the following result, as in (3.3) of ,
Then with Lemma 8.1, as in (3.6) of , we obtain (8.57).
As in the proof of Lemma 8.5, with a continuity argument and Lemma 8.4, we can obtain (8.56) from (8.62) and complete the proof.
Verifying Assumption III
The following theorem gives the estimate (2.13) for the singular values of the sample covariance matrices.
Assume that the single site distribution of the entries of satisfies the logarithmic Sobolev inequality (3.10). Recall that in (3.14) denotes the density in Marchenko-Pastur law. Define , with the relation
Denote the singular values of . Then there exists , such that,
Remark. An analogous result holds for the eigenvalues of the Wigner matrices; the proof is similar and we will not give the details here. We only point out that the key ingredients of the argument below are: (i) apriori bound on the extreme eigenvalues (see Lemma 9.2 and the remark afterwards); (ii) concentration of the local density of states (used in Lemma 9.3). We also critically use the fact that the density of states (semicircle law or Marchenko-Pastur law) has a square root singularity at the edges.
The real numbers denote the eigenvalues of and the nonnegative real numbers are given by
The proof of (9.3) is a straightforward computation. This identity is the key to extend our results on the local semicircle law for quaternion self-dual matrices without any further modifications.
Now we return to the the proof of Theorem 9.1 and we start with some preparatory lemmas. First, we recall the following result from .
then and for some . Therefore for any ,
Remark. In fact, the error term in is instead of but we will use only the weaker bound (9.6) in order to indicate that our proof goes through for Wigner matrices with not necessarily symmetric distributions as well. In the latter case only has been proven by Vu in for compactly supported distribution with an effective dependence of the constant on the support. This effective dependence is necessary to remove the compact support condition as in Lemma 4.1 of . Strictly speaking, the result in was stated only for symmetric Wigner matrices but it holds for hermitian and quaternion self-dual Wigner matrices as well, since the key estimate (equation (5) in ) is independent of the matrix ensemble.
Recall that in (3.14) denotes the density in Marchenko-Pastur law. Let
Proof of Lemma 9.3. To prove (9.8), with Lemma 9.2, one only needs to prove
To prove (9.9), for fixed , we can assume and denote . Because is an increasing function and the derivative of is bounded by , we have:
Using (9.8), it follows that .
Similarly to the calculation in Theorem 3.1 of , by using the logarithmic Sobolev inequality, we have
hold with an extremely high probability, where we introduced the notations and . Then for some , we have
Proof. By symmetry, we only need to prove that (9.16) holds for the sum on the indices with . Introduce the notation
The estimate (9.13), with , implies that holds with an extremely high probability, for any positive . Therefore, we can bound from above by (for any )
Similarly, we can obtain the lower bound. Putting them together, we have that:
holds for any . The assumption (9.14) implies that
holds with an extremely high probability, for any or . For the other ’s, for which may appear in , we use (9.15) and obtain the following improved bound on : when ,
Let be a continuous and differentiable function, such that , for , for , for or and . Combining (9.21) and (9.19), we obtain
for any , such that, . Therefore we can write
where in the second line we used the fact that the difference between and must be a multiple of . Since holds with an extremely high probability, using (9.22), we can replace with in (9.24), i.e.,
Change the variable from to . With , we obtain and . Thus
where is the inverse function of . Note, when in (9.26), we have . Define the inverse function of as , and for . Then
Inserting this inequality into (9.26) and performing the integration, we can see that
where we expanded (9.28) into four terms:
Since , when and for any , we obtain , for some . Next, from (9.8) and for any , we can see .
To prove , we start with writing as
The first term on the r.h.s. of (9.29) is equal to zero, since is constant outside . The second term can be bounded by , for some , using the facts and , i.e.,
Now we prove that the third and fourth term of (9.29) are less , for some .
From the explicit definition of , an easy calculation shows that, for all ,
Combining this with the fact , we obtain that the third and fourth terms of (9.29) are less than , for some .
To bound the last term of (9.29), we use, once again the bound . From (9.31), we find therefore that
At last, we prove . We rewrite as
where . When , from (9.9) and (9.32), one can see that
Here we also used the fact that for large , decays exponentially fast to zero (see, for example, Lemma 7.3 of , which is stated for matrices with complex entries, but can be trivially extended to the case of real entries). Together with (9.8), we obtain that
When , with (9.32) and (9.8), we can see
Combining (9.35) and (9.38), we obtain for some . Together with (9.28), this compeletes the proof of Lemma 9.5.
Next, we show that the assumptions (9.14) and (9.15) in Lemma 9.5 always hold. First we prove (9.14) in the next Lemma 9.6 with an analogous proof as Lemma 9.5. Then in Lemma 9.7 we show that (9.15) holds when (9.14) holds.
There exist small positive numbers and , such that,
hold with an extremely high probability, where we recall the notations and .
Proof. As in (9.19), for any , and sufficiently large , we have
So without any other assumptions, one can obtain (9.28), if we set instead of defined in the proof of Lemma 9.5. With a similar argument as in the proof of Lemma 9.5 but with this redefined , we have
We prove this claim by contradiction; assume that for some we have . By symmetry we can assume that , the case is analogous. We start with the case . Then and in this case must be larger than , otherwise would contradict to , see (9.6). Using
for any and that is monotone, we obtain that
for any such that . Then
with some positive which would contradict to (9.41). Now we consider the case . The previous argument remains unchanged if . If , then we use
for any such that and we obtain
which again contradicts to (9.41). This completes the proof of (9.42).
On the other hand, the estimate (9.13), with , implies holds with an extremely high probability. Combining (9.42) with this fact, we can see that for any small enough , there exists such that (9.39) holds, which completes the proof of Lemma 9.6.
The next Lemma guarantees the assumption (9.15) in Lemma 9.5, given (9.14).
If there exist sufficiently small positive numbers and , such that
holds with an extremely high probability, then there exists such that,
holds with an extremely high probability, where we recall the notations and .
Proof. For simplicity, we only prove the case of , the case is analogous. Using (9.13), for any , , with , we have
Now we claim that, for , ,
holds with an extremely high probability, which implies
Suppose now that . With the assumption (9.45), we have that, for ,
with an extremely high probability. Divide this interval into small intervals with the length . By the local Marchenko-Pastur law, i.e., Corollary 8.2, the event that the number of the eigenvalues in each piece is larger than holds with an extremely high probability. On the other hand, if , then the total number of eigenvalues in at least one of these intervals is less than , which implies that holds with an extremely high probability. Together with (9.50), we have (9.48). Then combining (9.48), (9.49) and (9.47), we obtain (9.46) and complete the proof.
Proof of Theorem 9.1. Note that the assumptions in Lemma 9.5 are proved in Lemma 9.6 and 9.7. Combining Lemma 9.5, 9.6 and 9.7, we obtain (9.16), i.e.,
holds with an extremely high probability. To see (9.53), first notice that (9.13), with , implies that, for any and
holds with an extremely high probability. The estimate (9.46) shows that there exist and such that
holds with an extremely high probability. Combining this with (9.54) for the remaining indices or , we obtain (9.53). Together with (9.52), we have:
for some . Using the definition , one has
Inserting (9.57) into (9.56), we obtain (9.2) and complete the proof of Theorem 9.1.
Appendix A Existence and restriction of the dynamics
Moreover, by approximating by functions and using that is contraction in (Section 1.5 in ), the differential equation holds even if the initial condition is only in . In this case the convergence , as , holds only in . We remark that is also a contraction on , by duality.
To establish the relation between and , we first define the symmetrized version of
Note that both with are local operators and is symmetric with respect to the permutation of the variables. For any function defined on , we define its symmetric extension onto by . Clearly for any . Since the generator is uniquely determined by its action on its core, and the generator uniquely determines the dynamics, we see that for any , one can determine by computing and restricting it to . In other words, the dynamics (2.1) is well defined when restricted to .
as . It is sufficient to consider only the positive semi-axis, i.e., . Extending the functions to two dimensional radial functions, , , this statement is equivalent to the fact that a point in two dimensions has zero capacity.
Appendix B Bakry-Emery argument on a subdomain
The estimate (4.14) in Theorem 4.2 is based on the Bakry-Emery argument for the dissipation of the Dirichlet form. This method uses a lower bound on the Hessian of and an integration by parts. Since the dynamics is restricted to , we need to check that the boundary term in the integration by parts vanishes.
In our application, this argument will be used for the Hamiltonian (see (4.5)) and its generator , but for simplicity, we omit the tilde from the notation below. With a standard calculation (see (5.8) of with somewhat different notations) shows that
assuming that the quantities in each step are well defined and that the boundary term
in the integration by parts in the third line vanishes. In we argued with a somewhat specific form of , an information not directly available here.
The rigorous proof in the general case uses a regularization and a cutoff argument. First we regularize the function , , by defining
for some . This has the advantage that the derivatives of can be bounded by those of . We consider a cutoff function to be specified later and we insert in the calculation (B.1). Since is an elliptic operator with smooth coefficients away from the boundary , by standard parabolic regularity we know that and thus are smooth functions inside . Thus each step in the cutoff version of (B.1) is justified with an additional term coming from the derivative hitting in the integration by parts. After repeating the steps in (B.1), we obtain
We now show that, by an appropriate choice of a sequence of cutoff functions, the second term in (B.3) vanishes. We first define the set of higher order coalescences where at least three point coincide as
We remark that in case of Assumption I’ we formally introduce to this definition, so that will include also three point singularities of the type . For any we define the set
is the -neighborhood of the three-point singularity set within . Introduce an additional small positive parameter . We now choose the cutoff function of the form , depending on the parameters and , such that
(i) if , if and ;
(ii) if , if and .
We state two estimates on the solution of (4.13) that will be proven at the end of the section.
Assume that . Then the solution of (4.13) satisfies a uniform supremum bound on the closure of ,
Furthermore, is regular away from the higher order coalescence singularities with the estimate
where is a compact set and the constant depends only on the indicated parameters. In particular, is regular up to the boundary , i.e., at the two-point coalescence points away from higher order coalescences.
Using this lemma, we can treat the second term on the r.h.s. of (B.3). We split the integration into two regimes. First we consider the regime where , i.e., an -neighborhood of . On this set we note the local density scales at as , thanks to the term in . Thus the measure of the support of near scales as , while (assuming ). Since (B.5) guarantees that the derivatives of remain locally bounded (with a bound depending on and ), the boundary term near vanishes as .
To estimate the integral on the support of , i.e., on a subset of , we use that can be replaced with after taking the limit. Since we have and with , the integrand scales at most . Since the local density scales at least as due to a factor of the type , the total measure of is of order . Hence the integral on scales at most as in the parameter and therefore the contribution of the neighborhood of higher order singularities to the second term in (B.3) vanishes as .
After having removed and the second term from (B.3), we let and this gives the desired result (4.14).
To complete the argument, finally, we need to prove Lemma B.1.
Proof of Lemma B.1. The bound (B.4) follows immediately, since and the semigroup a contraction in (see Appendix A).
The second statement of Lemma B.1 follows from a standard regularization argument for a typical two-point singularity at that was already outlined in . Fix a point and assume that , but for all other pairs . We remark that the neighborhood of two (or more) independent singularities, e.g., and , , can be treated similarly by applying the same regularization argument separately. We omit these details here.
where is an elliptic operator with second derivatives in the variables and with coefficients regular on the scale (since all other singularities are at least at a distance away from ).
For the case, by introducing a function of variables, we see that satisfies , where
i.e., becomes an elliptic operator with bounded and regular coefficients in the new variables. A similar transformation is possible for any integer , where is considered as the radial part of a -dimensional variable.
We claim that the singular point becomes a removable singularity in the variables around . Note that the singular set is a codimension two subspace in the space-time coordinate system which becomes a line segment in the space-time system if we disregard the variable . Note that plays no role in this argument since every coefficient is regular in . The parabolic equation holds in a strong sense away from the origin in these two variables, and moreover is bounded by (B.4). We can thus apply Theorem II of with , to see that must coincide with the regular solution obtained by using the fundamental solution to the equation in a small space-time neighborhood of the singular set. This proves that , and hence , is a smooth function up to the boundary .
To obtain the quantitative estimate (B.5), we consider the regularity of the coefficients of . Due to the special structure of , every term in is either regular on any small scales, or it scales as . Since the neighborhood is at least at distance away from the other singularities, the coefficients of are regular on scale . Therefore the solution is regular on scale on and this gives the -scaling of the estimate (B.5). This completes the proof of Lemma B.1.
Acknowledgement. The authors thank Alice Guionnet for pointing out some errors in the preliminary version of the manuscript.