Eigenvector localization for random band matrices with power law band width
Jeffrey Schenker
Introduction
Random band matrices, with entries that vanish outside a band of width around the diagonal, have been suggested as a model to study the crossover between a strongly disordered “insulating” regime, with localized eigenfunctions and weak eigenvalue correlations, and a weakly disordered “metallic” regime, with extend eigenfunctions and strong eigenvalue repulsion. Such a crossover is believed to occur in the spectra of certain random partial differential (or difference) operators as the spectral parameter (energy) is changed.
In this paper, the strong disorder side of the band matrix crossover is analyzed. It is shown here that certain ensembles of random matrices whose entries vanish in a band of width around the diagonal satisfy a localization condition in the limit that the size of the matrix tends to infinity provided . This result requires certain assumptions on the distribution of the entries of the matrix, and the proof given here has technical requirements that may not be necessary. Nonetheless, the conditions imposed below (see §3) allow for a large family of interesting examples. In particular, one may consider a Gaussian distributed band matrix, with distribution
with and independent families of i.i.d. real and complex unit Gaussian variables, respectively.
If has distribution (1.1), or more generally a distribution satisfying assumptions 1, 2, and 3 in §3 below, then there exists and such that given and there are and such that
for all and all . For the Gaussian band ensemble (1.1) and .
Remarks: For the Gaussian Band Ensemble (1.1), the density of states, in the regime , , is known to be the Wigner semi-circle law (see §5 below). For outside the support of the semi-circle law, one could obtain (1.3) with using Lifschitz tail type estimates. This will be dealt with in a separate paper.
Theorem 1 estimates the decay of matrix elements of the resolvent away from the diagonal. Using techniques developed in the context of discrete random Schrödinger operators one may obtain from (1.3) estimates on eigenvectors.
Let have distribution (1.1), or more generally a distribution satisfying assumptions 4 and 5 in §5.
With probability one all eigenvalues of are simple.
If (1.3) holds for all in an interval and if , , are the eigenvalues of with corresponding eigenvectors , , then there are , , and
For the proof of this theorem, the reader is directed to the corresponding result in the context of random Schroedinger operators. See for example for the non-degeneracy of the eigenvalues and [2, Theorem A.1] for a derivation of (1.4) from Green’s function decay (1.3). In both cases, the proof involves only averaging over the coupling of a rank one perturbation and can be applied in the present context.
The proof of Theorem 1 is based on two observations, which may be summarized as follows.The idea to study localization via these two complementary estimates was suggested in the context of random Schroedinger operators by Michael Aizenman, and is inspired by the Dobrushin-Shlosman proof of the Mermin-Wagner Theorem on the absence of continuous symmetry breaking in classical statistical mechanics of dimension 2. Let . Then
The random variable is rarely large. This may be expressed through a bound (uniform in ) on the tails of the distribution of
If has distribution (1.1), or more generally a distribution with the properties outlined in §3 below, then there exist and such that
The fluctuations of grow at least linearly with . One would typically express the growth of fluctuations by an inequality like
If has distribution (1.1), or more generally a distribution with the properties outlined in §3 below, then there is such that if and then
with . For the Gaussian band ensemble (1.1) .
Lemmas W and F together easily imply Theorem 1. Indeed, it suffices to show that the second factor on the right hand side of (1.6) is uniformly bounded. But it follows from Lemma W that
This observation, which is the basis of the fractional moment analysis of random Schrödinger operators , follows easily from (1.5) since
It may not be immediately clear what Lemma F has to do with large fluctuations. Towards understanding this, let . By the Hölder inequality,
with unless is non random.
If were Gaussian with variance (and arbitrary mean), then would be proportional to the variance
For a general random variable , the associated quantity may be taken as a measure of the fluctuations of . In place of (1.11), we have the following identity for in terms of the variance of in weighted ensembles:
The identity (1.12) follows from Taylor’s formula with remainder. Indeed, the second derivative of at is equal to the weighted variance . Thus,
Taking a convex combination of these identities, chosen so the first order terms cancel, gives
Thus Lemma F may be understood as giving a lower bound on the fluctuations of , as measured by the improvement to Hölder’s inequality. The proof of this result will be accomplished using a product formula for that expresses this quantity as a matrix element of a product of matrices of size . Prop. (3) will be applied to factors in this product, with each factor contributing a term of size to . Since there are terms, this produces the claimed decay.
The strategy taken below in proving Lemmas W and F has two parts. First we identify certain axioms for the distribution of which lead naturally to the lemmas. Second, we verify that the Gaussian band ensemble (1.1) satisfies these axioms. To motivate the form of the axioms for the distribution of , we begin in §2 with a self contained sketch of the argument in the tri-diagonal case . In §3 we state the assumptions needed to adapt the proof to , state the associated general results and prove Lemma W. In §4 we get to the heart of the matter and prove Lemma F. In §5, we discuss examples of ensembles, including the Gaussian band ensemble (1.1), satisfying the axioms of §3. In an appendix, an elementary probability lemma used below is stated and proved.
2. Remarks on the literature and open problems
In it was observed, based on numerical evidence, that the localization of eigenfunctions and eigenvalue statistics of the Gaussian band ensemble (1.1) are essentially determined by the parameter . When the eigenfunctions are strongly localized and the eigenvalue process is close to a Poisson process. When the eigenfunctions are extended and the eigenvalue statistics are well described by the Gaussian unitary ensemble (GUE). A theoretical physics explanation of these numerical results was given by Fyodorov and Mirlin . They considered a slightly different ensemble in which a full GUE matrix is modified by multiplying each element by a factor which decays exponentially in the distance from the diagonal. For this model, on the basis of super-symmetric functional integrals, they obtain an effective -model approximation which, at the level of saddle point analysis, shows a localization/delocalization transition at .
Theorem 1 is consistent with the above picture. However, suggest that proper exponent on the r.h.s. of (1.3) would be .
What is the optimal value of in (1.3)? In particular, does this equation hold with ?
In the physics literature, the nature of eigenvalue processes in the large limit is generally expected to be related to localization properties of the eigenfunctions, with Poisson statistics corresponding to localized eigenfunctions and Wigner-Dyson statistics corresponding to extended eigenfunctions. Let us call this idea the “statistics/localization diagnostic.” (In the context of band random matrices, a vector is a function on the index set , namely coordinate of . The statistics/localization diagnostic suggests that the eigenvalues of a random matrix should be approximately uncorrelated if a typical eigenvector is essentially supported on a vanishing fraction of , and should show strong correlations if it is typically spread over more or less the entire index set.)
The extreme cases and of the Gaussian band ensemble (1.1) are consistent the statistics/localization diagnostic. Indeed, with , the matrix is diagonal and the eigenvalues, which are just the diagonal entries , are independent. After suitable rescaling the eigenvalue process converges to a Poisson process in the large limit. (This is essentially the definition of a Poisson process.) Likewise the eigenfunctions are the elementary basis vectors , which are localized on single sites. On the other hand, with the matrix is sampled from the GUE. In this case, the eigenfunctions together form a uniformly distributed orthonormal frame, so they are completely extended, and a suitable rescaling of the eigenvalue process converges in distribution to an explicit determinental point process as calculated by Dyson .
Based on the statistics/localization diagnostic, it is reasonable to conjecture that Poisson statistics hold for local fluctuations of the eigenvalues of in a limit with provided . (One must be a little careful with the diagnostic, as it is easy to concoct random matrices with totally extended eigenfunctions and arbitrary statistics: put random numbers with any given joint distribution on the diagonal of a matrix and conjugate the result with a random unitary! Of course, in that ensemble the matrix elements will most likely be highly correlated. Thus, it remains plausible that the statistics/localization diagnostic is correct, at least, for matrices with independent matrix elements.)
For random Schrödinger operators, Minami has derived Poisson statistics for the local correlations of the eigenvalue process from exponential decay of the resolvent . Some aspects of Minami’s proof translate to the present context. Most notably, the so-called Minami estimate which bounds the probability of having two eigenvalues in a small interval,
where is the length of the interval and are the eigenvalues of , holds with
Here is as in Thm. 1. (The proof of this fact may be accomplished by following Minami’s argument or by one of the various alternatives that have appeared recently in the literature .)
(In fact, the Minami estimate is proved in a similar way, by showing that the expected number of eigenvalue pairs in is bounded by the r.h.s. of (1.18).)
which has mean spacing . We say that the eigenvalue process has Poisson statistics near , in some limit and , if the point process converges to a Poisson process. The density of this Poisson process would then be given by the limit . The difficulty is we do not know that this limit exists.
Now, for a fairly general class of matrix ensembles with independent centered entries, e.g., for the Gaussian ensemble (1.1), it is known that the density of states converges weakly to the semi-circle law, provided or (see ). That is,
However, as indicated this is a weak convergence result, and it does not follow that
which would in fact be sufficient to control the density of the putative limit process.
In this regard, let us state a couple of open problems.
Improve the estimate (1.21). In particular, does this bound hold with ? (The interpretation of as the mean eigenvalue spacing and the convergence (1.24) suggests that should be bounded.)
Acknowledgments
I would like to thank Tom Spencer and Michael Aizenman for many interesting discussions related to this and other works, and to express my gratitude for the hospitality extended me by the Institute for Advanced study, where I was member when this project started, and more recently by the Isaac Newton Institute during my stay associated with the program Mathematics and Physics of Anderson Localization: 50 Years After.
Tridiagonal matrices
with and two given mutually independent sequences of independent random variables, real and complex valued respectively. For such matrices, exponential decay of the Green’s function and localization of eigenfunctions can be obtained by the transfer matrix approach, see . Here we use a different method, which is closely related to the technique of Kunz and Souillard .
To facilitate the fluctuation argument proposed above we will suppose the common distribution of has a density with the following property:
Our goal in this section is to prove the following result:
Let and be two sequences of i.i.d. random variables, real and complex valued respectively. Suppose that the common distribution of has a density which is bounded and fluctuation regular. Then for and there are and such that for all ,
We restrict to a compact set to facilitate the fluctuation argument below. In fact, for large the rate of exponential decay will improve, although the mechanism will be somewhat different. One could construct a proof in this context along the lines of . Thus the dependence of the mass of decay may be dropped.
Suppose that the distribution of , satisfies
for any interval , with a finite constant. Then
The second part of the argument is to establish large fluctuations for — this is Lemma F above. In the present context we have
Together Lemmas 2.1 and 2.2 prove Theorem 4.
Let us fix for the moment and drop it from the notation: . Suppose without loss of generality that . A preliminary observation is that
which may be established using the resolvent identity, writing as a perturbation of the corresponding matrix with set equal to zero (which decouples into two distinct blocks). Iteration of this identity gives
suggesting that if either or were to exhibit fluctuations of order one, then the variance of would be of order and Lemma 2.2 would follow. However, there are substantial correlations between the various terms, making it difficult to proceed directly along this line of argument.
To make a precise analysis, let us consider the random variables
which are related by a recursion relation
These identities may be established using the Schur-complement formula. In a similar way, the Schur-complement formula may be used to show that
where with the matrix obtained from by setting :
In particular, is a function of the variables and .
We now make a change of variables in our probability space. The Jacobian is triangular with ones on the diagonal and therefore has determinant one. Thus
So are a chain of variables with nearest neighbor couplings — thinking of as a time parameter, is a Markov chain. In terms of these variables, we have
where may be written as a function of and , since .
We now fix , and consider the conditional distribution of , which carries some information on the distribution of . A key point is that the variables remain independent after conditioning. They are, however, no longer identically distributed. Instead,
Using the conditional independence of once again to reassemble inside the expectation on the r.h.s., we find that
where we have set for .
After averaging and applying the Hölder inequality, we conclude that
Eq. (2.30) is the key result. The exponent in the first factor is a sum of non-negative terms, each presumably and positive with positive probability. It will not be so surprising to find that the term itself is with good probability. The rest is estimates.
To proceed with the estimates, let us take the a priori distribution of , before coupling and conditioning, to be uniform in an interval centered at the origin:
The r.h.s. still carries some dependence on , through the density . We may eliminate the dependence on entirely by bounding the right hand side from below:
and similarly for the term in the denominator and the term with index . Finally, the r.h.s. is no larger if we factor the infimum on the right hand side,
Plugging this estimate into eq. (2.30), we obtain
such that where is the indicator function of the event:
In turn, since and (by assumption), we see that
with and any positive numbers.
We estimate the probability of from below by integrating eq. (2.41) over , , , , and in that order. (The need to integrate over three consecutive variables is the reason we introduced only for .) To begin,
Looking now at , since , we see that
with . Since the density is bounded, it follows that
Combining these estimates with eq. (2.41) and integrating over the identically distributed variables and , we find
The key things to observe is that the r.h.s. of eq. (4.28) is independent of and can be made arbitrarily close to by suitable choice of large , and small .
So, for sufficiently small we have , say. Since , we find that
by integrating successively over from (see Lemma A.1 below). Combined with (2.39) this completes the proof of Lemma 2.2. ∎
Band matrices
Band matrix ensembles such as the Gaussian band ensemble (1.1) are of this form, with lower triangular matrices. However, for the argument presented below it is not necessary that be lower triangular. (Also, neither strict independence nor identicality of distribution are needed. Nonetheless, to keep things simple, let us stick to the i.i.d. case.)
for small , , where is the density of the distribution of (assumed to be absolutely continuous with respect to Lebesgue measure on some vector space of matrices). In the scalar case, this change of variables was useful for all fluctuation regular densities. In the matrix case, an additional complication arises. Unless falls in the vector space supporting the distribution of there will be constraints on the matrix elements of which manifest themselves as functions after the change of variables. However, is formed from and via non-linear operations, so there is no reason to expect it to fall in this vector space. (For example when are diagonal, will in general have off-diagonal components.) To guarantee closure under non-linear operations we suppose that the vector space supporting the distribution of is a matrix algebra:
Let the set of hermitian elements of . We require that and ,
Note that is closed under conjugation:
There is a good deal of flexibility in the choice of algebras. Of course, we may take all complex matrices, so is complex Hermitian. On the other hand, we could restrict to be the set of matrices with real entries, so is real symmetric. In this case is not a complex vector space. Similarly we could take to be the set of matrices with quaternion entries, where the quaternions units are represented by matrices, so would by Hermitian but anti-symmetric under transposition . In this last case, would be the set of even integers.
An important consequence of assuming that and , is that we have some a priori information on the block matrices making up the resolvent of .
Suppose is an matrix that is block tri-diagonal,
where is the algebra generated by and .
The off diagonal blocks of need not be in . This is apparent already for , where, by the Schur complement formula,
In each expression on the right, the first and last factors are in but the middle factor, , is not.
The proof is by induction on . The result is clear for . So, suppose we know that it holds if is a tridiagonal block matrix of size no larger than .
First consider (3.5). By the Schur complement formula,
As and are of size no larger than and , it follows that
By Prop. 5
Now consider (3.6). Suppose (the other case is similar). Let
But by the induction hypothesis and as we have just shown. It follows that the r.h.s. is in ∎
and let denote the set of eigenvalues of a matrix. Recall, if is self-adjoint, that
(Wegner-type estimates): There are and such that for all , ,
and for all and , ,
If is suitably scaled so as to have mean eigenvalue spacing of order , this suggests that we should be able to take . That has not been proved, however, for the random matrix ensembles studied here. For Wigner type matrices, in particular for the Gaussian band ensemble (1.1), we will obtain the estimates (3.8, 3.9) with in §5,.
The parameters and are not independent. If we rescale via this results in a shift and . Nonetheless it is convenient to keep both parameters since the natural scaling of is to choose the eigenvalue spacing to be of order . This typically leads to , but if the entries of have heavy tails then one may have .
Lemma W for follows easily from part (2) of assumption 1.
Let us first consider the case . The Schur complement formula shows that
with and the restrictions of to the blocks above and below . By Lemma 3.1 . (Note that it is self adjoint.) It follows from (3.8) that is an eigenvalue of with probability and that (3.14) holds for .
The argument for is similar. In this case, we first estimate
where and are formed from blocks of the resolvents of restrictions of . One may verify that and Thus the result follows from (3.9). ∎
for any two vectors . (See (1.8).)
Below we will apply the result with a semi-norm such as the the absolute value of a matrix element or the norm . However the proof does not make use of the triangle inequality, so the result also applies, for example, to spectral radius or smallest singular value of .
Under rescaling of the matrix elements the localization length should not change. That this is indeed so follows since , , and , so the combination is invariant under rescaling.
where and denote elementary basis vectors. Then
and given there are constants such that for any
Putting (3.20) and (3.19) together we have
If the diagonal blocks are Wigner matrices, as in assumption 4 in §5 below, one may obtain the estimate
resulting in a very slight improvement on the estimate on the r.h.s. of (3.21),
This improvement is not very significant, as the main point here is the exponential factor, which dominates any power of as long as .
Fluctuations
We now prove Lemma 3.3. Following the proof of Lemma 2.2, let us fix and set
Since , in estimating we may assume without loss that . We have, by the resolvent identity,
Let us define random matrices
As in the case, these identities may be established using the Schur-complement formula — compare with (2.9) and (2.13). Similarly,
where with the matrix obtained from by setting . Thus is a function of the matrix variables and .
We now make the change of variables in our probability space. By Lem. 3.1 and Prop. 5, . As in the tri-diagonal case, the Jacobian determinant is , so
where is a function of and (since ).
The matrix product in (4.7) is non-commutative, so it is not clear if the heuristic analysis that the “log of is a sum of terms with only local correlations” is valid. Nonetheless, we may use the trick employed above of coupling the system to a family of independent identically distributed scalar variables , each with absolutely continuous distribution
with to be chosen below. We define
As in the tri-diagonal case, the variables remain independent after conditioning on . Also, the
is a function of . Since are conditionally independent, it follows that
By propostion 3 and the Hölder inequality, we conclude that (compare with (2.30)):
with as in (2.28), and we have set for .
with , and as in assumption (3). In turn, since , we see that
with as in assumption 2 and as in assumption 3. This allows us to estimate the probability of from below by successively integrating over , , , , and in that order. To begin, by assumption 2,
Since , we see from the Wegner estimate (3.8) that
Combining these estimates and using assumption 3 to integrate over and , we find
Taking with , we may choose sufficiently small to make the r.h.s. larger than , say. Since we find, integrating successively over from (see Lemma A.1), that
Increasing , if necessary, so that completes the proof. ∎
Ensembles
In this section, we consider several examples of band matrix ensembles satisfying assumptions 1, 2, and 3 of section 3. Assumption 1 is simply the choice of an algebra to support the distribution of the diagonal blocks, and the corresponding set for the off-diagonal blocks. In this regard, we will consider two cases:
matrices with real entries,
matrices with complex entries.
In each case the dimension of the algebra is comparable to and .
We shall suppose that the diagonal blocks of are Wigner matrices:
Under assumption 4, the Wegner estimates (3.8) and (3.9) hold with and
This result, which is obtained by averaging over the diagonal variables only, is a standard estimate from the theory of random Schrödinger operators, first obtained by Wegner . For completeness, we sketch the proof.
Note that if and only if has an eigenvalue in the interval . It follows that
where is a function of all matrix elements of except . Thus is a random variable independent of , so
The proof of (3.9) is analogous. However in that case the trace is over a dimensional space, resulting in the additional factor of on the r.h.s. of that equation. ∎
The scaling factor that appears in (5.11) is natural, as with this scaling the matrix has a finite density of states in the large limit :
Let be a random matrix of the form (5.11), with and mutually independent sets of independent random variables. If
This follows from Theorem A of ref. , which gives the convergence of , the largest eigenvalue of , to with probability one. Symmetrizing the assumptions of Theorem A and applying the result also to show that , the smallest eigenvalue of , converges to , this result follows. (The proof in is written out in the real symmetric case, but carries over to the complex hermitian case with only very minor modifications.)
Under assumption 4, we may find such that
We require very little of the off diagonal blocks . They need only satisfy the estimate (3.13) analogous to (3.10) and (5.10). In particular, they could be deterministic, say for all or given by a Toeplitz matrix. In this section we consider a few examples of random off-diagonal blocks modeled on the blocks for the Gaussian band ensemble (1.1). In that case, the off-diagonal blocks are lower triangular matrices with Gaussian entries. More generally we may suppose
Under assumption 5, we may find such that assumption 3 holds with , i.e.,
2. Fluctuation regularity
If is a GUE or GOE matrix of size then assumption 2 of section 3 holds with , and .
Assumptions 1, 2, and 3 hold for the Gaussian band ensemble (1.1).
We have already derived the Wegner estimates (Thm. 7). It remains only to show the fluctuation regularity. For the Gaussian ensembles, we have
If we have
Letting and be as in Cor. 9, we set . Then and if , we have
whenever . ∎
To obtain fluctuation regularity for general Wigner matrices (3.11) we require additional assumptions on and . For instance, we have the following
If satisfies assumption 4 with and uniformly Hölder continuous with exponent , then assumption 2 of section 3 holds with , and .
If , then
This estimate holds for every , so in particular for all in . ∎
Theorem 13 cannot apply if or has compact support. Nonetheless compactly supported densities can be handled. A general result of this type would somewhat involve to state, so let us simply note that assumption 2 holds if and are characteristic functions of open neighborhoods of the origin.
Suppose that satisfies assumption 4 and that
with and . Then assumption 2 of section 3 holds with , and .
Clearly the moment conditions of assumption 4 hold. Thus by Cor. 9 we can find and so that (5.10) holds.
Now suppose . Suppose also that the matrix elements of satisfy , for all . Then
3. Summary
Putting the results of this section together with Thm. 6 we have:
In particular, (5.29) holds for the Gaussian band ensemble (1.1).
If and are uniformly Hölder continuous with exponent , then given and there are and such that
If and are proportional to characteristic functions of open neighborhoods of the origin, then given and there are and such that (5.30) holds with
Appendix A A lemma on conditional averages
In the proofs of the various versions of Lemma F above, a key step was to estimate averages of the form
in which are non-negative, strictly positive with good probability, but not independent. The following Lemma gives the relevant estimate, which can be seen as a simple version of stochastic domination. As the proof shows, under appropriate assumptions, we can estimate (A.1) in terms of the same expression with replaced by i.i.d. non-negative Bernoulli variables taking with probability less than .
Let be a sequence of -algebras of events on a probability space and let be a sequence of non-negative random variables with measurable with respect to for . If for some ,