On Bernoulli Decompositions for Random Variables, Concentration Bounds, and Spectral Localization

Michael Aizenman, Francois Germinet, Abel Klein, Simone Warzel

Introduction

This article has a twofold purpose. As a general observation it is noted that in any random variable one may find a Bernoulli component. A decomposition which is based on the above observation allows then to extend results which for systems of Bernoulli variables are available by combinatorial methods to systems of random variables of arbitrary distribution.

A Bernoulli decomposition of a real-valued random variable XX is a representation of the form

where Y(⋅)Y(\cdot) and δ(⋅)≥0\delta(\cdot)\geq 0 are functions on (0,1)(0,1), the variable tt is uniformly distributed in (0,1)(0,1), and η\eta is a Bernoulli random variable taking the values {0,1}\{0,1\} with probabilities {1−p,p}\{1-p,p\} independently of tt. The relation in (1.1) is to be understood as expressing equality of the distributions of the corresponding random variables.

Bernoulli decompositions are constructed here for arbitrary random variables of non-degenerate distributions. For certain purposes it is useful to have positive uniform conditional variance of the Bernoulli term, i.e.,

We present such a representation below and discuss related issues of optimality.

Two applications mentioned here: i. anti-concentration bounds for monotone, though not necessarily linear, functions of independent random variables, and ii. a proof, based on the Bernoulli case [BK], of spectral localization for random Schrödinger operators with arbitrary probability distributions for the single site coupling constants.

In the first application, we consider functions Φ(X1,…,XN)\Phi(X_{1},\ldots,X_{N}) of independent non-degenerate random variables {Xj}\{X_{j}\} whose distributions are either identical or, in a sense explained below, are of widths greater than some common bX>0b_{X}>0. It is shown here that if for some ε>0\varepsilon>0 the function satisfies

with a constant CX<∞C_{X}<\infty which depends on the uniform bounds on the distributions of {Xj}\{X_{j}\}. The proof employs the Bernoulli representation along with the combinatorial bounds of Sperner [S], and the more general LYM lemma [E].

The use of combinatorial estimates for concentration bounds first appeared in the context of Bernoulli variables in P. Erdös’ variant of the Littlewood-Offord Lemma [Er]. The presence of a Bernoulli component in any random variable was noted implicitly in the work of A. N. Kolmogorov [Ko] where it was put to use in an improvement of the earlier concentration bounds of W. Doeblin and P. Lévy [DoL, Do] on linear functions of independent random variables. Initially, Kolmogorov did not extract the maximal benefit from the method by not connecting it with Sperner theory, and in particular the concentration bound in [Ko] includes an unnecessary logarithmic factor; the corresponding improvement was made by B. A. Rogozin [R1]. The bounds were further improved in a series of works, in particular [Es, K, R2] where use was also made of other methods. One may note here that perhaps quite naturally a general method like the Bernoulli decomposition is not optimized for specific applications. Nevertheless, it has the benefit of providing a simple perspective on a number of topics.

In our second application, we establish spectral localization for a broad class of continuum, alloy-type random Schrödinger operators (cf. (4.1)), building on a result of J. Bourgain and C. Kenig [BK] for the Bernoulli case. The model and the results are presented more explicitly in Section 4. The main point to be made here is that the understanding of spectral localization for the Bernoulli case can be extended through the Bernoulli decomposition to random operators with single site coupling parameters of arbitrary distribution (cf. Theorem 4.2).

Bernoulli decompositions for random variables

Randomness often is in the eyes of the beholder, as probability measures are used to express averages over specified sets of rather varied nature. However, it may be true that the most elementary model underlying the basic popular perception of probability is the simple ‘coin toss’, with two possible outcomes: heads or tails, which is modeled by a Bernoulli random variable: a binary variable equal to 11 with probability pp and equal to with probability 1−p1-p.

The following statement assert that any real valued random variable has a Bernoulli component, which can even be chosen to be of uniformly positive variance.

Given a real random variable XX by default we shall denote its probability distribution by μ\mu and let G:(0,1)→(−∞,∞)G:(0,1)\to(-\infty,\infty) be the function defined by

One may observe that GG is the ‘inverse’ distribution function of μ\mu, which takes values in the essential range of XX. It can alternatively be described by

Let XX be a non-degenerate real-valued random variable with a probability distribution μ\mu. Then, for each p∈(0,1)p\in(0,1), XX admits a decomposition of the form:

in the sense of equality of the corresponding probability distributions, where:

η\eta and tt are independent random variables, with η\eta a binary variable taking the values {0,1}\{0,1\} with probabilities {1−p,p}\{1-p,p\}, correspondingly, and tt having the uniform distribution in (0,1)(0,1),

Yp: (0,1)↦(−∞,∞)Y_{p}:\,(0,1)\mapsto(-\infty,\infty) is the monotone non-decreasing function

δp+: (0,1)↦[0,∞)\delta_{p}^{+}:\,(0,1)\mapsto[0,\infty) is the function

for at least one value of p∈(0,1)p\in(0,1) we have

Some explicit expressions for β+(p,μ)\beta^{+}(p,\mu) are mentioned in Remark 2.1 below. The Bernoulli component of the measure is not a uniquely defined notion, and other representations similar to (2.3) but with different distributions for the conditional variance of the Bernoulli component, i.e., for δ(t)\delta(t), can also be obtained. In the following construction its uniform positivity may be lost but one gains the feature that the range of values which δ\delta assumes reaches up to the diameter of the support of the measure μ\mu.

Let XX be a non-degenerate real-valued random variable with probability distribution μ\mu. Then, for each p∈(0,1)p\in(0,1), XX admits a decomposition of the form:

where tt, η\eta and the function YpY_{p} are as in Theorem 2.1, satisfying the above (1) and (2), but instead of (3) and (4) the following holds

δp−: (0,1)↦[0,∞)\delta_{p}^{-}:\,(0,1)\mapsto[0,\infty) is the non-increasing function:

for any x−<x+x_{-}<x_{+} and p±>0p_{\pm}>0 such that

at the particular value p=p+p−+p+p=\frac{p_{+}}{p_{-}+p_{+}} we have

where the probability is with respect to the uniform random variable tt.

In the proofs we employ two versions of what is called here the Pac-Man algorithm for the construction of a joint distribution ρ\rho of a pair of random variables, of the form {Y1(t),Y2(t)}\{Y_{1}(t),Y_{2}(t)\}, whose marginal probability measures, ρ1\rho_{1}, ρ2\rho_{2}, satisfy

The representations (2.3) and (2.7) correspond to letting:

The two Theorems will be proven in reverse order.

This relation allows to represent, in terms similar to (2.7), as

with tt the random variable with the uniform distribution in $$.

Extending the above representation, we now define a pair of coupled random variables through the following functions of t∈t\in:

By (2.16) the random variable seen on the right side of (2.7) has the same distribution as XX. The statement (2) readily follows from the definition (2.1) and (2.1).

For a proof of (4’) we note that (2.9) is equivalent to

This implies δp+(t)>x+−x−\delta_{p}^{+}(t)>x_{+}-x_{-} for all t≤p++p−t\leq p_{+}+p_{-}, and hence (2.10) holds true. ∎

For the representation (2.3) we shall employ the following variant of (2.1):

In this case, both Y1Y_{1} and Y2Y_{2} are monotone non-decreasing in tt and

for all t∈(0,1)t\in(0,1), where G(1−p+0)=lim⁡ε↓0G(1−p+ε)G(1-p+0)=\lim_{\varepsilon\downarrow 0}G(1-p+\varepsilon). Moreover, for any T∈(0,1)T\in(0,1) we have the lower bound

For a sufficient condition for the uniform positivity of δp+(t)=Y2(t)−Y1(t)\delta_{p}^{+}(t)=Y_{2}(t)-Y_{1}(t) let us consider the arrival/departure times:

The times T1,T2T_{1},T_{2} are non-random and depend on pp and μ\mu only. If

then for each T∈(T2,T1)T\in(T_{2},T_{1}) we have

The collection of p∈(0,1)p\in(0,1) such that (2.22) is not empty whenever the support of the measure includes more than one point. ∎

Explicit lower bounds on β+\beta^{+}. For the Bernoulli decomposition which is presented in Theorem 2.1 (i.e., based on the ‘chasing Pac-Men’ algorithm), an expression for the lower bound β+(p,μ)\beta^{+}(p,\mu) in terms of the distribution function of μ\mu is given in (2.34) below. A simple lower bound can be obtained in terms of just the “half-time” points for the two markers. i.e., from (2.20) with T=12T=\frac{1}{2}:

This shows that for continuous measures μ\mu one has β+(p,μ)>0\beta^{+}(p,\mu)>0, i.e., (2.6), for any p∈(0,1)p\in(0,1).

At the particular value p=1−μ((−∞,x^))p=1-\mu\left((-\infty,\hat{x})\right) we then have G(1−p)=x^G(1-p)=\hat{x} and

An alternative form. For another form of a Bernoulli decomposition, with a binary random variable σ=±1\sigma=\pm 1, let

When such a substitution is made in (2.3) the two resulting functions W(t)W(t) and δp+(t)\delta_{p}^{+}(t) are monotone non-decreasing in tt and δp+(⋅)\delta_{p}^{+}(\cdot) is constant over each interval of constancy of W(⋅)W(\cdot). It follows that the value of δp+(t)\delta_{p}^{+}(t) can be expressed in terms of W(t)W(t), and thus one obtains a representation of the form:

with WW and σ\sigma independent random variables, and b(⋅)b(\cdot) a measurable function which is determined by μ\mu and pp.

2. Optimality of the Pac-Man algorithm

In applications of the decomposition it is desirable to maximize the conditional variance of the binary term. We shall now address related questions from an optimal transport perspective, and in particular establish optimality, in a certain limited sense, of the ‘chasing Pac-Men’ construction.

In addition to the explicit choices presented in Theorems 2.1 and 2.2 there are other possibilities for a Bernoulli decomposition of the form (2.3). With a change of variables as in (2.1), such representations can alternatively be expressed in terms of joint distributions of the variables Y1Y_{1}, Y2Y_{2} with the properties listed in the following definition.

where F(x)=μ((−∞,x])F(x)=\mu((-\infty,x]), and Fj(x)=ρ({Yj≤x})F_{j}(x)=\rho(\{Y_{j}\leq x\}) for j=1,2j=1,2.

The minimal conditional variation β∗(p,ρ)\beta_{*}(p,\rho) is maximized by the ‘chasing Pac-Men’ algorithm which is presented in the proof of Theorem 2.1, i.e. for any Bernoulli decomposition

where ess inf⁡t∈(0,1){\operatorname{ess\,inf}}_{t\in(0,1)} yields the same value as inft∈(0,1)\rm{inf}_{t\in(0,1)}.

The maximal conditional variation β#(p,ρ)\beta^{\#}(p,\rho) is maximized by the ‘colliding Pac-Men’ algorithm of Theorem 2.2, for which β#(p,ρ)\beta^{\#}(p,\rho) equals the diameter of the essential support of μ\mu.

The equality: essinft∈δp+(t)=inft∈δp+(t)\rm{essinf}_{t\in}\delta_{p}^{+}(t)=\rm{inf}_{t\in}\delta_{p}^{+}(t) is a simple consequence of the left-continuity property of the chasing Pac-Men algorithm, where Yj(t)=Yj(t−0)Y_{j}(t)=Y_{j}(t-0) and hence also δ+(t)=δ+(t−0)\delta^{+}(t)=\delta^{+}(t-0).

To prove (2.32) let us first establish a helpful expression for β+(p,μ)\beta^{+}(p,\mu). Denoting by Fj+F^{+}_{j} the distribution functions corresponding to Y1Y_{1} and Y2Y_{2} of (2.1) we have:

For the ‘chasing Pac-Men’ construction, of Theorem 2.1:

The statements (2.33) follow directly from the definition of the Pac-Man process (2.1). In the derivation of (2.34), we shall use the fact that for all t∈(0,1)t\in(0,1) and ε>0\varepsilon>0:

It follows that for any t∈(F1(x),F2(x+u))t\in(F_{1}(x),F_{2}(x+u)):

and therefore δ+(t)=Y2(t)−Y1(t) ≤ u\delta^{+}(t)=Y_{2}(t)-Y_{1}(t)\ \leq\ u. Thus: inft∈(0,1)δp+(t)≤S\rm{inf}_{t\in(0,1)}\delta_{p}^{+}(t)\leq S.

which implies that Y1+(t)+u≤Y2+(t)Y_{1}^{+}(t)+u\leq Y_{2}^{+}(t). Therefore

It follows that inft∈δp+(t)≥S\rm{inf}_{t\in}\delta_{p}^{+}(t)\geq S, which completes the proof of (2.34). ∎

The second assertion is an elementary consequence of (2.8). To prove (1) we shall show that for any b>β+(p,μ)b>\beta^{+}(p,\mu) it is also true that b>β∗(p,ρ)b>\beta_{*}(p,\rho).

The condition (2.30) readily implies that (1−p) F1(u)≤F(u)(1-p)\,F_{1}(u)\leq F(u), or F1(x) ≤ min⁡{(1−p)−1F(x), 1}F_{1}(x)\ \leq\ \min\left\{(1-p)^{-1}F(x),\,1\right\}, and hence

Eq. (2.42) means that ρ({Y1≤u})≤t\rho\left(\{Y_{1}\leq u\}\right)\leq t and ρ({Y2>u+b})<1−t\rho\left(\{Y_{2}>u+b\}\right)<1-t. Since the probabilities of the two events add to less than 11 the complement of their union is of positive probability, and this implies:

and hence b>β∗(p,ρ)b>\beta_{*}(p,\rho). This concludes the proof of (2.32). ∎

The idea of seeking optimal joint realizations of random variables with constrained marginals has allowed to present a wide range of analytical results from a common ‘optimal transport’ perspective (see, e.g., [V]). The most familiar variants of the problem concern couplings which minimize a distance function between the two coupled variables. As our discussion demonstrates, it may also be of interest to seek couplings which maximize the difference between the two variables with constrained marginals.

Concentration Bounds

We shall now demonstrate how the Bernoulli decomposition yields probabilistic bounds from combinatorial results. If there is any novelty in this section it is in the formulation of the bounds for the non-linear case, as the two main ideas were noted before in the context of linear functions: P. Erdös [Er] observed that concentration bounds for linear functions of Bernoulli variables can be derived from the combinatorial theory of E. Sperner [S], and B. A. Rogozin [R1] has used the Bernoulli decomposition of A. N. Kolmogorov [Ko] for the further extension of these bounds to arbitrary random variables.

First, we present some essentially known results of Sperner theory; in the second subsection these results will be combined with the Bernoulli decomposition to yield concentration bounds for functions of independent random variables.

The configuration space {0,1}N\{0,1\}^{N} for a collection of Bernoulli random variables η={η1,...,ηN}\boldsymbol{\eta}=\{\eta_{1},...,\eta_{N}\} is partially ordered by the relation defined by:

A set A⊂{0,1}N{\mathcal{A}}\subset\{0,1\}^{N} is said to be an antichain if it does not contain any pair of configurations which are compatible in the sense of “≺\prec”. The original Sperner Lemma states that for any such set: ∣A∣≤(N[N2])\left\lvert\mathcal{A}\right\rvert\leq\binom{N}{[\frac{N}{2}]}. A more general result is the LYM inequality for antichains (cf. [An]):

where ∣η∣=∑ηj\left\lvert\boldsymbol{\eta}\right\rvert=\sum\eta_{j}.

The LYM inequality has the following probabilistic implication.

Let {ηj}\{\eta_{j}\} be independent copies of a Bernoulli random variable η\eta with

where p ∈(0,1)p\,\in(0,1). Then for any antichain A⊂{0,1}N\mathcal{A}\subset\left\{0,1\right\}^{N}:

where η=(η1,…,ηN)\boldsymbol{\eta}=(\eta_{1},\ldots,\eta_{N}), ση=pq\sigma_{\eta}=\sqrt{pq} is the standard deviation of η\eta, and Θ\Theta is an independent constant which does not exceed 222\sqrt{2}.

Let AkA_{k} be the subset of A\mathcal{A} consisting of configurations with ∣η∣=k\left\lvert\boldsymbol{\eta}\right\rvert=k. Then:

where b(k;N,p):=pkqN−k(Nk)b(k;N,p):=p^{k}q^{N-k}\binom{N}{k} is the binomial distribution, and the inequality is by (3.2). The maximum of b(k;N,p)b(k;N,p) over kk, which is known to occur near k=pNk=pN (cf. [F, Theorem 1 on p. 140]) yields (3.4). ∎

The bound (3.4) has the virtue of being valid for all NN; for N→∞N\to\infty it holds with a smaller constant which tends to the asymptotic value Θ→1/2π\Theta\to 1/\sqrt{2\pi} (implied by (3.5) and Stirling’s formula).

Following is an extension of Lemma 3.1 to the case of non-identically distributed random variables.

Let η=(η1,…,ηN)\boldsymbol{\eta}=(\eta_{1},\ldots,\eta_{N}), where {ηj}\{\eta_{j}\} are independent Bernoulli random variables with possibly different values of pjp_{j}, and set

Then, for any antichain A⊂{0,1}N\mathcal{A}\subset\left\{0,1\right\}^{N}:

where Θ~\widetilde{\Theta} is an independent constant which does not exceed 44.

The proof gives us the chance to introduce the technique of ‘double sampling’.

We start from the observation that any Bernoulli variable η\eta with parameter pηp_{\eta} as in (3.3) may be decomposed in terms of two independent Bernoulli variables χ\chi and ξ\xi as

By the definition of α\alpha, eq. (3.6), pj∈[α,1−α]p_{j}\in[\alpha,1-\alpha] for all j=1,2,…,Nj=1,2,\ldots,N. Hence the variables η\boldsymbol{\eta} may be represented as in (3.8) with independent identically distributed (iid) Bernoulli variables {χj}\{\chi_{j}\} with common pχ:=1−αp_{\chi}:=1-\alpha. We abbreviate this representation as ξ χ:=(ξ1χ1,…,ξNχN)\boldsymbol{\xi}\,\boldsymbol{\chi}:=(\xi_{1}\chi_{1},\ldots,\xi_{N}\chi_{N}). Evaluating the probability by first conditioning on the values of ξ\boldsymbol{\xi}, one has

For specified values of the variables χ\boldsymbol{\chi} , the event A\mathcal{A} depends only on the values of χj\chi_{j} with jj in the set Jξ:={j : ξj≠0}J_{\boldsymbol{\xi}}:=\{j\,:\,\xi_{j}\neq 0\}, and as such it is an antichain in {0,1}Jξ\{0,1\}^{J_{\boldsymbol{\xi}}}. Bounding its conditional probability by Lemma 3.1 we obtain

where σχ=α(1−α)\sigma_{\chi}=\sqrt{\alpha(1-\alpha)} is the common standard deviation of χj\chi_{j}.

The event {∣Jξ∣ ≤ αN/2(1−α)}\left\{|J_{\boldsymbol{\xi}}|\ \leq\ \alpha N/2(1-\alpha)\right\} is of exponentially small probability, as can be seen by a standard large deviation estimate for independent variables. It then readily follows that

with a constant for which elementary estimates yield Θ~≤4{\widetilde{\Theta}}\leq 4. ∎

For completeness it should be added that in addition to the anti-concentration upper bounds it is of interest to know the asymptotic behavior. That is covered by known results, such as is presented in Engel [E, Theorem 7.2.1]:

which amounts to a ‘local’ central limit theorem (CLT).

2. Concentration Bounds for Functions of Independent Random Variables

We shall now employ the Bernoulli decomposition of Section 2, along with the results presented in the previous subsection, for an upper bound on the concentration probability

where {Xj}\left\{X_{j}\right\} are independent random variables.

Let X=(X1,…,XN)\boldsymbol{X}=(X_{1},\ldots,X_{N}) be a collection of independent random variables whose distributions satisfy, for all j∈{1,...,N}j\in\{1,...,N\}:

where 44 can also be replaced by the constant Θ~\widetilde{\Theta} of (3.7).

We start by selecting p∈(0,1)p\in(0,1) by the condition p=p+p++p−p=\frac{p_{+}}{p_{+}+p_{-}}. Next, we represent the variables {Xj}\{X_{j}\} using Theorem 2.2:

with η=(η1,…,ηN)\boldsymbol{\eta}=(\eta_{1},\ldots,\eta_{N}) a collection of iid Bernoulli variables taking values {0,1}\{0,1\} with probability {1−p,p}\{1-p,p\}. From (2.10) one may conclude that for all j∈{1,…,N}j\in\{1,\ldots,N\}:

By virtue of (3.17), the set At\mathcal{A}_{\boldsymbol{t}} is an antichain in its dependence on {ηj}j∈Jt\{\eta_{j}\}_{j\in J_{\boldsymbol{t}}} with Jt:={j : δj(tj)≥x+−x−}J_{\boldsymbol{t}}:=\left\{j\,:\,\delta_{j}(t_{j})\geq x_{+}-x_{-}\right\}. Lemma 3.1 thus yields

with ση=p(1−p)\sigma_{\eta}=\sqrt{p(1-p)}. We conclude by the large-deviation argument used in the proof of Lemma 3.2. Using (3.20) the expected value of ∣Jt∣=∑j=1N1{j : δj(tj)≥x+−x−}|J_{\boldsymbol{t}}|=\sum_{j=1}^{N}1_{\{j\,:\,\delta_{j}(t_{j})\geq x_{+}-x_{-}\}} is bounded below:

Therefore {∣Jt∣≤12(p++p−)N}\{|J_{\boldsymbol{t}}|\leq\frac{1}{2}(p_{+}+p_{-})N\} is a large deviation event and its probability is exponentially bounded. Elementary estimates lead to

with the same constant Θ~\widetilde{\Theta} as in (3.7). ∎

A simpler proof for iid variables. For iid non-degenerate random variables X1,…,XNX_{1},\ldots,X_{N} the theorem has a simpler proof using the binary decomposition of Theorem 2.1; there is no need for the large deviation argument. The constants in the theorem will then depend on the value of pp and its corresponding lower bound in (2.6).

concentration inequalities as in (3.18) go back to W. Doeblin, P. Lévy [DoL, Do], P. Erdös [Er] (for the Bernoulli case, where it reduces to the Littlewood-Offord problem), A. N. Kolmogorov [Ko], B. A. Rogozin [R1], H. Kesten [K] and C. G. Esseen [Es]. In this case, sharper inequalities than (3.18) are known, e.g. [R3],

where Θ\Theta is some constant. A recent application of the discrete case of the concentration bounds is found in [TV].

An extension. As it is already true for (3.27), the statement of Theorem 3.1 has an immediate extension to functions which in some variables are monotone increasing and in some are monotone decreasing, satisfying the natural analog of (3.17). For this extension, one only needs to replace p+p_{+} and p−p_{-} in (3.18) by p^=min⁡{p+,p−}\hat{p}=\min\{p_{+},p_{-}\}.

and for which Φ(u)=0\Phi(\boldsymbol{u})=0 if and only in u∈A\boldsymbol{u}\in\mathcal{A}.

An Application to Random Schrödinger Operators

As a demonstration of a possible uses of the elementary observations which are made in this article, let us present the case of spectral localization under random iid single site potential for an arbitrary probability distribution.

The (continuum) Anderson Hamiltonian is the random Schrödinger operator given by

Although definitions of localization may come in several flavors, they all include (or imply) spectral localization (i.e., pure point spectrum), as given in the following definition.

This property is clearly invariant under translations. The defining condition is equivalent to the requirement that for a spanning set of vectors the spectral measure is pure-point within II. The set of ω\boldsymbol{\omega} for which this holds for the random operator HωH_{\boldsymbol{\omega}} is known to be measurable.

In the one-dimensional case the continuous Anderson Hamiltonian has been long known to exhibit spectral localization in the whole real line for any non-degenerate μ\mu, i.e. when the random potential is not constant [GoMP, DSS]. In the multidimensional case, localization at the bottom of the spectrum is already known at great, but nevertheless not all-inclusive, generality; cf. [St, Kl, BK] and references therein. The Bernoulli decomposition presented here allows to prove localization for general non-degenerate single site distributions μ\mu.

More explicitly, the simplest case to deal with, for the different approaches which yield proofs of localization, has been when the single site probability distribution is absolutely continuous with bounded derivative. The absolute continuity condition can be relaxed to Hölder continuity of μ\mu, both in the approach based on the multiscale analysis which was introduced in [FrS] and is discussed in [Kl], and in the one based on the fractional moment method of [AM, AE+]. (The basis in the former case is an improved analysis of the Wegner estimate, which can be found in [St, CHK].) However, techniques relying on the regularity of μ\mu seem to reach their limit with log-Hölder continuity. In particular, until recently the Bernoulli random potential had been beyond the reach of analysis in more than one dimension. For that extreme case, i.e., of HωH_{\boldsymbol{\omega}} with μ{1}=μ{0}=12\mu\left\{1\right\}=\mu\left\{0\right\}=\frac{1}{2}, localization at the bottom of the spectrum was recently proven by Bourgain and Kenig [BK]. A crucial step in the analysis of [BK] is the estimation of the probabilities of energy resonances using Sperner’s Lemma, i.e., the p=12p=\tfrac{1}{2} version of (3.4).

The point which we would like to make here is that the Bernoulli decomposition of random variables enables one to turn the latter result of Bourgain and Kenig [BK] into a tool for a general proof of localization at the edge of the spectrum for arbitrary non-degenerate μ\mu.

where u(⋅)u(\cdot) is as in (4.2), satisfying the above condition (1), but instead of (2):

Due to the presence of the background potential UU the spectrum of HηH_{\boldsymbol{\eta}} need not be deterministic, i.e., equal to some fixed set with probability one. For our main purpose it would suffice to restrict attention to UU for which the spectrum of HηH_{\boldsymbol{\eta}} is almost surely [0,∞)[0,\infty). Such restriction is not included in the following statement but instead there is a caveat in the conclusion.

The extended BK result, whose proof is presented in [GK], is:

Given a function u(⋅)u(\cdot) as above, and: p∈(0,1)p\in(0,1), b±>0b_{\pm}>0 and U+<∞U_{+}<\infty, there exist E0>0E_{0}>0 such that any random operator HηH_{\boldsymbol{\eta}} of the form (4.4), satisfying conditions (1), (2’) and (3), for otherwise arbitrary external potential UU, with probability one, either exhibits spectral localization in [0,E0][0,E_{0}] or σ(Hη)∩[0,E0]=∅\sigma(H_{\boldsymbol{\eta}})\cap[0,E_{0}]{=}\emptyset.

Theorem 2.1 allows now to deduce the following general statement from the above non-trivial Bernoulli result.

Let Hω=−Δ+VωH_{\boldsymbol{\omega}}=-\Delta+V_{\boldsymbol{\omega}} be a Schrödinger operator with the random potential given by (4.2), satisfying the above conditions (1) and (2). Then for some E0>0E_{0}>0 the operator HωH_{\boldsymbol{\omega}}, with probability one, exhibits spectral localization in [0,E0][0,E_{0}].

The Bernoulli decomposition (2.3) allows to write the coefficients in the random potential in the form:

As a consequence, the random operator can be written as:

This implies that when conditioned on the values of t\boldsymbol{t} the operator Ht,ηH_{\boldsymbol{t},\boldsymbol{\eta}} is of the form (4.4), with pp, U+U_{+} and b±b_{\pm} independent of t\boldsymbol{t}. Thus, by Theorem 4.1 there exists E0>0E_{0}>0 such that when conditioned on t\boldsymbol{t} with probability one, Ht,ηH_{\boldsymbol{t},\boldsymbol{\eta}} either exhibits spectral localization or has no spectrum in [0,E0][0,E_{0}]. However, the latter is excluded (almost surely, also with respect to the conditional probability) by (4.3) and Fubini. ∎

In addition to the spectral localization it is also of interest to establish the existence of uniform localization length, i.e., to prove that all eigenfunctions ϕ\phi of HωH_{\boldsymbol{\omega}} with eigenvalue in [0,E0][0,E_{0}] satisfy

This can be accomplished in the following two ways, for which the details are presented in [GK].

To establish uniform localization length under the hypotheses of Theorem 4.2 one may use the Bernoulli decomposition (4.6) before performing the multiscale analysis which is behind the proof of Theorem 4.1. The multiscale analysis is then executed for the random Schrödinger operator Ht,ηH_{\boldsymbol{t},\boldsymbol{\eta}} in (4.7), in such a way that all events in the analysis are jointly measurable in t\boldsymbol{t} and η\boldsymbol{\eta}.

An alternative proof of Theorem 4.2, which yields also uniform localization length, can be based on the concentration bound of Theorem 3.1. Namely, the Bourgain-Kenig proof can be extended to arbitrary single site probability distribution μ\mu, with the probabilities of energy resonance estimated by the concentration bound instead of by Sperner’s Lemma as in [BK] (see [GK]).

Acknowledgements

We thank the Oberwolfach center for hospitality at a meeting where the four-way collaboration started, and the Isaac Newton Institute where some of the work was done. We also thank B. Sudakov for an instructive review of the recent results in Sperner’s theory, S. Molchanov for alerting us to relevant references, and M. Cranston for many helpful discussions of results in probability theory. This work was supported in parts by the NSF grants DMS-0602360 (MA), DMS-0457474 (AK) and DMS-0701181 (SW), and a Rothschild Fellowship at INI (MA).

References