On generic chaining and the smallest singular value of random matrices with heavy tails

Shahar Mendelson, Grigoris Paouris

Introduction

The main goal of this article is to obtain a non-asymptotic version of the Bai-Yin Theorem on the largest and smallest singular values of certain random matrices. The Bai-Yin theorem asserts the following:

Let A=AN,nA=A_{N,n} be an N×nN\times n random matrix with independent entries, distributed according to a random variable ξ\xi, for which

If N,n→∞N,n\to\infty and the aspect ratio n/Nn/N converges to β∈(0,1]\beta\in(0,1], then

almost surely, where smax⁡s_{\max} and smin⁡s_{\min} denote the largest and smallest singular value of AA.

Also, without the fourth moment assumption, smax⁡(A)/Ns_{\max}(A)/\sqrt{N} is almost surely unbounded.

The main result of this article is a quantitative version of the Bai-Yin Theorem.

We will focus on the following questions:

Observe that the two questions are very similar. For example, it is straightforward to verify that if μ\mu is isotropic, then both parts can be resolved by estimating the supremum of the empirical process

And, in view of the second part of Question 1.2, we will be especially interested in the case N∼nN\sim n, that is, while keeping the aspect ratio n/Nn/N constant.

To formulate the moment assumption we will use here, recall that for α≥1\alpha\geq 1, the ψα\psi_{\alpha} Orlicz norm of random variable ZZ is defined by

and there are obvious extensions for 0<α<10<\alpha<1. It is standard to verify that for every α>0\alpha>0, ∥Z∥ψα\|Z\|_{\psi_{\alpha}} is equivalent to sup⁡q≥1∥Z∥Lq/q1/α\sup_{q\geq 1}\|Z\|_{L_{q}}/q^{1/\alpha}.

For p,q≥2p,q\geq 2, a symmetric measure μ\mu satisfies a pp-small diameter, LqL_{q} moment assumption with constants κ1\kappa_{1} and κ2\kappa_{2}, if a random vector XX distributed according to μ\mu satisfies that

μ\mu satisfies a small diameter ψα\psi_{\alpha} moment assumption if the ψα\psi_{\alpha} norm replaces the LqL_{q} one in (1.2).

One should note that with very few exceptions, both parts of Assumption 1.3 are needed if one wishes to address Question 1.2.

and c1,c2c_{1},c_{2} are constants that depend only on κ1\kappa_{1}.

It is straightforward to verify that this bound is optimal by considering the uniform measure on the set of coordinate vectors {ne1,...,nen}\{\sqrt{n}e_{1},...,\sqrt{n}e_{n}\}, which results in the coupon-collector problem. Thus, given ε>0\varepsilon>0, one requires at least c(ε)nlog⁡nc(\varepsilon)n\log n random points to ensure that the sample covariance matrix ε\varepsilon-approximates the true covariance. Of course, does not lead to a nontrivial estimate in the second part of Question 1.2, i.e. if the aspect ratio n/N→β∈(0,1]n/N\to\beta\in(0,1] and n→∞n\to\infty, and in particular, (1.3) can not yield a Bai-Yin type of bound. Any hope of getting the desired bounds in Question (1.2) requires additional assumptions on XX.

Unfortunately, when one has a weaker moment estimate than a ψ2\psi_{2} one, the situation becomes considerably more difficult. The complexity of the set one has to control remains the same, but the individual concentration deteriorates, because N^{-1}\sum\bigl{<}X_{i},x\bigr{>}^{2} does not exhibit a strong enough concentration around its mean to balance the concentration-complexity tradeoff at the level of n/N\sqrt{n/N}. Therefore, with a weaker moment assumption than a ψ2\psi_{2} one, a combination of individual tail bounds and a “global” assumption, like the small diameter information, is required in both parts of Question 1.2.

Partial results in the isotropic, log-concave case have been obtain by Bourgain , yielding an estimate on the covariance operator for N=c(ε)nlog⁡3nN=c(\varepsilon)n\log^{3}n, which was improved by Rudelson to N=c(ε)nlog⁡2nN=c(\varepsilon)n\log^{2}n. Subsequent improvements were N=c(ε)nlog⁡nN=c(\varepsilon)n\log n for unconditional convex bodies in and for general log-concave measures in . Finally, the optimal estimate of N=c(ε)nN=c(\varepsilon)n was obtained for an unconditional, log-concave measures by Aubrun , and for an arbitrary log-concave measure in Adamczak et al. , where the following result was proved:

There exist absolute constants c1c_{1} and c2c_{2} for which the following holds. If μ\mu is an isotropic, log-concave measure, then with probability at least 1−exp⁡(−c1n)1-\exp(-c_{1}\sqrt{n}),

Naturally, Question 1.2 becomes even harder when one assumes that linear functionals have heavy tails, because sums of independent random variable exhibit very limited concentration – far below the level required for the proof of Theorem 1.4. Recently, Vershynin proved the following remarkable fact:

For every q>4q>4, δ>0\delta>0 and constants κ1\kappa_{1} and κ2\kappa_{2}, there exist constants c1c_{1} and c2c_{2} that depend on qq, δ\delta and κ1,κ2\kappa_{1},\kappa_{2} for which the following holds.

If μ\mu satisfies a 22-small diameter, LqL_{q} moment assumption with constants κ1\kappa_{1} and κ2\kappa_{2}, then for every δ>0\delta>0, with probability at least 1−δ1-\delta,

In particular, if μ\mu is isotropic then

Moreover, very recently Strivastava and Vershynin , obtained the following result:

If (Xi)i=1N(X_{i})_{i=1}^{N} are independent random vectors distributed according to μ\mu then for every N⩾c1nN\geqslant c_{1}n,

Moreover, only under a qq-moment assumption,

It should be noted that the boundedness assumption in Theorem 1.6 is satisfied by a vector with independent components X=(ξi)i=1nX=(\xi_{i})_{i=1}^{n}, if ξ∈Lq\xi\in L_{q} for q>4q>4, and thus both parts may be used in the i.i.d situation. However, for any η>0\eta>0, c3<12c_{3}<\frac{1}{2} (1/21/2 being the power in the Bai-Yin Theorem).

Our main result gives a version of Theorem 1.5 for an unconditional measure with “heavy tails”.

Theorem A. Let μ\mu be an unconditional measure that satisfies the pp-small diameter, LqL_{q} moment assumption with constants κ1\kappa_{1} and κ2\kappa_{2} for some p>2p>2.

1. For every q>4q>4 and δ<1/2−1/2(p−1)\delta<1/2-1/2(p-1), there exist constants c0c_{0}, c1c_{1} and c2c_{2} that depend on qq, pp, κ1\kappa_{1}, κ2\kappa_{2} and δ\delta, such that, for every n≤N≤exp⁡(c0nδ)n\leq N\leq\exp(c_{0}n^{\delta}), with probability at least 1−exp⁡(−c1nδ)1-\exp(-c_{1}n^{\delta}),

2. For every 2<q≤42<q\leq 4, if p>(1−2/q)−1p>(1-2/q)^{-1} and δ<1/2−1/2(p−1)\delta<1/2-1/2(p-1), there exist constants c3c_{3} and c4c_{4} that depend on qq, pp, δ\delta, κ1\kappa_{1} and κ2\kappa_{2}, such that, for every n≤N≤exp⁡(c0nδ)n\leq N\leq\exp(c_{0}n^{\delta}), with probability at least 1−exp⁡(−c3nδ)1-\exp(-c_{3}n^{\delta}),

In both cases, for every ε>0\varepsilon>0, with probability at least 1−2exp⁡(−cnδ)1-2\exp(-cn^{\delta}), ∥ΣN−Σ∥2→2≤ε\|\Sigma_{N}-\Sigma\|_{2\to 2}\leq\varepsilon provided that N≳q,p,δ,κ1,κ2nN\gtrsim_{q,p,\delta,\kappa_{1},\kappa_{2}}n. Moreover, if μ\mu is isotropic and q>4q>4, then

for an arbitrary class of functions HH – not even for H_{T}=\{\bigl{<}t,\cdot\bigr{>}:t\in T\} when TT is not the sphere or close to the sphere in some sense.

The proof of Theorem A does just that, since it is based on a bound on (1.4) in terms of a certain notion of “complexity” of the class HH. It is not tailored to the case HB2nH_{B_{2}^{n}}, nor does it relay on the fact that the indexing class consists of linear functionals. Rather, the proof is based on a chaining scheme which is much more general than the applications that will be presented here.

The second application we chose to present as an illustration of the potential this empirical processes based method has, is the following.

and thus μ\mu is weakly dominated by GG.

Moreover, the results of show that if TT is centrally symmetric and μ\mu is isotropic and LL-subgaussian, then

Theorem B has many standard applications, leading to embedding results of a similar nature to the Johnson-Lindenstrauss Lemma and to “low M∗M^{*}” estimates that hold for unconditional, log-concave ensembles. Deriving these and other outcomes from Theorem B is standard and will not be presented here. One should also note that a log-concave Chevet type inequality, i.e., upper estimates on the operator norm ∥Γ∥X→Y\|\Gamma\|_{X\to Y} for finite dimensional normed spaces XX and YY has recently been established in .

Preliminaries

Throughout, all absolute constants are positive numbers, denoted by c,c0,c1,...c,c_{0},c_{1},... and their value may change from line to line. κ0,κ1,...\kappa_{0},\kappa_{1},... denote constants whose value will remain unchanged. By A∼BA\sim B we mean that there are absolute constants cc and CC such that cB≤A≤CBcB\leq A\leq CB, and by A≲BA\lesssim B that A≤CBA\leq CB. A∼γBA\sim_{\gamma}B (resp. A≲γBA\lesssim_{\gamma}B) denotes that the constants depend only on γ\gamma.

Next, let us turn to the complexity parameters that motivated our method of analysis – Talagrand’s γ\gamma-functionals.

For a metric space (T,d)(T,d), an admissible sequence of TT is a collection of subsets of TT, {Ts:s≥0}\{T_{s}:s\geq 0\}, such that for every s≥1s\geq 1, ∣Ts∣≤22s|T_{s}|\leq 2^{2^{s}} and ∣T0∣=1|T_{0}|=1. For β≥1\beta\geq 1, define the γβ\gamma_{\beta} functional by

where the infimum is taken with respect to all admissible sequences of TT. For an admissible sequence (Ts)s≥0(T_{s})_{s\geq 0} we denote by πst\pi_{s}t a nearest point to tt in TsT_{s} with respect to the metric dd.

One should note that our chaining approach is based on a slightly less restrictive definition, giving one more freedom; for example, the cardinality of the sets will not necessarily be 22s2^{2^{s}}, the metric may change with ss, etc. (see Section 3).

When considered for a set T⊂L2T\subset L_{2}, γ2\gamma_{2} has close connections with properties of the canonical gaussian process indexed by TT, and we refer the reader to for detailed expositions on these connections. One can show that under mild measurability assumptions, if {Gt:t∈T}\{G_{t}:t\in T\} is a centered gaussian process indexed by a set TT, then

Decomposition of sets

Let ϕ\phi be an increasing function which will be chosen according to additional information one will have on the given class. Examples that one should have in mind are ϕβ(x)∼βxlog⁡1/β(eN/x)\phi_{\beta}(x)\sim_{\beta}\sqrt{x}\log^{1/\beta}(eN/x), resulting from a bound on the ψβ\psi_{\beta} diameter of HH, or ϕq,ε∼q,εN(1+ε)/qx1/2−(1+ε)/q\phi_{q,\varepsilon}\sim_{q,\varepsilon}N^{(1+\varepsilon)/q}x^{1/2-(1+\varepsilon)/q} for q>2q>2 and ε\varepsilon in the right range, arising from an LqL_{q} moment assumption.

1. sup⁡v∈V(θ0(π0v)+∑s>0θs(Δsv))≤γ\sup_{v\in V}\left(\theta_{0}(\pi_{0}v)+\sum_{s>0}\theta_{s}(\Delta_{s}v)\right)\leq\gamma.

2. For every v∈Vv\in V and every I⊂{1,...,N}I\subset\{1,...,N\},

3. If ηs≤N\eta_{s}\leq N then for every v∈Vv\in V and every I⊂{1,...,N}I\subset\{1,...,N\}

and if ηs≥N\eta_{s}\geq N then for every v∈Vv\in V and every I⊂{1,...,N}I\subset\{1,...,N\},

Although this definition seems artificial at first glance, we will show that it captures the geometry of a typical coordinate projection PσH={(h(Xi))i=1N:h∈H}P_{\sigma}H=\{(h(X_{i}))_{i=1}^{N}:h\in H\}.

To formulate the estimate on the Bernoulli process, set

For 2<q≤42<q\leq 4 and 0≤ε<(q/2)−10\leq\varepsilon<(q/2)-1, let

As will become clearer, the most important of the Bq,εB_{q,\varepsilon} parameters is

which, under the standard choice of η0=0\eta_{0}=0 and ηs=2s\eta_{s}=2^{s} for s≥1s\geq 1, corresponds to γ2(V,∥ ∥)\gamma_{2}(V,\|\ \|).

Before presenting the proof, let us consider the two main examples which will interest us, namely, the families ϕβ=xlog⁡1/β(eN/x)\phi_{\beta}=\sqrt{x}\log^{1/\beta}(eN/x) for any β>0\beta>0 and ϕq,ε=x(N/x)(1+ε)/q\phi_{q,\varepsilon}=\sqrt{x}(N/x)^{(1+\varepsilon)/q} for any q>2q>2 (and for ε\varepsilon selected appropriately).

In both cases ϕ(N)∼N\phi(N)\sim\sqrt{N} and for any β>0\beta>0, Φ∼βN\Phi\sim_{\beta}\sqrt{N}. If q>4q>4 and 0≤ε≤q/4−10\leq\varepsilon\leq q/4-1, Φ≤(1−4(1+ε)/q)−1/2N\Phi\leq(1-4(1+\varepsilon)/q)^{-1/2}\sqrt{N}, and since Φs≤Φ\Phi_{s}\leq\Phi, then for β>0\beta>0 or q>4q>4,

with the constant depending either on β\beta or on qq and ε\varepsilon as above.

On the other hand, if 2<q≤42<q\leq 4 and 0<ε<q/2−10<\varepsilon<q/2-1 then

Next, since (ηs)s≥0(\eta_{s})_{s\geq 0} increases exponentially, then for q>2q>2

and the constant in (3.1) depends on β\beta or on qq and ε\varepsilon respectively. In particular, if 2<q≤42<q\leq 4 and 0<ε<q/2−10<\varepsilon<q/2-1, then

Finally, one has to control ∑{s:ηs≤N}ϕ2(ηs)∥Δsv∥\sum_{\{s:\eta_{s}\leq N\}}\phi^{2}(\eta_{s})\|\Delta_{s}v\|. Note that if β>0\beta>0 or q>4q>4, then

and if 2<q≤42<q\leq 4 and 0<ε<q/2−10<\varepsilon<q/2-1 then

We thus arrive to a more compact formulation of Theorem 3.2 in the cases we will be interested in.

For any β>0\beta>0 or q>4q>4, with probability at least 1−2exp⁡(−c1r2η0)1-2\exp(-c_{1}r^{2}\eta_{0}),

with a constant that depends on β\beta or on qq and ε\varepsilon respectively.

Also, if 2<q≤42<q\leq 4 and 0<ε<q/2−10<\varepsilon<q/2-1, then with probability at least 1−2exp⁡(−c1r2η0)1-2\exp(-c_{1}r^{2}\eta_{0}),

Let w⋅v=∑i=1Nwivieiw\cdot v=\sum_{i=1}^{N}w_{i}v_{i}e_{i}, and since

one has to control increments of the form ∑i=1Nεi(Δsv)i(πsv+πs−1v)i\sum_{i=1}^{N}\varepsilon_{i}(\Delta_{s}v)_{i}(\pi_{s}v+\pi_{s-1}v)_{i}.

Observe that if ηs≥N\eta_{s}\geq N then with probability 11,

Next, if ηs≤N\eta_{s}\leq N we will decompose the vectors one has to control according to the size of their coordinates, because, with probability 1−2exp⁡(−r2/2)1-2\exp(-r^{2}/2),

Consider the following two cases. If is,v=ηsi_{s,v}=\eta_{s} then

Therefore, summing the three terms over {s>0:ηs≤N}\{s>0:\eta_{s}\leq N\},

then splitting each w∈Vw\in V to w++w−w^{+}+w^{-} as above,

Recall that ∣ΔsV∣,∣Vs∣≤10⋅2ηs+1|\Delta_{s}V|,|V_{s}|\leq 10\cdot 2^{\eta_{s+1}} and that ηs+1≤10ηs\eta_{s+1}\leq 10\eta_{s}. Given r≥c0r\geq c_{0}, then applying (3) for ts=10rηs1/2t_{s}=10r\eta_{s}^{1/2} and summing over {s:ηs≤N}\{s:\eta_{s}\leq N\}, it follows that sup⁡v∈V∣∑i=1Nεi(v2−(π0v)2)i∣\sup_{v\in V}\left|\sum_{i=1}^{N}\varepsilon_{i}(v^{2}-(\pi_{0}v)^{2})_{i}\right| is bounded by the desired quantity with probability at least 1−2exp⁡(−c1r2η0)1-2\exp(-c_{1}r^{2}\eta_{0}).

Coordinate projections of Function classes

The aim of this section is to show that under very mild assumptions, empirical processes have well behaved coordinate projections in the sense of Definition 3.1. A first result in this direction was established in , in which the main observation, formulated in the language of Section 3, was that if η0=0\eta_{0}=0 and ηs=2s\eta_{s}=2^{s} for s≥1s\geq 1, then for the choice of θs((h(Xi))i=1N)=2s/2∥h∥ψ2\theta_{s}((h(X_{i}))_{i=1}^{N})=2^{s/2}\|h\|_{\psi_{2}}, α=u\alpha=\sqrt{u}, ∥(h(Xi))i=1N∥=∥h∥ψ1\|(h(X_{i}))_{i=1}^{N}\|=\|h\|_{\psi_{1}} and ϕ(x)∼xlog⁡(eN/x)\phi(x)\sim\sqrt{x}\log(eN/x), the set V={(h(Xi))i=1N:h∈H}V=\{(h(X_{i}))_{i=1}^{N}:h\in H\} has a good decomposition with high probability. Hence, the Bernoulli process indexed by V2V^{2} satisfies the following:

There exist absolute constants c1c_{1}, c2c_{2} and c3c_{3} for which the following holds. If HH is a class of functions, then for every r,u≥c1r,u\geq c_{1}, with μN\mu^{N}-probability at least 1−2exp⁡(−c2u)1-2\exp(-c_{2}u), V=PσHV=P_{\sigma}H satisfies that

with probability at least 1−2exp⁡(−c3r2)1-2\exp(-c_{3}r^{2}) with respect to the Bernoulli random variables.

Theorem 4.1 is rather restricted because the ψ2\psi_{2}-based complexity parameter seems too strong in many situations, as does the assumption that HH is a bounded subset of Lψ1L_{\psi_{1}}. Here, we will try to impose as few assumptions as possible on HH.

Let HH be a class of functions on (Ω,μ)(\Omega,\mu). For every u>0u>0 we will define three events in the product space ΩN\Omega^{N}, which will be denoted by Ω1,u\Omega_{1,u}, Ω2,u\Omega_{2,u} and Ω3,u\Omega_{3,u}. On the event Ω1,u∩Ω2,u∩Ω3,u\Omega_{1,u}\cap\Omega_{2,u}\cap\Omega_{3,u}, the random set PσH={(h(Xi))i=1N:h∈H}P_{\sigma}H=\{(h(X_{i}))_{i=1}^{N}:h\in H\} will be well behaved for the right choice of functionals θs\theta_{s} and ϕ\phi. We will then study cases in which the event Ω1,u∩Ω2,u∩Ω3,u\Omega_{1,u}\cap\Omega_{2,u}\cap\Omega_{3,u} has high probability.

For (ηs)s≥0(\eta_{s})_{s\geq 0} as above, set s0≥0s_{0}\geq 0 to be the first integer for which ηs≥log⁡(eN)\eta_{s}\geq\log(eN).

For an admissible sequence (Hs)s≥0(H_{s})_{s\geq 0} and a sequence of functionals (θu,s)s≥s0(\theta_{u,s})_{s\geq s_{0}}, let Ω1,u\Omega_{1,u} be the event for which, for every h∈Hh\in H, the following holds:

2. for every ηs>N\eta_{s}>N, (∑i=1N((Δsh)2(Xi))∗)1/2≤θu,s(Δsh)\left(\sum_{i=1}^{N}((\Delta_{s}h)^{2}(X_{i}))^{*}\right)^{1/2}\leq\theta_{u,s}(\Delta_{s}h).

Formally, to define the set Ω2,u\Omega_{2,u}, first fix a random variable YY, an integer NN and ε>0\varepsilon>0. For every j≤Nj\leq N let δj=(j/eN)(1+ε)\delta_{j}=(j/eN)^{(1+\varepsilon)}, set

and without loss of generality, we will assume that the infimum is attained.

where κ3\kappa_{3} is a suitable chosen absolute constant.

The motivation for this definition is the following observation, showing that with high probability, the “tail” of a sum of i.i.d random variables can be controlled using ff.

Proof. Since Pr(∣Y∣≥yj)≤δj=(j/eN)1+εPr(|Y|\geq y_{j})\leq\delta_{j}=(j/eN)^{1+\varepsilon} then for u≥1u\geq 1,

where the last inequality is evident by a change of variables.

We will also need the following “global” counterpart of the functional ff.

Given a class of functions HH, an integer NN and ε>0\varepsilon>0, set

Clearly, for every h∈Hh\in H and every kk, fu(h,k)≤Fu(k)f_{u}(h,k)\leq F_{u}(k).

The final set, Ω3,u\Omega_{3,u} is very close in nature to Ω2,u\Omega_{2,u}. It is needed to control the coordinates of “very small” increments – when s<s0s<s_{0}, if such an integer exists.

If η0<log⁡(eN)\eta_{0}<\log(eN), let Ω3,u\Omega_{3,u} be the event on which for every h∈Hh\in H, every 0≤s<s00\leq s<s_{0} and 1≤j≤N1\leq j\leq N,

If η0≥log⁡(eN)\eta_{0}\geq\log(eN) set Ω3,u=ΩN\Omega_{3,u}=\Omega^{N}.

It turns out that on the event Ω1,u∩Ω2,u∩Ω3,u\Omega_{1,u}\cap\Omega_{2,u}\cap\Omega_{3,u}, the set PσHP_{\sigma}H is indeed well behaved. Let

with the infimum is taken with respect to all (ηs)(\eta_{s})-admissible sequences. From here on we will assume that (Hs)s≥0(H_{s})_{s\geq 0} is an almost optimal (ηs)s≥0(\eta_{s})_{s\geq 0}-admissible sequence.

There exists absolute constants c1c_{1} and c2c_{2} for which the following holds. Let (θu,s)s≥s0(\theta_{u,s})_{s\geq s_{0}} be functionals, and for s<s0s<s_{0} set θu,s=0\theta_{u,s}=0. For every u≥c1u\geq c_{1}, on the event Ω1,u∩Ω2,u∩Ω3,u\Omega_{1,u}\cap\Omega_{2,u}\cap\Omega_{3,u}, for every h∈Hh\in H and I⊂{1,...,N}I\subset\{1,...,N\},

where Rs0,I(h)=θu,0(π0h)R_{s_{0},I}(h)=\theta_{u,0}(\pi_{0}h) if s0=0s_{0}=0 and Rs0(h,I)=min⁡{θu,s0(πs0h),Fu(∣I∣)}R_{s_{0}}(h,I)=\min\{\theta_{u,s_{0}}(\pi_{s_{0}}h),F_{u}(|I|)\} otherwise.

and the claim is evident from the definition of the function fuf_{u} and the set Ω2,u\Omega_{2,u}.

If, on the other hand, ηs<log⁡(eN)\eta_{s}<\log(eN) then s0>0s_{0}>0 and the assertion follows from the definition of Ω3,u\Omega_{3,u}.

The second part of (1) follows from the definition of Ω1,u\Omega_{1,u}.

and thus, for every h∈Hh\in H and every I⊂{1,...,N}I\subset\{1,...,N\},

For Lemma 4.8 to have any meaning, one has to identify the functionals fuf_{u}, FuF_{u} and θu,s\theta_{u,s} in the cases one is interested in. Our next goal is to study the functions fuf_{u} and FuF_{u} under various tail assumptions on functions in HH, and naturally, the two families of tail estimates we will be interested in are when HH has a bounded diameter in LψβL_{\psi_{\beta}} or in LqL_{q} for q>2q>2.

If H⊂LψβH\subset L_{\psi_{\beta}}, then for every h∈Hh\in H, Pr(∣h∣≥y)≤exp⁡(−(y/∥h∥ψβ)β)Pr(|h|\geq y)\leq\exp(-(y/\|h\|_{\psi_{\beta}})^{\beta}). Thus, for ε≥1\varepsilon\geq 1 and every jj,

Hence, if dψβ=sup⁡h∈H∥h∥ψβd_{\psi_{\beta}}=\sup_{h\in H}\|h\|_{\psi_{\beta}}, then

Using the same argument, if h∈Lqh\in L_{q} then Pr(∣h∣≥∥h∥Lqy)≤1/yqPr(|h|\geq\|h\|_{L_{q}}y)\leq 1/y^{q} and for any 0<ε<q/2−10<\varepsilon<q/2-1, yj=∥h∥Lq(N/j)(1+ε)/qy_{j}=\|h\|_{L_{q}}(N/j)^{(1+\varepsilon)/q}. If sup⁡h∈H∥h∥Lq=dLq\sup_{h\in H}\|h\|_{L_{q}}=d_{L_{q}}, q>2q>2 and cq,ε=1−2(1+ε)/qc_{q,\varepsilon}=1-2(1+\varepsilon)/q then

Combining these observations with the estimates of Lemma 4.8 and noting that if s0>0s_{0}>0 then Rs0,I(h)≲udLqϕq,ε(∣I∣)R_{s_{0},I}(h)\lesssim\sqrt{u}d_{L_{q}}\phi_{q,\varepsilon}(|I|), one reaches the following corollary.

Let (θu,s)s≥s0(\theta_{u,s})_{s\geq s_{0}} be a sequence of functionals and for s<s0s<s_{0} let θu,s=0\theta_{u,s}=0. If HH is bounded in LqL_{q} for q>2q>2, then on Ω1,u∩Ω2,u∩Ω3,u\Omega_{1,u}\cap\Omega_{2,u}\cap\Omega_{3,u}, for every h∈Hh\in H and every I⊂{1,...,N}I\subset\{1,...,N\}

A similar bound holds when HH is bounded in LψβL_{\psi_{\beta}}.

We will begin by showing that Ω2,u\Omega_{2,u} is a large set, almost regardless of any assumptions on ϕ\phi, an observation that is based on the same idea as Lemma 4.4.

There exist absolute constants c1c_{1} and c2c_{2} such that, for every ε>0\varepsilon>0 and u≥c1/εu\geq c_{1}/\varepsilon, Pr(Ω2,u)≥1−2exp⁡(−c2εuηs0)Pr(\Omega_{2,u})\geq 1-2\exp(-c_{2}\varepsilon u\eta_{s_{0}}).

Since Ω2,u\Omega_{2,u} is always large, and since Ω3,u\Omega_{3,u} will behave in a very similar way when s0>0s_{0}>0, the crucial point in the construction of a good decomposition of PσHP_{\sigma}H is a correct choice of θu,s\theta_{u,s} and estimates on Ω1,u\Omega_{1,u}.

The functionals θu,s\theta_{u,s} capture the geometry of HH, and thus have to be selected according to the information one has on the class. We will present two examples of such choices, each leading to one of our two main results. The first one will be based on “global” structure like metric entropy, while the second uses accurate estimates on each “chain”.

Let κ4≥10\kappa_{4}\geq 10 be an absolute constant to be fixed later, set 2s1∼nδ2^{s_{1}}\sim n^{\delta} for δ<1/2−1/2(p−1)\delta<1/2-1/2(p-1), and put

Note that s0=0s_{0}=0 as long as η0∼2s1log⁡(en/2s1)≥log⁡(eN)\eta_{0}\sim 2^{s_{1}}\log(en/2^{s_{1}})\geq\log(eN), i.e., if nδlog⁡(n)≳log⁡(eN)n^{\delta}\log(n)\gtrsim\log(eN) - which we will assume is the case, since our main interest in when N∼nN\sim n.

For every κ1\kappa_{1}, p>2p>2 and δ<1/2−1/2(p−1)\delta<1/2-1/2(p-1) there exist constants c1,c2c_{1},c_{2} and c3c_{3} that depend only on κ1\kappa_{1}, pp and δ\delta for which the following holds. There is an (ηs)s≥0(\eta_{s})_{s\geq 0}-admissible sequence of B2nB_{2}^{n}, for which, if u≥c1u\geq c_{1}, then Pr(Ω1,u)≥1−exp⁡(−c2nδ)Pr(\Omega_{1,u})\geq 1-\exp(-c_{2}n^{\delta}) and

Indeed, by the unconditionality of μ\mu, (x1,...,xn)(x_{1},...,x_{n}) has the same distribution as (ε1x1,...,εnxn)(\varepsilon_{1}x_{1},...,\varepsilon_{n}x_{n}). Hence, for every r≥1r\geq 1

If I⊂{1,...,n}I\subset\{1,...,n\} then for every ε>0\varepsilon>0, log⁡N(B2I,εBψ2)≲M∣I∣2/ε2\log N(B_{2}^{I},\varepsilon B_{\psi_{2}})\lesssim M_{|I|}^{2}/\varepsilon^{2}. Moreover, for ε≤1\varepsilon\leq 1, log⁡N(B2n,εBψ2)≲nlog⁡(2/ε)\log N(B_{2}^{n},\varepsilon B_{\psi_{2}})\lesssim n\log(2/\varepsilon).

for a suitable absolute constant cc, proving the first part.

For the second part, note that N(B2n,εBψ2)≤N(B2n,Bψ2)⋅N(Bψ2,εBψ2)N(B_{2}^{n},\varepsilon B_{\psi_{2}})\leq N(B_{2}^{n},B_{\psi_{2}})\cdot N(B_{\psi_{2}},\varepsilon B_{\psi_{2}}). By the first part, log⁡N(B2n,Bψ2)≲n\log N(B_{2}^{n},B_{\psi_{2}})\lesssim n, while a standard volumetric estimate shows that N(Bψ2,εBψ2)≤(5/ε)nN(B_{\psi_{2}},\varepsilon B_{\psi_{2}})\leq(5/\varepsilon)^{n}.

Next, let us define the sets TsT_{s}. If 2s+s1>n2^{s+s_{1}}>n, let TsT_{s} be a maximal εs\varepsilon_{s} separated subset of B2nB_{2}^{n} relative to the ψ2\psi_{2} norm and of cardinality 2ηs2^{\eta_{s}}. If 2s+s1≤n2^{s+s_{1}}\leq n, let TsT_{s} be a maximal εs\varepsilon_{s} separated subset of U2s+s1={x∈B2n:∣supp(x)∣≤2s+s1}U_{2^{s+s_{1}}}=\{x\in B_{2}^{n}:|{\rm supp}(x)|\leq 2^{s+s_{1}}\} with respect to the ψ2\psi_{2} norm, and of cardinality 2ηs2^{\eta_{s}}. Given a vector t∈B2nt\in B_{2}^{n}, we will define the functions πs\pi_{s} as follows. If 2s+s1>n2^{s+s_{1}}>n, πst\pi_{s}t is a best ψ2\psi_{2} approximation of tt in TsT_{s}. For 2s+s1≤n2^{s+s_{1}}\leq n one combines approximation and dimension reduction. Set s∗s_{*} to satisfy that 2s∗+s1=n2^{s_{*}+s_{1}}=n (and without loss of generality we will assume that such an integer exists). If v=πs∗tv=\pi_{s_{*}}t, let In/2I_{n/2} be the set of the largest n/2n/2 coordinates of vv, and put πs∗−1t\pi_{s_{*}-1}t to be the best approximation of the coordinate projection PIn/2vP_{I_{n/2}}v in Ts∗−1T_{s_{*}-1}, and so on.

There exists an absolute constant cc such that for every t∈B2nt\in B_{2}^{n}, if s>s∗s>s_{*} (i.e., if ηs>κ3n\eta_{s}>\kappa_{3}n), then \|\bigl{<}\Delta_{s}t,\cdot\bigr{>}\|_{\psi_{2}}\leq c2^{-2^{s+s_{1}}/n}, and if 0<s≤s∗0<s\leq s_{*} then \|\bigl{<}\Delta_{s}t,\cdot\bigr{>}\|_{\psi_{2}}\leq c2^{-(s+s_{1})/2}M_{2^{s+s_{1}}}.

Proof. First consider s>s∗s>s_{*}. Note that \|\bigl{<}\Delta_{s}t,\cdot\bigr{>}\|_{\psi_{2}}\leq\|\bigl{<}t-\pi_{s}t,\cdot\bigr{>}\|_{\psi_{2}}+\|\bigl{<}t-\pi_{s-1}t,\cdot\bigr{>}\|_{\psi_{2}}\leq\varepsilon_{s}+\varepsilon_{s-1}, and by the covering numbers estimate from Lemma 5.3, in that range εs≲2−2s+s1/n\varepsilon_{s}\lesssim 2^{-2^{s+s_{1}}/n}.

In the range s≤s∗s\leq s_{*}, Δst=u+w\Delta_{s}t=u+w, where ww consists of the smallest 2s+s1−12^{s+s_{1}-1} coordinates of πst∈B2I\pi_{s}t\in B_{2}^{I} for some ∣I∣=2s+s1|I|=2^{s+s_{1}}, and uu is an εs−1\varepsilon_{s-1}-approximation of the largest 2s+s1−12^{s+s_{1}-1} coordinates of πst\pi_{s}t. Therefore, \|\bigl{<}\Delta_{s}t,\cdot\bigr{>}\|_{\psi_{2}}\leq\|\bigl{<}w,\cdot\bigr{>}\|_{\psi_{2}}+\varepsilon_{s-1}. Recall that for every such ss, U2s+s1U_{2^{s+s_{1}}} is a union of (n2s+s1)\binom{n}{2^{s+s_{1}}} balls of dimension 2s+s12^{s+s_{1}}, then

Proof of Theorem 5.2. Observe that ∥Y∥ψ22=∥Y2∥ψ1\|Y\|_{\psi_{2}}^{2}=\|Y^{2}\|_{\psi_{1}}, and thus, by a standard application of Bernstein’s inequality, for every integer mm,

Also, with probability at least 1−2exp⁡(−c2w2ηs)1-2\exp(-c_{2}w^{2}\eta_{s}), if ηs≥N\eta_{s}\geq N then

Using Lemma 5.4 and summing the probability estimates, it is evident that with probability at least 1−2exp⁡(−c3w2η0)1-2\exp(-c_{3}w^{2}\eta_{0}), the following holds: if ηs≥N\eta_{s}\geq N then

and if s>0s>0 and ηs≤κ4n\eta_{s}\leq\kappa_{4}n then

Note that if 2s1∼nδ2^{s_{1}}\sim n^{\delta} for δ<1/2−1/2(p−1)\delta<1/2-1/2(p-1), then

There exist absolute constants c1c_{1}, c2c_{2} and c3c_{3} and c4c_{4} that depend on κ1,κ2,p,δ\kappa_{1},\kappa_{2},p,\delta, for which the following holds. If μ\mu is as above and ε>0\varepsilon>0, then B2nB_{2}^{n} has an (ηs)s≥0(\eta_{s})_{s\geq 0}-admissible sequence (Ts)s≥0(T_{s})_{s\geq 0} for which, for u≥c1/εu\geq c_{1}/\varepsilon, with probability at least 1−2exp⁡(−c2εun)−2exp⁡(−c3nδ)1-2\exp(-c_{2}\varepsilon un)-2\exp(-c_{3}n^{\delta}), for every t∈B2nt\in B_{2}^{n} and every I⊂{1,...,N}I\subset\{1,...,N\},

For every κ1\kappa_{1}, κ2\kappa_{2}, q>4q>4, p>2p>2 and δ<1/2−1/2(p−1)\delta<1/2-1/2(p-1), there exist constants c0c_{0}, c1c_{1}, c2c_{2} and c3c_{3} which depend on κ1\kappa_{1}, κ2\kappa_{2}, pp, qq and δ\delta, and an absolute constant c4c_{4} for which the following holds. If μ\mu is as above, and N≤exp⁡(c0nδ)N\leq\exp(c_{0}n^{\delta}), then for every u≥c1u\geq c_{1}, with μN\mu^{N}-probability at least 1−2exp⁡(−c2nδ)1-2\exp(-c_{2}n^{\delta}), Pσ(B2n)P_{\sigma}(B_{2}^{n}) satisfies that

with probability at least 1−2exp⁡(−c4nr2)1-2\exp(-c_{4}nr^{2}) relative to the Bernoulli random variables.

Turning to the case 2<q≤42<q\leq 4, recall that for 0<ε<q/2−10<\varepsilon<q/2-1, Bq,ε=∑{s:ηs≤N}ηs1−2(1+ε)/q∥Δsv∥LqB_{q,\varepsilon}=\sum_{\{s:\eta_{s}\leq N\}}\eta_{s}^{1-2(1+\varepsilon)/q}\|\Delta_{s}v\|_{L_{q}}. Assume that μ\mu is as above and satisfies the pp-small diameter assumption for p>q/(q/2−1)p>q/(q/2-1). Then, for 0<ε<q/2−1−q/p0<\varepsilon<q/2-1-q/p (i.e. if 1−(2(1+ε)/q)−1/p>01-(2(1+\varepsilon)/q)-1/p>0),

Let 2<q≤42<q\leq 4, p>(1−2/q)−1p>(1-2/q)^{-1} and 0<ε<2/q−1−q/p0<\varepsilon<2/q-1-q/p. If μ\mu and δ\delta are as above, u≳1/εu\gtrsim 1/\varepsilon, and N≲exp⁡(c0nδ)N\lesssim\exp(c_{0}n^{\delta}), then with μN\mu^{N} probability at least 1−2exp⁡(−c1εun)−2exp⁡(−c2nδ)1-2\exp(-c_{1}\varepsilon un)-2\exp(-c_{2}n^{\delta}), Pσ(B2n)P_{\sigma}(B_{2}^{n}) satisfies that

with probability at least 1−2exp⁡(−c3nr2)1-2\exp(-c_{3}nr^{2}) relative to the Bernoulli random variables.

In particular, taking ε∼1/log⁡(eN/n)\varepsilon\sim 1/\log(eN/n), then for every such NN satisfying that N≳κ1,κ2,q,pnN\gtrsim_{\kappa_{1},\kappa_{2},q,p}n, and any u≥κ1,κ2,q,plog⁡(eN/n)u\geq_{\kappa_{1},\kappa_{2},q,p}\log(eN/n),

2 Unconditional log-concave measures

We will now present a different way of bounding Ω1,u\Omega_{1,u} (and Ω3,u\Omega_{3,u} if needed) by estimating the moments of the increments Δsh\Delta_{s}h, and selecting the functionals θu,s\theta_{u,s} accordingly.

In light of Theorem B, we will assume that HH is a bounded subset of Lψ1L_{\psi_{1}} (although what we do here can be extended to other moment assumptions), and thus one may control Ω2,u\Omega_{2,u} using ϕβ\phi_{\beta} for β=1\beta=1 and ε\varepsilon which will be selected later.

There exist absolute constants c1c_{1}, c2c_{2} and c3c_{3} for which the following holds. For u≥c1u\geq c_{1}, with probability at least 1−2exp⁡(−c1uηs0)1-2\exp(-c_{1}u\eta_{s_{0}}), for every s≥s0s\geq s_{0} and every h∈Hh\in H, Zs(h)≤e∥Zs(h)∥L2uηs+1Z_{s}(h)\leq e\|Z_{s}(h)\|_{L_{2u\eta_{s+1}}}.

Next, one has to control the moments appearing in Lemma 5.8, which is based on the following result, due to Latała .

Let X1,...,XmX_{1},...,X_{m} be independent, distributed according to a nonnegative random variable XX. Then for every p≥1p\geq 1,

If XX is a random variable, for every p≥1p\geq 1 set

The (p)(p)-norms are a local version of the ψ2\psi_{2} norm, and clearly ∥X∥(p)≲∥X∥ψ2\|X\|_{(p)}\lesssim\|X\|_{\psi_{2}}. Using those norms one may obtain a more compact expression for the required moments.

There exist an absolute constant cc such that for every h∈Hh\in H, every s>s0s>s_{0} and every u>0u>0,

There exist absolute constants c1c_{1} and c2c_{2} for which the following holds. If, for s>s0s>s_{0},

then Pr(Ω1,u)≥1−2exp⁡(−c2uηs0)Pr(\Omega_{1,u})\geq 1-2\exp(-c_{2}u\eta_{s_{0}}).

Next, assume that s0>0s_{0}>0, and thus one has to bound Pr(Ω3,u)Pr(\Omega_{3,u}).

There exists absolute constants c1c_{1}, c2c_{2} and c3c_{3} such that, for every u≥c1u\geq c_{1}, with probability at least 1−2exp⁡(−c2ulog⁡N)1-2\exp(-c_{2}u\log N), for every 0≤s<s00\leq s<s_{0} and every h∈Hh\in H,

and a similar bound holds for πs0h\pi_{s_{0}}h.

Proof. Recall that for a fixed ε>0\varepsilon>0 and every ii, Pr(Yi∗≥yi)≤exp⁡(−εilog⁡(eN/i))Pr(Y_{i}^{*}\geq y_{i})\leq\exp(-\varepsilon i\log(eN/i)). Let ε∼u≥1\varepsilon\sim u\geq 1 and observe that if Y∈Lψ1Y\in L_{\psi_{1}}, then yi≲u∥Y∥ψ1log⁡(eN/i)y_{i}\lesssim u\|Y\|_{\psi_{1}}\log(eN/i) and

Since the cardinality of the set ∪s<s0ΔsH\cup_{s<s_{0}}\Delta_{s}H is at most ∑s<s02ηs+1≲Nc2\sum_{s<s_{0}}2^{\eta_{s+1}}\lesssim N^{c_{2}}, (5.6) holds uniformly with probability at least 1−exp⁡(−c3ulog⁡N)1-\exp(-c_{3}u\log N) for u≥c4u\geq c_{4}. Therefore, on that event, for every 0≤s<s00\leq s<s_{0} and every jj,

An identical argument holds for (∑i=1j((πs0h)2(Xi))∗)1/2\left(\sum_{i=1}^{j}((\pi_{s_{0}}h)^{2}(X_{i}))^{*}\right)^{1/2}.

Therefore, the event Ω1,u∪Ω2,u∪Ω3,u\Omega_{1,u}\cup\Omega_{2,u}\cup\Omega_{3,u} has high probability, leading to the following decomposition result.

There exist absolute constants c1c_{1} and c2c_{2} for which the following holds. For every u≥c1u\geq c_{1}, with probability at least 1−2exp⁡(−c2ulog⁡N)1-2\exp(-c_{2}u\log N), for every h∈Hh\in H and every I⊂{1,...,N}I\subset\{1,...,N\},

where Rs0(h)≲uη1∥π0h∥(2uη1)R_{s_{0}}(h)\lesssim\sqrt{u}\eta_{1}\|\pi_{0}h\|_{(2u\eta_{1})} if s0=0s_{0}=0 and Rs0(h)≲udψ1ϕ1(∣I∣)R_{s_{0}}(h)\lesssim ud_{\psi_{1}}\phi_{1}(|I|) otherwise.

Note that ∥Δsh∥(2uηs+1)≲∥Δsh∥ψ2\|\Delta_{s}h\|_{(2u\eta_{s+1})}\lesssim\|\Delta_{s}h\|_{\psi_{2}}, and thus one may take θu,s∼uηs+11/2∥Δsh∥ψ2\theta_{u,s}\sim\sqrt{u}\eta^{1/2}_{s+1}\|\Delta_{s}h\|_{\psi_{2}}. If η0=0\eta_{0}=0 and ηs=2s\eta_{s}=2^{s} for s≥1s\geq 1, then for an almost optimal admissible sequence,

Although this estimate leads to an alternative proof of Theorem 4.1, it is not sharp enough to prove Theorem B, as the latter requires more accurate bounds on ∥Δsh∥(2uηs+1)\|\Delta_{s}h\|_{(2u\eta_{s+1})}.

From here on we will assume that η0=0\eta_{0}=0 and that ηs=2s\eta_{s}=2^{s} for s≥1s\geq 1. If s≥s0∼log⁡Ns\geq s_{0}\sim\log N, set θu,s(Δst)=uηs+11/2∥Δsh∥(2uηs+1)\theta_{u,s}(\Delta_{s}t)=\sqrt{u}\eta_{s+1}^{1/2}\|\Delta_{s}h\|_{(2u\eta_{s+1})}.

There exist absolute constants c1c_{1} and c2c_{2} for which the following holds. If μ\mu is an isotropic, unconditional log-concave measure, H_{T}=\{\bigl{<}t,\cdot\bigr{>}:t\in T\} and (Ts)s≥0(T_{s})_{s\geq 0} is an admissible sequence of TT, then for every u≥c1u\geq c_{1},

where K1K_{1} is an isotropic image of B1nB_{1}^{n}.

The moments of every linear functional \bigl{<}t,\cdot\bigr{>} relative to the volume measure of an isotropic position of B1nB_{1}^{n} are well known : namely, for 1≤p≤n1\leq p\leq n,

Note that for an almost optimal admissible sequence,

There exist absolute constants c1c_{1}, c2c_{2}, c3c_{3} and c4c_{4} for which the following holds. For every u≥c1u\geq c_{1}, With μN\mu^{N}-probability at least 1−2exp⁡(−c2ulog⁡N)1-2\exp(-c_{2}u\log N), the set V=PσTV=P_{\sigma}T satisfies that

with probability at least 1−2exp⁡(−c4r2)1-2\exp(-c_{4}r^{2}) with respect to the Bernoulli random variables.

3 Proofs of Theorems A and B

The final step we need for the proofs of Theorem A and Theorem B is a version of the Ginè-Zinn symmetrization Theorem (see, e.g. ), which enables one to pass from the Bernoulli process indexed by random coordinate projections of a class of functions, to the empirical process indexed by the class.

Let HH be a class of functions which is bounded in LqL_{q} and consider the empirical process indexed by F={h2:h∈H}F=\{h^{2}:h\in H\}. If q≥4q\geq 4 and x≳dLq2Nx\gtrsim d_{L_{q}}^{2}\sqrt{N} then βN(x)≥1/2\beta_{N}(x)\geq 1/2 and the same holds if 2<q<42<q<4 and x≳qdLq2N2/qx\gtrsim_{q}d_{L_{q}}^{2}N^{2/q}.

Proof. The first part of the claim follows from an application of Chebyshev’s inequality, and is omitted. For the second part, fix r>0r>0, set V=(h2(Xi))i=1NV=(h^{2}(X_{i}))_{i=1}^{N}, and since Pr(∣h2(X)∣≥dLq2(rN/i)2/q)≤i/(rN)Pr(|h^{2}(X)|\geq d_{L_{q}}^{2}(rN/i)^{2/q})\leq i/(rN) then

Moreover, for r0∼c2r_{0}\sim c_{2}, Pr(∃i:∣h2(Xi)∣≥r0dLq2N2/q)≤1/10Pr(\exists i:|h^{2}(X_{i})|\geq r_{0}d_{L_{q}}^{2}N^{2/q})\leq 1/10. Hence, a truncation argument shows that without loss of generality we may assume that ∥h2∥L∞≤r0dLq2N2/q\|h^{2}\|_{L_{\infty}}\leq r_{0}d_{L_{q}}^{2}N^{2/q}. Applying the L∞L_{\infty} estimate for the largest two coordinates of VV and (5.7) for the rest, it follows that

showing that it suffices to take x∼qdLq2N2/qx\sim_{q}d_{L_{q}}^{2}N^{2/q} as claimed.

Since x/Nx/N is well within our range, one may complete the proofs of Theorem A and Theorem B.

Proof of Theorem A. For q>4q>4, let ρr,u∼κ1,κ2,qru(n/N+n/N)\rho_{r,u}\sim_{\kappa_{1},\kappa_{2},q}ru(\sqrt{n/N}+n/N) for u≳κ1,κ2,qc1u\gtrsim_{\kappa_{1},\kappa_{2},q}c_{1}, and r≥c2r\geq c_{2}. If 2<q<42<q<4 set ρr,u∼κ1,κ2,qru(n/N)1−2/q\rho_{r,u}\sim_{\kappa_{1},\kappa_{2},q}ru(n/N)^{1-2/q} for u≳κ1,κ2,qlog⁡(eN/n)u\gtrsim_{\kappa_{1},\kappa_{2},q}\log(eN/n) and r≥c3r\geq c_{3}. Then,

Proof of the quantitative Bai-Yin Theorem.

Hence, the final step in the proof of our version of the Bai-Yin Theorem is to show that if ξ∈Lq\xi\in L_{q} for q>2q>2, there is some p>2p>2 for which A{\cal A} has a large measure.

For every q>4q>4 and 2<p<q2<p<q, there exist constants c1c_{1} and c2c_{2} that depend on qq and pp for which the following holds. If ξ∈Lq\xi\in L_{q}, X=(ξ1,...,ξn)X=(\xi_{1},...,\xi_{n}) and X1,...,XNX_{1},...,X_{N} are independent copies of XX, then

Proof. If A=∥ξ∥LqA=\|\xi\|_{L_{q}} then Pr(∣ξ∣≥At)≤t−qPr(|\xi|\geq At)\leq t^{-q}, and for every 1≤k≤n1\leq k\leq n, Pr(ξk∗≥t)≤(nk)(Pr(∣ξ∣≥t))kPr(\xi_{k}^{*}\geq t)\leq\binom{n}{k}(Pr(|\xi|\geq t))^{k}. Therefore, if p<qp<q and y>ey>e then

Combining Lemma 5.21 with Theorem A concludes the proof of the quantitative Bai-Yin Theorem.

Proof of Theorem B. If r∼ur\sim u, with probability at least 1−2exp⁡(−c3u2)1-2\exp(-c_{3}u^{2}) with respect to the Bernoulli random variables,

Since d2(T)E(T)/Nd_{2}(T)E(T)/\sqrt{N} is a “legal” choice in the Giné-Zinn symmetrization theorem, the proof is concluded.

References