Entropy and the fourth moment phenomenon
Ivan Nourdin, Giovanni Peccati, Yvik Swan
Introduction
The study of the classical CLT for sums of independent random elements by entropic methods dates back to Linnik’s seminal paper . Among the many fundamental contributions to this line of research, we cite (see the monograph for more details on the history of the theory). All these influential works revolve around a deep analysis of the effect of analytic convolution on the creation of entropy: in this respect, a particularly powerful tool are the ‘entropy jump inequalities’ proved and exploited e.g. in . As discussed e.g. in , entropy jump inequalities are directly connected with challenging open questions in convex geometry, like for instance the Hyperplane and KLS conjectures. One of the common traits of all the above references is that they develop tools to control the Fisher information and use the aforementioned de Bruijn’s formula to translate the bounds so obtained into bounds on the relative entropy.
One of the main motivations of the present paper is to initiate a systematic information-theoretical analysis of a large class of CLTs that has emerged in recent years in connection with different branches of modern stochastic analysis. These limit theorems typically involve: (a) an underlying infinite dimensional Gaussian field (like for instance a Wiener process), (b) a sequence of rescaled centered random vectors , , having the form of some highly non-linear functional of the field . For example, each may be defined as some collection of polynomial transformations of , possibly depending on a parameter that is integrated with respect to a deterministic measure (but much more general forms are possible). Objects of this type naturally appear e.g. in the high-frequency analysis of random fields on homogeneous spaces , fractional processes , Gaussian polymers , or random matrices .
In view of their intricate structure, it is in general not possible to meaningfully represent the vectors in terms of some linear transformation of independent (or weakly dependent) vectors, so that the usual analytical techniques based on stochastic independence and convolution (or mixing) cannot be applied. To overcome these difficulties, a recently developed line of research (see for an introduction) has revealed that, by using tools from infinite-dimensional Gaussian analysis (e.g. the so-called Malliavin calculus of variations – see ) and under some regularity assumptions on , one can control the distance between the distribution of and that of some Gaussian target by means of quantities that are no more complex than the fourth moment of . The regularity assumptions on are usually expressed in terms of the projections of each on the eigenspaces of the Ornstein-Uhlenbeck semigroup associated with .
The estimate (1.1) is obtained by combining the Malliavin calculus of variations with the Stein’s method for normal approximations . Stein’s method can be roughly described as a collection of analytical techniques, allowing one to measure the distance between random elements by controlling the regularity of the solutions to some specific ordinary (in dimension 1) or partial (in higher dimensions) differential equations. The needed estimates are often expressed in terms of the same Stein factors that lie at the core of the present paper (see Section 2.3 for definitions). It is important to notice that the strength of these techniques significantly breaks down when dealing with normal approximations in dimension strictly greater than .
For instance, in view of the structure of the associated PDEs, for the time being there is no way to directly use Stein’s method in order to deduce bounds in the multidimensional total variation distance. (This fact is demonstrated e.g. in references , where multidimensional bounds are obtained for distances involving smooth test functions, as well as in , containing a quantitative version of the classical multidimensional CLT, with bounds on distances involving test functions that are indicators of convex sets: all these papers use some variations of the multidimensional Stein’s method, but none of them achieves bounds in the total variation distance.)
where stands for the Euclidean norm, and assume that , as . Then,
where stands for a bounded numerical sequence, depending on and on the sequence .
As in the one-dimensional case, one has always that for as in the previous statement. The quantity of the left-hand-side of (1.2) equals of course the relative entropy of . In view of the Csiszar-Kullback-Pinsker inequality (see ), according to which
relation (1.2) then translates in a bound on the square of the total variation distance between and , where the dependence in hinges on the order of the chaoses only via a multiplicative constant. This bound agrees up to a logarithmic factor with the estimates in smoother distances established in (see also ), where it is proved that there exists a constant such that
where stands for the usual Wasserstein distance of order 1. Relation (1.2) also drastically improves the bounds that can be deduced from , yielding that, as ,
where is any strictly positive number verifying , and the symbol stands again for some bounded numerical sequence. The estimate (1.2) seems to be largely outside the scope of any other available technique. Our results will also show that convergence in relative entropy is a necessary and sufficient condition for CLTs involving random vectors whose components live in a fixed Wiener chaos. As in , an important tool for establishing our main results is the Carbery-Wright inequality , providing estimates on the small ball probabilities associated with polynomial transformations of Gaussian vectors. Observe also that, via the Talagrand’s transport inequality , our bounds trivially provide estimates on the 2-Wasserstein distance between and , for every . Notice once again that, in Theorem 1.1 and its generalisations, no additional regularity (a part from the fact of being elements of a fixed Wiener chaos) is required from the components of the vector . One should contrast this situation with the recent work by Hu, Lu and Nualart , where the authors achieve fourth moment bounds on the supremum norm of the difference , under very strong additional conditions expressed in terms of the finiteness of negative moments of Malliavin matrices.
We stress that, although our principal motivation comes from asymptotic problems on a Gaussian space, the methods developed in Section 2 are general. In fact, at the heart of the present work lie the powerful equivalences (2.45)– (2.46) (which can be considered as a new form of so-called Stein identities) that are valid under very weak assumptions on the target density; it is also easy to uncover a wide variety of extensions and generalizations so that we expect that our tools can be adapted to deal with a much wider class of multidimensional distributions.
The connection between Stein identities and information theory has already been noted in the literature (although only in dimension 1). For instance, explicit applications are known in the context of Poisson and compound Poisson approximations , and recently several promising identities have been discovered for some discrete as well as continuous distributions . However, with the exception of , the existing literature seems to be silent about any connection between entropic CLTs and Stein’s identities for normal approximations. To the best of our knowledge, together with (which however focusses on bounds of a completely different nature) the present paper contains the first relevant study of the relations between the two topics.
In order to simplify the discussion, we shall sometimes use the shorthand notation
To enhance the readability of the text, the next Subsection 1.2 contains an intuitive description of our method in dimension one.
2 Illustration of the method in dimension one
Recall also that, in view of the Pinsker-Csiszar-Kullback inequality, one has that
for every smooth test function . Specifying implies, in particular, that . It is easily seen that, under standard regularity assumptions, a version of is given by , for in the support of (in particular, the Stein factor of is 1). The relevance of the factor in comparing with is actually revealed by the following Stein’s bound , which is one of the staples of Stein’s method:
providing a formal meaning to the intuitive fact that the distributions of and are close whenever is close to , that is, whenever is close to 1. To motivate the reader, we shall now present a simple illustration of how the estimate (1.7) applies to the usual CLT.
Let be a sequence of i.i.d. copies of , set and assume that (a simple sufficient condition for this to hold is e.g. that has compact support, and is bounded from below inside its support). Then, using e.g. [51, Lemma 2],
Since (by definition) for all we get
In particular, writing for the density of , we deduce from (1.7) that
We shall demonstrate in Section 3 and Section 4 that the quantity (as well as its multidimensional generalisations) can be explicitly controlled whenever is a smooth functional of a Gaussian field. In view of these observations, the following question is therefore natural: can one bound by an expression analogous to the right-hand-side of (1.7)?
for every smooth test function . We also write, for ,
for the Fisher information of , and we observe that
where is the so-called standardised Fisher information of (note that ). With this notation in mind, de Bruijn’s formula (in an integral and rescaled version due to Barron ) reads
(see Lemma 2.3 below for a multidimensional statement).
Using the standard relation (see e.g. [21, Lemma 1.21]), we deduce the upper bound
a result which is often proved by using entropy power inequalities (see also Shimizu ). Formula (1.13) is a quantitative counterpart to the intuitive fact that the distributions of and are close, whenever is close to zero. Using (1.5) we further deduce that closeness between the Fisher informations of and (i.e. ) or between the entropies of and (i.e. ) both imply closeness in terms of the total variation distance, and hence in terms of many more probability metrics. This observation lies at the heart of the approach from where a fine analysis of the behavior of over convolutions (through projection inequalities in the spirit of (1.8)) is used to provide explicit bounds on the Fisher information distance which in turn are transformed, by means of de Bruijn’s identity (1.12), into bounds on the relative entropy. We will see in Section 4 that the bound (1.13) is too crude to be of use in the applications we are interested in.
Let the previous notation prevail. We have
Proof. Using we see that, for any function , one has
and the desired conclusion follows from de Bruijn’s identity (1.12).
To properly control the integral on the right-hand-side of (1.14), we need to deal with the fact that the mapping is not integrable in , so that we cannot directly apply the estimate E\big{[}E[Z(1-\tau_{F}(F))|F_{t}]^{2}\big{]}\leqslant{\rm Var}(\tau_{F}(F)) to deduce the desired bound. Intuitively, one has to exploit the fact that the mapping satisfies , thus in principle compensating for the singularity at .
As we will see below, one can make this heuristic precise provided there exist three constants such that
Under the assumptions appearing in condition (1.16), the following strategy can indeed be implemented in order to deduce a satisfactory bound. First split the integral in two parts: for every ,
the last inequality being a consequence of . To deal with the second term in (1.17), let us observe that, by using in particular the Hölder inequality and the convexity of the function , one deduces from (1.16) that
By virtue of (1.17) and (1.18), the term is eventually amenable to analysis, and one obtains:
Assuming finally that (recall that, in the applications we are interested in, such a quantity is meant to be close to 0) we can optimize over and choose , which leads to
Clearly, combining (1.19) with (1.5), one also obtains an estimate in total variation which agrees with (1.7) up to the square root of a logarithmic factor.
The problem is now how to identify sufficient conditions on the law of for (1.16) to hold; we shall address this issue by means of two auxiliary results. We start with a useful technical lemma, that has been suggested to us by Guillaume Poly.
where the supremum is taken over all such that .
Proof. Since we have, by using e.g. Lusin’s Theorem,
To see the reversed inequality, observe that, for any bounded by 1,
Then (1.16) holds, with and .
Proof. Take such that . Then, by independence of and ,
so that, since and ,
Inequality (1.16) now follows by applying Lemma 1.5.
As anticipated, in Section 4 (see Lemma 4.4 for a precise statement) we will describe a wide class of distributions satisfying (1.21). The previous discussion yields finally the following statement, answering the original question of providing a bound on that is comparable with the estimate (1.7).
then, provided ,
3 Plan
The rest of the paper is organised as follows. In Section 2 we will prove that Theorem 1.7 can be generalised to a fully multidimensional setting. Section 3 contains some general results related to (infinite-dimensional) Gaussian stochastic analysis. Finally, in Section 4 we shall apply our estimates in order to deduce general bounds of the type appearing in Theorem 1.1.
Entropy bounds via de Bruijn’s identity and Stein matrices
In Section 2.1 and Section 2.2 we discuss some preliminary notions related to the theory of information (definitions, notations and main properties). Section 2.3 contains the proof of a new integral formula, allowing one to represent the relative entropy of a given random vector in terms of a Stein matrix. The reader is referred to the monograph , as well as to [1, Chapter 10], for any unexplained definition and result concerning information theory.
As discussed above, we are interested in estimating the distance between the law of and the law of a -dimensional centered Gaussian vector , where is the associated covariance matrix. Our measure of the discrepancy between the distributions of and is the relative entropy (often called Kullback-Leibler divergence or information entropy)
where is the density of . It is easy to compute the Gaussian entropy (where is the determinant of ), from which we deduce the following alternative expression for the relative entropy
where ‘tr’ stands for the usual trace operator. If and have the same covariance matrix then the relative entropy is simply the entropy gap between and so that, in particular, one infers from (2.26) that has maximal entropy among all absolutely continuous random vectors with covariance matrix .
We stress that the relative entropy does not define a bona fide probability distance (for absence of a triangle inequality, as well as for lack of symmetry): however, one can easily translate estimates on the relative entropy in terms of the total variation distance, using the already recalled Pinsker-Csiszar-Kullback inequality (1.3). In the next subsection, we show how one can represent the quantity as the integral of the standardized Fisher information of some adequate interpolation between and .
2 Fisher information and de Bruijn’s identity
(with components for ), and is customarily called the Fisher information matrix of . Focussing on the case , one sees immediately that the Gaussian vector has linear score function and Fisher information .
Fix . Using formula (2.28) one deduces that a version of is given by the conditional expectation , from which we infer that the matrix is well-defined and its entries are all finite.
For , we define the standardized Fisher information matrix of as
where is the identity matrix, and the last equality holds because . Note that the positive semidefinite matrix is the difference between the Fisher information matrix of and that of a Gaussian vector having distribution . Observe that
is completely characterized (up to sets of -measure 0) by the equation
holding for every smooth test function .
Of course the above information theoretic quantities are not defined only for Gaussian mixtures of the form but more generally for any random vector satisfying the relevant assumptions (which are necessarily verified by the ). In particular, if has covariance matrix and differentiable density then, letting be the score function for , the standardized Fisher information of is
The following fundamental result is known as the (multidimensional) de Bruijn’s identity: it shows that the relative entropy can be represented in terms of the integral of the mapping with respect to the measure on . It is one of the staples of the entire paper. We refer the reader e.g. to for proofs in the case . Our multidimensional statement is a rescaling of [23, Theorem 2.3] (some more details are given in the proof). See also .
Let the above notation and assumptions prevail. Then,
Proof. In [23, Theorem 2.3] it is proved that
(note that the definition of standardized Fisher information used in is different from ours). The conclusion is obtained by using the change of variables , as well as the fact that
which follows from the scale-invariance of standardized Fisher information mentionned in Remark 2.2.
Assume that , , is a sequence of nonsingular covariance matrices such that for every , as . Then, the second and third summands of (2.35) (with replacing ) converge to 0 as .
For future reference, we will now rewrite formula (2.35) for some specific choices of , , and .
Assume . Then, (null matrix) for every , and formula (2.35) becomes
Assume that and that and have variances , respectively. Defining , relation (2.35) becomes
Relation (2.37) in the case corresponds to the integral formula (1.12) proved by Barron in [8, Lemma 1].
In the special case where , one has that
of which (1.12) is a particular case ().
In the univariate setting, the general variance case () follows trivially from the standardized one () through scaling; the same cannot be said in the multivariate setting since the appealing form (2.39) cannot be directly achieved for when the covariance matrices are not the identity because here the dependence structure of needs to be taken into account. In Lemma 2.6 we provide an estimate allowing one to deal with this difficulty in the case , for every . The proof is based on the following elementary fact: if are two symmetric matrices, and if is semi-positive definite, then
where and stand, respectively, for the maximum and minimum eigenvalue of . Observe that , the operator norm of .
Fix , and assume that . Then, , and one has the following estimates
Proof. Write and apply (2.40) first to and , and then to and .
In the next section, we prove a new representation of the quantity in terms of Stein matrices: this connection will provide the ideal framework in order to deal with the normal approximation of general random vectors.
3 Stein matrices and a key lemma
The centered -dimensional vectors are defined as in the previous section (in particular, they are stochastically independent).
The entries of the random matrix are called the Stein factors of .
Selecting in (2.43), one deduces that, if is a Stein matrix for , then . More to this point, if , then the covariance matrix is itself a Stein matrix for . This last relation is known as the Stein’s identity for the multivariate Gaussian distribution.
Assume that and that has density and variance . Then, under some standard regularity assumptions, it is easy to see that is a Stein factor for .
Let the above notation and framework prevail, and assume that is a Stein matrix for such that for every . Then, for every , the mapping
is a version of the score of . Also, the mapping
is a version of the function defined in formula (2.32).
Proof. Remember that is the score of , and denote by the mapping defined in (2.45). Removing the conditional expectation and exploiting the independence of and , we infer that, for every smooth test function ,
To simplify the forthcoming discussion, we shall use the shorthand notation: , , and . The following statement is the main achievement of the section, and is obtained by combining Lemma 2.9 with formulae (2.38) and (2.41)–(2.42), in the case where .
Let the above notation and assumptions prevail, assume that , and introduce the notation
The next subsection focusses on general bounds based on the estimates (2.47)–(2.51).
4 A general bound
The following statement provides the announced multidimensional generalisation of Theorem 1.7. In particular, the main estimate (2.55) provides an explicit quantitative counterpart to the heuristic fact that, if there exists a Stein matrix such that is small (with denoting the usual Hilbert-Schmidt norm), then the distribution of and must be close. By virtue of Theorem 2.10, the proximity of the two distributions is expressed in terms of the relative entropy .
Proof. Take such that . Then, by independence of and and using (2.44), one has, for any ,
As a result, due to Lemma 1.5 and since , one obtains
Now, using among others the Hölder inequality and the convexity of the function , we have that
with given by (2.56). At this stage, we shall use the key identity (2.50). To properly control the right-hand-side of (2.48), we split the integral in two parts: for every ,
the second inequality being a consequence of (2.57) as well as . Since (2.54) holds true, one can optimize over and choose , which leads to the desired estimate (2.55).
5 A general roadmap for proving quantitative entropic CLTs
In Section 4, we will apply the content of Theorem 2.11 to deduce entropic fourth moment theorems on a Gaussian space. As demonstrated in the discussion to follow, this achievement is just one possible application of a general strategy, leading to effective entropic estimates by means of Theorem 2.11.
Description of the strategy. Fix , and consider a sequence , , of centered random vectors with covariance and such that has a density for every . Let , , be a sequence of Gaussian vectors, and assume that , as . Then, in order to show that one has to accomplish the following steps:
Write explicitly the Stein matrix for every . This task can be realised either by exploiting the explicit form of the density of , or by applying integration by parts formulae, whenever they are available. This latter case often leads to a representation of the Stein matrix by means of a conditional expectation. As an explicit example and as shown in Section 3.3 below, Stein matrices on a Gaussian space can be easily expressed in terms of Malliavin operators.
Check that, for some , the quantity E\big{[}|\tau_{F_{n}}^{j,k}(F_{n})|^{\eta+2}\big{]} is bounded by a finite constant independent of . This step typically requires one to show that the sequence is bounded in , and moreover that the elements of this sequence verify some adequate hypercontractivity property. We will see in Section 4 that, when considering random variables living inside a finite sum of Wiener chaoses, these properties are implied by the explicit representation of Stein matrices in terms of Malliavin operators, as well as by the fundamental inequality stated in Proposition 3.4.
Prove a total variation bound such as (2.53) for every , with a constant independent of . This is arguably the most delicate step, as it consists in an explicit estimate of the total variation distance between the law of and that of , when is close to one. It is an interesting remark that, for every (as ), , so that an estimate such as (2.53) basically requires one to explicitly relate the total variation and mean-square distances between and . We will see in Section 4 that, for random variables living inside a fixed sum of Wiener chaoses, adequate bounds of this type can be deduced by applying the Carbery-Wright inequalities , that we shall combine with the approach developed by Nourdin, Nualart and Poly in .
Show that . This is realised by using the explicit representation of the Stein matrices .
Once steps (a)–(d) are accomplished, an upper bound on the speed of convergence to zero of the sequence can be deduced from (2.55), whereas the total variation distance between the laws of and follows from the inequality
Note that the quantity can be assessed by using (2.36).
Before proceeding to the application of the above roadmap on Gaussian space we now first provide, in the forthcoming Section 3, the necessary preliminary results about Gaussian stochastic analysis.
Gaussian spaces and variational calculus
As announced, we shall now focus on random variables that can be written as functionals of a countable collection of independent and identically distributed Gaussian random variables, that we shall denote by
Note that our description of is equivalent to saying that is a Gaussian sequence such that for every and . We will write to indicate the class of square-integrable (real-valued) random variables that are measurable with respect to the -field generated by .
The reader is referred e.g. to for any unexplained definition or result appearing in the subsequent subsections.
We will now briefly introduce the notion of Wiener chaos.
The sequence of Hermite polynomials is defined as follows: , and, for ,
A multi-index is a sequence of nonnegative integers such that only for a finite number of indices . We use the symbol in order to indicate the collection of all multi-indices, and use the notation , for every .
It is easily seen that two random variables belonging to Wiener chaoses of different orders are orthogonal in . Moreover, since linear combinations of polynomials are dense in , one has that , that is, any square-integrable functional of can be written as an infinite sum, converging in and such that the th summand is an element of . This orthogonal decomposition of is customarily called the Wiener-Itô chaotic decomposition of .
where is given in (3.59). Another classical result (see e.g. ) is that, for every , the mapping (as defined in (3.60)) is onto, and defines an isomorphism between and the Hilbert space , endowed with the modified norm . This means that, for every ,
Finally, we observe that one can reexpress the Wiener-Itô chaotic decomposition of as follows: every admits a unique decomposition of the type
where the series converges in , the symmetric kernels , , are uniquely determined by , and . This also implies that
2 The language of Malliavin calculus: chaoses as eigenspaces
We let the previous notation and assumptions prevail: in particular, we shall fix for the rest of the section a real separable Hilbert space , and represent the elements of the th Wiener chaos of in the form (3.60). In addition to the previously introduced notation, indicates the space of all -valued random elements , that are measurable with respect to and verify the relation E\big{[}\|u\|_{\EuFrak{H}}^{2}\big{]}<\infty. Note that, as it is customary, is endowed with the Borel -field associated with the distance on given by .
Let be the set of all smooth cylindrical random variables of the form
In what follows, we denote by the adjoint of the operator , also called the divergence operator. A random element belongs to the domain of , written , if and only if it satisfies
for some constant depending only on . If , then the random variable is defined by the duality relationship (customarily called “integration by parts formula”):
Let be an independent copy of , and denote by the mathematical expectation with respect to . For every the operator is defined as follows: for every ,
in such a way that and . The collection verifies the semigroup property and is called the Ornstein-Uhlenbeck semigroup associated with .
The properties of the semigroup that are relevant for our study are gathered together in the next statement.
For every , the eigenspaces of the operator coincide with the Wiener chaoses , , the eigenvalue of being given by the positive constant .
The infinitesimal generator of , denoted by , acts on square-integrable random variables as follows: a random variable with the form (3.61) is in the domain of , written , if and only if is convergent in , and in this case
In particular, each Wiener chaos is an eigenspace of , with eigenvalue equal to .
In view of the previous statement, it is immediate to describe the pseudo-inverse of , denoted by , as follows: for every mean zero random variable of , one has that
For future reference, we record the following estimate involving random variables living in a finite sum of Wiener chaoses: it is a direct consequence of the hypercontractivity of the Ornstein-Uhlenbeck semigroup – see e.g. in [36, Theorem 2.7.2 and Theorem 2.8.12].
Let and . Then, there exists a finite constant such that, for every ,
In particular, all norms, , are equivalent on a finite sum of Wiener chaoses.
Since we will systematically work on a fixed sum of Wiener chaoses, we will not need to specify the explicit value of the constant . See again , and the references therein, for more details.
which is a well-defined element of .
3 The role of Malliavin and Stein matrices
The following statement is taken from , and provides a simple necessary and sufficient condition for a random vector living in a finite sum of Wiener chaoses to have a density.
(which is a direct consequence of for any ), one has that
To conclude, we present a result providing an explicit representation for Stein matrices associated with random vectors in the domain of .
Taking conditional expectations yields the desired conclusion.
The next section contains the statements and proofs of our main bounds on a Gaussian space.
Entropic fourth moment bounds on a Gaussian space
Our main result is the following entropic central limit theorem for sequences of chaotic random variables.
Let and be fixed integers. Consider vectors
where indicates a bounded numerical sequence depending on , as well as on the sequence .
One immediate consequence of the previous statement is the following characterisation of entropic CLTs on a finite sum of Wiener chaoses.
Let the sequence be as in the statement of Theorem 4.1, and assume that . Then, the following three assertions are equivalent, as :
converges in distribution to ;
The proofs of the previous results are based on two technical lemmas that are the object of the next section.
2 Estimates based on the Carbery-Wright inequalities
where are independent random variables with common distribution .
Fix , and let be a random vector such that with . Let denote the Malliavin matrix of , and assume that (which is equivalent to assuming that has a density by Theorem 3.6). Set with . Then, there exists a universal constant such that
Proof. Let be an orthonormal basis of . Since is a polynomial of degree in the entries of and because each entry of belongs to by the product formula for multiple integrals (see, e.g., [36, Chapter 2]), we have, by iterating the product formula, that . Thus, there exists a sequence of real-valued polynomials of degree at most such that the random variables converge in and almost surely to as tends to infinity (see [40, proof of Theorem 3.1] for an explicit construction). Assume now that . Then, for sufficiently large, . We deduce from the estimate (4.67) the existence of a universal constant such that, for any ,
from which (4.68) follows by letting tend to infinity.
In order to estimate the two last terms in (4.71), we decompose the expectation into two parts using the identity
Step 3. For all and using (4.68),
Choosing yields
with the usual adjugate matrix operator.
From the hypercontractivity property together with the equality
one immediately deduces the existence of such that
Substituting this estimate into (4.74) and assuming that , yields
Choosing and assuming , we obtain
It is worthwhile noting that the inequality (4.2) is valid for any , in particular for .
Step 6. From (4.71), (LABEL:7) and (4.2) we obtain
By plugging this inequality into (4.70) we thus obtain that, for every , and :
Choosing , and , one obtains the desired conclusion (4.69).
3 Proof of Theorem 4.1
In the proof of [41, Theorem 4.3], the following two facts have been shown:
Using Proposition 3.7, one infers immediately that
defines a Stein’s matrix for , which moreover satisfies the relation
Now let denote the Malliavin matrix of . Thanks to [43, Lemma 6], we know that, for any ,
Since lives in a finite sum of chaoses (see e.g. [36, Chapter 5]) and is bounded in , we can again apply the hypercontractive estimate (3.64) to deduce that is actually bounded in for every , so that the convergence in (4.79) is in the sense of any of the spaces . As a consequence, and there exists large enough so that
This means that relation (2.53) is satisfied uniformly on . Concerning (2.52), again by hypercontractivity and using the representation (4.77), one has that, for all ,
Finally, since and because (4.78) holds true, the condition (2.54) is satisfied for large enough. The proof of (4.66) is concluded by applying Theorem 2.11.
4 Proof of Corollary 4.2
In view of Theorem 4.1, one has only to prove that (b) implies (a). This is an immediate consequence of the fact that the covariance converges to , and that the sequence lives in a finite sum of Wiener chaoses. Indeed, by virtue of (3.64) one has that
yielding in particular that, if converges in distribution to , then or, equivalently, . The proof is concluded.
Acknowledgement. We heartily thank Guillaume Poly for several crucial inputs that helped us achieving the proof of Theorem 4.1. We are grateful to Larry Goldstein and Oliver Johnson for useful discussions. Ivan Nourdin is partially supported by the ANR grant ANR-10-BLAN-0121.