Adaptive density estimation for directional data using needlets

P. Baldi, G. Kerkyacharian, D. Marinucci, D. Picard

Introduction

There is an abundant literature about this type of problems. In particular, minimax L2L^{2} results have been obtained (see [Kle99],[Kle03]). These procedures are generally obtained using either kernel methods (but in this case the manifold structure of the sphere is not well taken into account), or using orthogonal series methods associated with spherical harmonics (and in this case the ’local performances of the estimator are quite poor, since spherical harmonics are spread all over the sphere).

In our approach we focus on two important points. We aim at a procedure of estimation which is efficient from a L2L^{2} point of view (as it is a tradition in statistics to evaluate the procedure with the mean square error). On the other hand, we would like it to perform satisfactorily also from a local point of view (in infinity norm, for instance). To have these two requirements together seems to us a warrant to have good results in practice. In effect, it is very difficult to produce a loss function which reflects at the same time the requirement of clearly seeing the bumps of the density, of being able to well estimate different level sets, of testing whether there is a difference between the northern and southern hemispheres and so on.

In addition, we require this procedure to be simple to implement, as well as adaptive to inhomogeneous smoothness.

This type of requirements is generally well handled using thresholding estimates associated to wavelets. The problem requires a special construction adapted to the sphere, since usual tensorized wavelets will never reflect the manifold structure of the sphere and will necessarily create unwanted artifacts. Recently in ([NPW06b],[NPW06a]) a tight frame (i.e. a redundant family) was produced which enjoys enough properties to be successfully used for density estimation.

The fundamental properties of wavelets are their concentration in the Fourier domain as well as in the space domain. Here, obviously the ’space’ domain is the sphere itself whereas the Fourier domain is now obtained by replacing the ’Fourier’ basis by the basis of Spherical Harmonics which plays an analogous role on the sphere.

The construction [NPW06b],[NPW06a] produces a family of functions which very much resemble to wavelets, the needlets, and in particular have very good concentration properties.

We use these needlets to construct an estimation procedure, and prove that this procedure attains optimal rates over various spaces of regularity.

Again, the problem of choosing appropriated spaces of regularity on the sphere in a serious question, and we decided to consider the spaces which may be the closest to our natural intuition: those which generalize to the sphere case the classical Hölder spaces.

In the first section we present ([NPW06b]) needlets, and describe spaces of regularity on the sphere. In the second one we define our estimation procedure, and describe its properties.

The novelties of this paper lie in the application of thresholding to the needlet coefficients, which gives a very simple and adaptive procedure which works on the sphere. We also focus here on giving the results in L∞L_{\infty} norm, and obtain the rates of convergence for many other loss functions as a consequence of the previous ones.

Our results are motivated by many recent developments in the area of observational astrophysics. As an example, we refer to experiments measuring incoming directions of Ultra High Energy Cosmic Rays, such as the AUGER Observatory (http://www.auger.org). Here, efficient estimation of the density function of these directional data may yield crucial insights into the physical mechanisms generating the observations. More precisely, a uniform density would suggest the High Energy Cosmic Rays are generated by cosmological effects, such as the decay of massive particles generated during the Big Bang; on the other hand, if these Cosmic Rays are generated by astrophysical phenomena (such as acceleration into Active Galactic Nuclei), then we should observe a density function which is highly non-uniform and tightly correlated with the local distribution of nearby Galaxies. Massive amount of data in this area are expected to be available in the next few years. The Auger observatory will be based on two arrays of detectors; the first one covers an area larger than 3000 Km2 in Pampa Amarilla (Argentina), and has already started to collect observations: some preliminary evidence was provided in [Col08], and a non-uniform distribution seems to be favored. The whole celestial sphere will actually be covered only when the construction of the northern hemisphere array, due to be built in eastern Colorado, will be completed, a few years from now. Hence, in the immediate future efficient statistical techniques will be eagerly requested for the analysis of the forthcoming datasets.

A survey of statistical methodologies dealing with directional data on the sphere may be found in [Mar72], [Jup95], [MJ00]. The generalization of estimation using orthogonal series methods to the case of compact Riemannian manifold can be found in [Hen03]. See related works in [HK96], [Ruy89], [HJR93], [Pel05], [Jup08]. Kernel methods on the sphere have been investigated in [HWC87]. Minimax rates for the equivalent of Sobolev spaces on the sphere associated can be found in [Kle99], [Kle00], [Kle03].

The plan of the paper is as follows. In §2 and §3 we review some background material on needlets and Besov spaces. §4 introduces our thresholding estimator, whose minimax performances are stated in §5. §6 shows the performance of the estimators on some simulated data. §7–§9 contain the proofs.

Needlets

This construction is due to Narcowich, Petrushev and Ward [NPW06b]. Its aim is essentially to build a very well localized tight frame constructed using spherical harmonics, as discussed below. It was recently extended with fruitful statistical applications to more general Euclidean settings (see [KPPW07]) and already exploited for estimation and testing problems in [BKMP06], [BKMP07].

For the main situation of interest, d=2d=2, the right hand side above is equal to 2l+18π2\frac{2l+1}{8\pi^{2}}. Recall that if d=2d=2, the usual normalization of the Legendre polynomial (Ll(1)=1L_{l}(1)=1) gives the square of their L2L^{2} norm equal to 22l+1\frac{2}{2l+1}. Therefore these must be multiplied by (2l+1)/(4π)(2l+1)/(4\pi), in order to satisfy (3).

Let us point out the following reproducing property of the projection operators:

The construction of needlets is based on the classical Littlewood-Paley decomposition and a subsequent discretization.

Remark that b(ξ)≠0b(\xi)\not=0 only if 12≤∣ξ∣≤2\tfrac{1}{2}\leq|\xi|\leq 2. Let us now define the operator Λj=∑l≥0b2(l2j)Ll\Lambda_{j}=\sum_{l\geq 0}b^{2}(\tfrac{l}{2^{j}})L_{l} and the associated kernel

Moreover, if Mj(x,y)=∑l≥0b(l2j)Ll(⟨x,y⟩)M_{j}(x,y)=\sum_{l\geq 0}b(\frac{l}{2^{j}})L_{l}(\langle x,y\rangle), then

Then the operator MjM_{j} defined in the subsection above is such that: z↦Mj(x,z)∈P[2j+1]z\mapsto M_{j}(x,z)\in\mathscr{P}_{[2^{j+1}]}, so that

The choice of the sets Zj{\mathscr{Z}}_{j} of cubature points is not unique, but one can impose the conditions

for some c>0c>0. Actually in the simulations of §6 we make use of some sets of cubature points for d=2d=2 such that #Zj=22j+4\#\mathscr{Z}_{j}=2^{2j+4} exactly (the corresponding weights being however not identical). We have, using (6)

where dd is the natural geodesic distance on the sphere (for d=2d=2, d(ξ,η)=arccos⁡⟨η,ξ⟩d(\xi,\eta)=\arccos\langle\eta,\xi\rangle ). In other words needlets are almost exponentially localized around any cubature point, which motivates their name. From this localization property it follows (see [NPW06b]) that for 1≤p≤+∞1\leq p\leq+\infty there exists positive constants cp,Cpc_{p},C_{p} such that

Let us prove (13) for p=+∞p=+\infty. Using (11) and Lemma 6 of [BKMP06]

If 1≤p<+∞1\leq p<+\infty, by Hölder inequality, if 1p+1p′=1\frac{1}{p}+\frac{1}{p^{\prime}}=1 so that pp′=p−1\frac{p}{p^{\prime}}=p-1,

where the last inequality comes again from (11) and Lemma 6 of [BKMP06]. Now integrating and using (12) for p=1p=1,

from which (13) follows. The remaining case 0<p≤10<p\leq 1 follows immediately by subadditivity, as

But, by Holder inequality, for p′p^{\prime} such that 1p+1p′=1,thenpp′=p−1\frac{1}{p}+\frac{1}{p}^{\prime}=1,then\frac{p}{p^{\prime}}=p-1 and

Relation (12) for p=2p=2 states that the L2L^{2} norm of ψjξ\psi_{j\xi} is bounded with respect to jj and also bounded away from from below. Assume d=2d=2. Then using (4) it is actually easy to see that, keeping in mind that Ll(1)=2l+14πL_{l}(1)=\frac{2l+1}{4\pi},

Assuming that the cubature points are of cardinality 22j+42^{2j+4} and that they sum up to 4π4\pi, λη∼4π⋅2−2j−4\lambda_{\eta}\sim 4\pi\cdot 2^{-2j-4} as j→∞j\to\infty. If the previous relation were an equality we could recognize in the right hand term the Riemann sum

that converges, as j→∞j\to\infty, to the integral

which depends on the choice of the function bb. This L2L^{2} norm shall appear frequently in the sequel. For instance, if we write down the development (10) or the function f=ψj0ξ0f=\psi_{j_{0}\xi_{0}}, then the coefficient βj0ξ0=⟨f,ψj,η⟩\beta_{j_{0}\xi_{0}}=\langle f,\psi_{j,\eta}\rangle would be exactly equal to ∥ψjξ∥22\|\psi_{j\xi}\|^{2}_{2}. As it is clear that it would be desirable for this coefficient to be as large as possible, the value of the integral above can be seen as a measure of the localization properties of the system of needlets and can be used as a criterion of goodness of the choice of the function bb. With the choice we made (see §6) the quantity II above is ≃0.107\simeq 0.107.

Besov spaces on the sphere and needlets

In this section we summarize the main properties of Besov spaces and needlets, as established in [NPW06b].

the infimum of the distances in LrL^{r} of ff from the polynomials of degree kk. Then the Besov space Br,qsB^{s}_{r,q} is defined as the space of functions such that

Remarking that k→Ek(f,r))k\to E_{k}(f,r)) is decreasing, by a standard condensation argument this is equivalent to

Let 1≤r≤+∞1\leq r\leq+\infty, s>0s>0, 0≤q≤+∞0\leq q\leq+\infty. Let ff a measurable function and define

provided the integrals exists. Then f∈Br,qsf\in B^{s}_{r,q} if and only if, for every j=1,2,…j=1,2,\dots,

for some positive constants c,Cc,C, the Besov space Br,qsB^{s}_{r,q} turns out to be a Banach space associate to the norm

In the sequel we shall denote by Br,qs(M)B^{s}_{r,q}(M) the ball of radius MM of the Besov space Br,qsB^{s}_{r,q}.

(The Besov embedding) If p≤r≤∞p\leq r\leq\infty then Br,qs⊆Bp,qsB^{s}_{r,q}\subseteq B^{s}_{p,q}. If s>d(1r−1p)s>d(\frac{1}{r}-\frac{1}{p}),

On the other hand, if r≤p≤∞r\leq p\leq\infty,

Needlet estimation of a density on the sphere

Let us suppose that we observe X1,…,XnX_{1},\ldots,X_{n}, i.i.d. random variables taking values on the sphere having common density ff with respect to dxdx. ff can be decomposed using the frame of needlets described above.

The needlet estimator is based on hard thresholding of a needlet expansion as follows. We start by letting:

The tuning parameters of the needlet estimator are:

∙\bullet The range J=J(n)J=J(n) of resolution levels (frequencies) where the approximation (17) is used:

We shall see that the choice 2^{J}=\big{(}\frac{n}{\log n}\big{)}^{\frac{1}{d}} is appropriate.

∙\bullet The threshold constant κ\kappa. Evaluations of κ\kappa are given in the following Section and also discussed in §6.

∙\bullet cnc_{n}: is a sample size-dependent scaling factor. We shall see that an appropriate choice is

which is a quantity already discussed in Example 3. As for the correlation between coefficients, it is given by the function θ→λη∑l≥0b2(l2j)Ll(cos⁡θ)\theta\to\lambda_{\eta}\sum_{l\geq 0}b^{2}(\tfrac{l}{2^{j}})L_{l}(\cos\theta), whose graph, for some values of jj is plotted in Figure 2.

Whereas coefficients associated to cubature points that are not too close are only slightly correlated, the random needlet coefficients β^j,η\widehat{\beta}_{j,\eta}, η∈Zj\eta\in{\mathscr{Z}}_{j} are not independent and they even satisfy the linear relation

This comes from the fact that, as y→Ll(⟨y,x⟩)y\to L_{l}(\langle y,x\rangle) for l≤2jl\leq 2^{j} is a polynomial of degree ≤22j\leq 2^{2j}, one has

For 0<r≤∞0<r\leq\infty, p≥1p\geq 1, s>drs>\frac{d}{r} we have

a) For any z>1z>1, there exist some constants c∞=c∞(s,p,r,A,M)c_{\infty}=c_{\infty}(s,p,r,A,M) such that if κ>z+16\kappa>\frac{z+1}{6},

b) For 1≤p<∞1\leq p<\infty there exist some constant cp=cp(s,r,p,A,M)c_{p}=c_{p}(s,r,p,A,M) such that if κ>p12\kappa>\frac{p}{12},

where αp=p−1+1{r=dp2s+d}\alpha_{p}=p-1+1_{\{r=\frac{dp}{2s+d}\}}, if r≤dp2s+dr\leq\frac{dp}{2s+d}, whereas

which, as remarked above gives an indication about the square of the value of the L2L^{2} norm of a needlet ψjξ\psi_{j\xi}. In both the examples below we considered samples of cardinality n=2000n=2000 and n=8000n=8000. The hint for the value of JJ of Theorem 8 is J=\frac{1}{2}\,\log_{2}\big{(}\frac{n}{\log n}\big{)}, which gives the values J∼4.02J\sim 4.02 and J∼4.9J\sim 4.9 respectively. One should keep in mind that at a given level jj it is necessary to have enough cubature points in order to integrate exactly all polynomials up to the degree 2(2j+1−1)=2j+2−22(2^{j+1}-1)=2^{j+2}-2, which means ∼22j+4\sim 2^{2j+4} cubature points with Womersley’s set (recall that on the sphere the polynomials of degree dd form a vector space of dimension (2d+1)2(2d+1)^{2}). This gives 210=10242^{10}=1024 cubature points for j=3j=3, 212=40962^{12}=4096 for j=4j=4 and 214=163842^{14}=16384 for j=5j=5. To avoid to have more coefficients than observations, we decided to set J=3J=3 for n=2000n=2000 and J=4J=4 for n=8000n=8000.

As for the value of κ\kappa, we shall give the result with κ=k00.107 M\kappa=k_{0}\sqrt{0.107}\,M, where MM is an a bound for ∥f∥∞\|f\|_{\infty}, trying different values of k0k_{0}. Recall that this means that the threshold kills all coefficients βjξ\beta_{j\xi} such that ∣βjξ∣<κlog⁡nn|\beta_{j\xi}|<\kappa\sqrt{\frac{\log n}{n}}

f=14πf=\frac{1}{4\pi}, the uniform density. In this case in the development (10) it holds βjξ=⟨f,ψjξ⟩L2=0\beta_{j\xi}=\langle f,\psi_{j\xi}\rangle_{L^{2}}=0 for every jj and ξ\xi. Therefore a first simple way of assessing the performance of the procedure is to count the number of coefficients that survive thresholding. Of course in this case a good estimate is such that the coefficients βj,ξ\beta_{j,\xi} fall below the threshold. Taking into account Lemma 2 the square root of the sum of the squares of the coefficients surviving thresholding gives an estimate of ∥f^−f∥2\|\hat{f}-f\|_{2}. Therefore a measure of the goodness of the fit is obtained by taking the sum of their squares. Tables 1 and 2 give the number of surviving coefficients for different values of the constant k0k_{0}.

In order to kill all the coefficients one should choose k0=5.4k_{0}=5.4 for n=2000n=2000 and k0=2.8k_{0}=2.8 for n=8000n=8000. The estimate of the L2L^{2} norm of the difference between f^\hat{f} and ff by taking the square root of the sum of the squares of the coefficients is

Let us consider a mixture ff of two densities of the form fi(x)=ciexp⁡(−ki∣x−xi∣2)f_{i}(x)=c_{i}\exp(-k_{i}|x-x_{i}|^{2}), i=1,2i=1,2, for k1=.7k_{1}=.7 and k2=2k_{2}=2 and with weights 0.650.65 and 0.350.35 respectively. Here the centers xix_{i} of the two bell-shaped densities were taken to be x1=(0,1,0)x_{1}=(0,1,0), x2=(0,−.8,.6)x_{2}=(0,-.8,.6). With these choices it turns out that ∥f∥∞=0.26\|f\|_{\infty}=0.26. The graph of ff in the coordinates (φ,θ)(\varphi,\theta) (φ=\varphi=longitude, θ=\theta=colatitude) is given in Figure 4.

The estimator f^\hat{f} obtained with the choice k0=1.1k_{0}=1.1 has the graph of Figure 5.

If one chooses k0=1.5k_{0}=1.5 the graph becomes the one of Figure 6. At a closer inspection it turns out that with this value of kk all coefficients at level j=3j=3 do not pass the thresholding. It looks very much like the graph of ff, even though some differences in shape are apparent. An estimate of the norm ∥f^−f∥∞\|\hat{f}-f\|_{\infty} computed on a grid gives ∥f^−f∥∞∼0.054\|\hat{f}-f\|_{\infty}\sim 0.054.

We repeated the simulation with n=8000n=8000 observations. The results are reported in Figures 7 and 8 and are to be considered rather satisfactory. It should be stressed that a very limited number of coefficients passes thresholding at a frequency j>2j>2. This behaviour is expected: ff being very regular, it belongs to a space B∞,∞∞B^{\infty}_{\infty,\infty} and thus its needlet coefficients decay very rapidly.

In the sequel we note t(β^j,ξ)=β^jη 1{∣β^jη∣≥ κ cn}t(\hat{\beta}_{j,\xi})=\hat{\beta}_{j\eta}\,{1}_{\{|\hat{\beta}_{j\eta}|\geq\,\kappa\,c_{n}\}}, so that the needlet estimator (17) is

The following proposition collects the main estimates needed in the proof.

Let J1≤JJ_{1}\leq J be such that, for all J1≤j≤JJ_{1}\leq j\leq J, ∣βjη∣≤κ2tn|\beta_{j\eta}|\leq\tfrac{\kappa}{2}t_{n} (possibly J1=JJ_{1}=J; obviously, when ff belongs to a Besov class, J1J_{1} depends on the “regularity” ss). Then for any γ>0,s>0,z≥1\gamma>0,s>0,z\geq 1, we have

1) if κ>13(γd+12)\kappa>\frac{1}{3}(\frac{\gamma}{d}+\frac{1}{2})

We delay the proof of Proposition 14 to §8 and derive from it the proof of Theorem 8. In this proof CC will denote an absolute constant which may change from line to line. Let us now prove that Proposition 14 yields to the statements of Theorem 8:

Let us prove the L∞L^{\infty} upper bound (19), first under the condition q=r=∞q=r=\infty

The term IIII is easy to analyze: as ff belongs to B∞,∞s(M)B^{s}_{\infty,\infty}(M), we have using (12) and 13,

Then we only need to remark that sd≥sd+2s\frac{s}{d}\geq\frac{s}{d+2s} for s>0s>0.

As for II, using the triangular inequality together with Hölder inequality, then (12), and (22) with γ=dz2,z>1\gamma=\frac{dz}{2},z>1, we get

As ff belongs to B∞,∞s(M)B^{s}_{\infty,\infty}(M), ∣βjη∣≤M2−j(s+d2)|\beta_{j\eta}|\leq M2^{-j(s+\frac{d}{2})} and we can see that

For arbitrary q,rq,r (19) is now easy to deduce from the previous computation by the Besov embedding (Theorem 5) Br,qs(M)⊂B∞,∞s−dr(M).B^{s}_{r,q}(M)\subset B^{s-\frac{d}{r}}_{\infty,\infty}(M). Let us prove (21), that is the regular case. We observe first that since Br,qs(M)⊂Bp,qs(M)B^{s}_{r,q}(M)\subset B^{s}_{p,q}(M) for r≥pr\geq p, this case will be assimilated to the case p=rp=r, and from now on, we only consider r≤pr\leq p. We follow the same arguments as above. (25) can be replaced by

For IIII using the embedding Br,qs(M)⊂Bp,qs−dp+dp(M)B^{s}_{r,q}(M)\subset B^{s-\frac{d}{p}+\frac{d}{p}}_{p,q}(M), for r≤pr\leq p, we have

And it is easy to verify that sd−1r+1p≥s2s+d\frac{s}{d}-\frac{1}{r}+\frac{1}{p}\geq\frac{s}{2s+d} on the zone that we are considering in this part. In effect as s≥p2(dr−dp)s\geq\frac{p}{2}(\frac{d}{r}-\frac{d}{p}), s2s+d≤srdp\frac{s}{2s+d}\leq\frac{sr}{dp} we have sd−1r+1p−srdp=(1r−1p)(sdr−1)≥0\frac{s}{d}-\frac{1}{r}+\frac{1}{p}-\frac{sr}{dp}=(\frac{1}{r}-\frac{1}{p})(\frac{s}{d}r-1)\geq 0.

For II, we have using the triangular inequality together with Hölder inequality,

Then we need only to use (24), with γ=dp2,z=p\gamma=\frac{dp}{2},z=p, to obtain

is adequate, and to observe that the first term in the sum has the right order. For the second term, it can be bounded (as p≥rp\geq r) by

as, sr−d2(p−r)≥0.sr-\frac{d}{2}(p-r)\geq 0. Now, this term obviously is of the right order.

Again we proceed as above and observe first that in order to have s>0s>0 as well as s≤pd2(1r−1p)s\leq\frac{pd}{2}(\frac{1}{r}-\frac{1}{p}), it is necessary that p≥rp\geq r:

For IIII using the embedding, Br,qs(M)⊂Bp,qs−dr+dp(M)B^{s}_{r,q}(M)\subset B^{s-\frac{d}{r}+\frac{d}{p}}_{p,q}(M), for r≤pr\leq p, we have:

And it is easy to verify that sd−1r+1p≥(s−d(1r−1p))2(s−d(1r−12))\frac{s}{d}-\frac{1}{r}+\frac{1}{p}\geq\frac{(s-d(\frac{1}{r}-\frac{1}{p}))}{2(s-d(\frac{1}{r}-\frac{1}{2}))}, since 2(s−d(1r−12))≥d2(s-d(\frac{1}{r}-\frac{1}{2}))\geq d, when s>drs>\frac{d}{r}.

For II, again, we have using the triangular inequality together with Hölder inequality,

Then we need to use (23), with γ=dp2,z=p\gamma=d\frac{p}{2},z=p, to obtain:

as we are in the sparse region. It is easy to realize that now, again because we are in the sparse region

is adequate, and to observe then that the first term in the sum has the right order. For the second term, let us introduce

We easily observe that p−m=p(s+dp−dr)s+d2−dr>0p-m=\frac{p(s+\frac{d}{p}-\frac{d}{r})}{s+\frac{d}{2}-\frac{d}{r}}>0, and that m−r=−sr+(p−r)d2s+d2−dr≥0.m-r=\frac{-sr+(p-r)\frac{d}{2}}{s+\frac{d}{2}-\frac{d}{r}}\geq 0. Then, as Br,qs(M)⊂Bm,qs−dr+dm(M)B^{s}_{r,q}(M)\subset B^{s-\frac{d}{r}+\frac{d}{m}}_{m,q}(M)

which gives the right order. Observe that the term JJ (which is of logarithmic order), can be avoided by choosing m~\widetilde{m} instead of mm in such a way that m~>m\widetilde{m}>m, but r<m~r<\widetilde{m}. This can be done except for the case where r=dp2s+dr=\frac{dp}{2s+d} where this logarithmic term is unavoidable.

The proof of Proposition 14 relies on the following lemma:

There exist constants σ2>0,  C,  c\sigma^{2}>0,\;C,\;c, such that, as soon as 2j≤[nlog⁡n]1d2^{j}\leq[\frac{n}{\log n}]^{\frac{1}{d}},

Proof of the lemma (28) is simply Bernstein inequality, noticing that

and ∥ψjη(Xi)∥∞≤c2jd2\|\psi_{j\eta}(X_{i})\|_{\infty}\leq c2^{\frac{jd}{2}}. The following inequality directly follows from (28), when 2jd≤n2^{jd}\leq n.

using the change of variables u=nvu=\sqrt{n}v.

(30) also follows from (32): take a=max⁡{8cd3,22dσ}a=\max\{\frac{8cd}{3},{2\sqrt{2}d\sigma}\}

Now, if v≥ajnv\geq\frac{aj}{\sqrt{n}}, 2jde−nv24σ2≤e−nv48σ2−nv48σ2+jd≤e−nv48σ22^{jd}e^{-\frac{nv^{2}}{4\sigma^{2}}}\leq e^{-\frac{nv^{4}}{8\sigma^{2}}-\frac{nv^{4}}{8\sigma^{2}}+jd}\leq e^{-\frac{nv^{4}}{8\sigma^{2}}}. Similarly 2jde−3nv4c≤e−3nv8c2^{jd}e^{-\frac{3\sqrt{n}v}{4c}}\leq e^{-\frac{3\sqrt{n}v}{8c}}, so that

Let us now turn to the proof of the Proposition. We partition our sum in four regions:

We use extensively Lemma 15 in order to bound separately each of the four terms Bb,Ss,Sb,BsBb,Ss,Sb,Bs.

where J1J_{1} is chosen such that for j≥J1j\geq J_{1}, ∣βjη∣≤κ2tn|\beta_{j\eta}|\leq\frac{\kappa}{2}t_{n}. Also

which gives the proper rate of convergence. Moreover, using (30) and (31),

where κ>13(γd+12)\kappa>\frac{1}{3}(\frac{\gamma}{d}+\frac{1}{2}). Finally, using (31), and the fact that for ff bounded, ∣βjη∣≤C2−jd2|\beta_{j\eta}|\leq C2^{-\frac{jd}{2}}

for κ>16(γd+1)\kappa>\frac{1}{6}(\frac{\gamma}{d}+1).

This proof follows along the lines of the previous one. (24) is a consequence of (23), and the two inequalities will be proved together. We again separate the four cases.

Let us now bound separately each of the four terms Bb,  Ss,  Sb,  BsBb,\;Ss,\;Sb,\;Bs. Using (29)

where J1J_{1} again is chosen such that for j≥J1j\geq J_{1}, ∣βjη∣≤κ2tn|\beta_{j\eta}|\leq\frac{\kappa}{2}t_{n}. To prove (23), we stop in (*), the next bound yields (24).

which gives the proper rate of convergence. Again, to prove (23), we stop in (*), the next bound yields (24). Moreover, using (30) and (31),

for κ>γ6d\kappa>\frac{\gamma}{6d}. Finally, using (31), and the fact that for ff bounded, ∣βjη∣≤C2−jd2|\beta_{j\eta}|\leq C2^{-\frac{jd}{2}}

Let us recall that given two probabilities PP, QQ on some measure space their Kullback-Leibler distance is

We make use of Fano’s lemma below, see [Tsy04] and the references therein. We use the point of view introduced in [Bir01].

(Fano’s Lemma) Let F\mathscr{F} be a sigma algebra on the space Ω.\Omega. Let Fi∈F,  i∈{0,1,…,m}F_{i}\in\mathscr{F},~{}~{}i\in\{0,1,\ldots,m\} such that ∀i≠j,Fi∩Fj=∅.\forall i\neq j,F_{i}\cap F_{j}=\emptyset. Let PiP_{i}, i=0,…,mi=0,\ldots,m be probability measures on (Ω,A).(\Omega,\mathscr{A}). If

∙\bullet We prove first that the minimax LpL^{p}-loss is ≳n−αp\gtrsim n^{-\alpha p} with α=s2s+d\alpha=\frac{s}{2s+d}. For every jj let us consider the family Aj{\mathscr{A}}_{j} of densities

where AjA_{j} is a subset of Zj{\mathscr{Z}}_{j} to be made precise later, εξ∈{0,1}\varepsilon_{\xi}\in\{0,1\} and γ\gamma is chosen so that all these functions are positive. We are going to show that for every estimator f^\hat{f},

Throughout this section we shall note x≲yx\lesssim y, x≳yx\gtrsim y whenever it holds x≤cyx\leq cy or x≥cyx\geq cy respectively, cc being a strictly positive constant independent of j,ξj,\xi. We shall note x≃yx\simeq y whenever both x≲yx\lesssim y and x≳yx\gtrsim y hold. Thanks to (12) for these functions to be positive it is enough that ∣γ∣≲2−jd/2|\gamma|\lesssim 2^{-jd/2}. Such a γ\gamma can even be chosen in such a way that all the densities (37) are bounded from below by a strictly positive constant. If the functions (ψjξ)ξ∈Zj(\psi_{j\xi})_{\xi\in{\mathscr{Z}}_{j}} were orthonormal we would have immediately that

Needlets are not a basis, but their scalar product is close enough to if the respective cubature points are far enough. Hence one can get the following lemma that states that a subset Aj⊂ZjA_{j}\subset{\mathscr{Z}}_{j} can be chosen so that it is quite large and inequalities (14) and (13), in a sense, can be reversed.

There exists a subset Aj⊂ZjA_{j}\subset{\mathscr{Z}}_{j} such that card Aj≳2jd{\rm card}\,A_{j}\gtrsim 2^{jd} and

Let us now impose conditions that ensure that fεf_{\varepsilon} belongs to the ball Br,qs(M)B^{s}_{r,q}(M). Now, recalling (15),

where we use the fact that ∣εξ∣=1|\varepsilon_{\xi}|=1. Therefore the condition ∥fε∥Br,qs≤M\|f_{\varepsilon}\|_{B^{s}_{r,q}}\leq M follows from

In order to apply Fano’s Lemma and get a lower bound of the left hand side let us first get an upper bound for the Kullback-Leibler distances K(fε;fε)K(f_{\varepsilon};f_{\varepsilon}), which comes from (33) and (12) for p=2p=2,

Thanks to Lemma 17, the set of functions Aj{\mathscr{A}}_{j} has a cardinality that is ≥2c2jd\geq 2^{c2^{jd}}. By the Varshanov-Gilbert Lemma ([Tsy04] e.g.) there exists a subset Aj′⊂Aj{\mathscr{A}}^{\prime}_{j}\subset{\mathscr{A}}_{j} such that card Aj′≥2c′2jd{\rm card}\,{\mathscr{A}}^{\prime}_{j}\geq 2^{c^{\prime}2^{jd}} and such that if fε,fε′∈Aj′f_{\varepsilon},f_{\varepsilon}^{\prime}\in{\mathscr{A}}^{\prime}_{j}, then ∑ξ∈Aj′∣εξ−εξ′∣>14 2jd\sum_{\xi\in A^{\prime}_{j}}|\varepsilon_{\xi}-\varepsilon^{\prime}_{\xi}|>\frac{1}{4}\,2^{jd}. Therefore, as ∣εξ−εξ′∣|\varepsilon_{\xi}-\varepsilon^{\prime}_{\xi}| can be =0=0 or =1=1 only and by (12),

which implies that the events {∥f^−fε′∥p≥12 2−js}\{\|\widehat{f}-f_{\varepsilon^{\prime}}\|_{p}\geq\tfrac{1}{2}\,2^{-js}\} are disjoint. The family of densities fε∈Aj′f_{\varepsilon}\in{\mathscr{A}}^{\prime}_{j} given by the Varshanov-Gilbert Lemma has cardinality m≃2c′2jdm\simeq 2^{c^{\prime}2^{jd}} and by (39) and (38)

We apply now Fano’s lemma to the probabilities Pε{\rm P}_{\varepsilon} that are the nn times product of fε dxf_{\varepsilon}\,dx and to the events Aε={∥f^−fε′∥p≥12 2−js}A_{\varepsilon}=\{\|\widehat{f}-f_{\varepsilon^{\prime}}\|_{p}\geq\tfrac{1}{2}\,2^{-js}\}. It is well known that

Now let jj be so that n2−2js≃2jdn2^{-2js}\simeq 2^{jd}, that is 2j≃n12s+d2^{j}\simeq n^{\frac{1}{2s+d}}. With this choice one has

∙\bullet We prove now that the minimax LpL^{p}-loss is ≳n−p(s+d(1p−1r))2(s+d(12−1r))\gtrsim n^{-\frac{p(s+d(\frac{1}{p}-\frac{1}{r}))}{2(s+d(\frac{1}{2}-\frac{1}{r}))}}. Let us consider the two densities

with γ\gamma such that the above are positive (∣γ∣≲2−jd/2|\gamma|\lesssim 2^{-jd/2} is enough). If ∣γ∣≤2−js2−jd(12−1r)M|\gamma|\leq 2^{-js}2^{-jd(\frac{1}{2}-\frac{1}{r})}M, then thanks to (15) both f0f_{0} and f1f_{1} belong to the ball Br,qs(M)B^{s}_{r,q}(M). Remark that this condition implies ∣γ∣≲2−jd/2|\gamma|\lesssim 2^{-jd/2}, as we assume s≥drs\geq\frac{d}{r}. Also

so that,if we denote by P0,P1{\rm P}_{0},P_{1} the nn-times product of f0 dxf_{0}\,dx and f1 dxf_{1}\,dx by itself respectively, K(P0,P1)≈nγ2K({\rm P}_{0},{\rm P}_{1})\approx n\gamma^{2}. By (13) and Lemma 17 we have

We choose γ=1n=2−j(s+d(12−1r))\gamma=\frac{1}{\sqrt{n}}=2^{-j(s+d(\frac{1}{2}-\frac{1}{r}))}, so that K(P0,P1)≈nK({\rm P}_{0},{\rm P}_{1})\approx n. Moreover with this choice of nn, j≈log⁡n((2(s+d(12−1r)))−1j\approx\log n((2(s+d(\frac{1}{2}-\frac{1}{r})))^{-1}, so that again by Fano’s lemma,

Thanks to (39) the events {∥f^−fi∥p≥δ}\{\|\hat{f}-f_{i}\|_{p}\geq\delta\} are disjoint if δ≲2−j(s+d(1p−1r))\delta\lesssim 2^{-j(s+d(\frac{1}{p}-\frac{1}{r}))}. Therefore by Fano’s lemma

Putting things together and checking for which values of the parameters one rate is larger than the other one concludes the proof of Theorem 10. Note that, as s>drs>\frac{d}{r}, if 1≤p≤21\leq p\leq 2