The inverse conjecture for the Gowers norm over finite fields via the correspondence principle

Terence Tao, Tamar Ziegler

Introduction

thus ∥f∥Ud+1(V)\|f\|_{U^{d+1}(V)} measures the average bias in dth⁡d^{\operatorname{th}} multiplicative derivatives of ff. We also define the weak Gowers norm ∥f∥ud(V)\|f\|_{u^{d}(V)} of ff to be the quantity

thus ∥f∥ud(V)\|f\|_{u^{d}(V)} measures the extent to which ff can correlate with a phase polynomial of degree at most d−1d-1.

where ∣x∣|x| denotes the number of indices 1⩽j⩽n1\leqslant j\leqslant n for which xj=1x_{j}=1, lies in P3(V){\mathcal{P}}_{3}(V) and has a large inner product with ff; indeed, since f(x)=+1f(x)=+1 when ∣x∣=0,1,2,3mod  8|x|=0,1,2,3\mod 8 and −1-1 otherwise, we easily check that

In particular, we see that ∥(−1)S4∥u4(V)\|(-1)^{S_{4}}\|_{u^{4}(V)} is bounded from below by a positive absolute constant for large nn.

The main result of this paper is to establish this conjecture in the high characteristic case.

In the low characteristic case we have a partial result:

One could in principle make the quantity k=C(d)k=C(d) in Theorem 1.10 explicit, but this would require analyzing the arguments in in careful detail. One should however be able to obtain reasonable values of kk for small dd (e.g. d=4d=4).

The proofs of Theorems 1.9, 1.10 rely on four additional ingredients:

The Furstenberg correspondence principle, combined with the random averaging trick of Varnavides;

A statistical sampling lemma (Proposition 3.13); and

Local testability of phase polynomials (Lemma 4.5), essentially established in .

Of these ingredients, the ergodic inverse theorem is the most crucial, and we now pause to describe it in detail.

12. The ergodic inverse conjecture in finite characteristic

By setting h1=…=hd+1=0h_{1}=\ldots=h_{d+1}=0 we see that every phase polynomial ϕ∈Pd(X)\phi\in{\mathcal{P}}_{d}(\mathbf{X}) has unit magnitude: ∣ϕ∣=1|\phi|=1 μ\mu-a.e..

If d>1d>1, then ∥ϕ∥Ud(X):=lim sup⁡n→∞(∥Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣hϕ∥Ud−1(X,μ,T)2d−1)1/2d\|\phi\|_{U^{d}(\mathbf{X})}:=\limsup_{n\to\infty}\left(\|{\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{h}\phi\|_{U^{d-1}(X,\mu,T)}^{2^{d-1}}\right)^{1/2^{d}}.

We also define the weak Gowers-Host-Kra seminorm ∥ϕ∥ud(X)\|\phi\|_{u^{d}(\mathbf{X})} as

If ϕ∈Pd−1(X)\phi\in{\mathcal{P}}_{d-1}(\mathbf{X}) is a phase polynomial of degree at most d−1d-1, then ∥ϕ∥Ud(X)=∥ϕ∥ud(X)=1\|\phi\|_{U^{d}(\mathbf{X})}=\|\phi\|_{u^{d}(\mathbf{X})}=1.

One can use the ergodic theorem to show that the limits here in fact converge, but we will not need this. The UdU^{d} are indeed seminorms, but we will not need this either.

In [2, Corollaries 1.26,1.27], the following ergodic theory analogues of Theorems 1.9, 1.10 was shown:

We will use Theorem 1.20 as a “black box”, and it will be the primary ingredient in our proof of Theorem 1.10, in much the same way that the Furstenberg recurrence theorem is the primary ingredient in Furstenberg’s proof of Szemerédi’s theorem in . Theorem 1.19 plays a similar role for Theorem 1.9.

As with any other argument using a Furstenberg-type correspondence principle, our bounds are ineffective, in that we do not obtain an explicit value of ε\varepsilon in terms of dd and δ\delta. In principle, one could finitise the arguments in (in the spirit of ) to obtain such an explicit value, but this would be extremely tedious (and not entirely straightforward), and would lead to an extremely poor dependence (such as iterated tower-exponential or worse). We will not pursue this matter here.

22. Acknowledgments

The first author is supported by a grant from the MacArthur Foundation, and by NSF grant CCF-0649473. The second author is supported by ISF grant 557/08, by a Landau fellowship - supported by the Taub foundations, and by an Alon fellowship. The authors are also greatly indebted to Ben Green for helpful conversations, and Vitaly Bergelson for encouragement.

Notation

We will rely heavily on asymptotic notation. Given any parameters x1,…,xkx_{1},\ldots,x_{k}, we use Ox1,…,xk(X)O_{x_{1},\ldots,x_{k}}(X) to denote any quantity bounded in magnitude by Cx1,…,xkXC_{x_{1},\ldots,x_{k}}X for some finite quantity Cx1,…,xkC_{x_{1},\ldots,x_{k}} depending only on x1,…,xkx_{1},\ldots,x_{k}. We also write Y≪x1,…,xkXY\ll_{x_{1},\ldots,x_{k}}X or X≫x1,…,xkYX\gg_{x_{1},\ldots,x_{k}}Y for Y=Ox1,…,xk(X)Y=O_{x_{1},\ldots,x_{k}}(X). Furthermore, given an asymptotic parameter nn that can go to infinity, we use on→∞;x1,…,xk(X)o_{n\to\infty;x_{1},\ldots,x_{k}}(X) to denote any quantity bounded in magnitude by cx1,…,xk(n)Xc_{x_{1},\ldots,x_{k}}(n)X, where cx1,…,xk(n)c_{x_{1},\ldots,x_{k}}(n) is a quantity which goes to zero as n→∞n\to\infty for fixed x1,…,xkx_{1},\ldots,x_{k}. Thus for instance, if r2>r1>1r_{2}>r_{1}>1, then exp⁡(r1)log⁡r2=or2→∞;r1(1)\frac{\exp(r_{1})}{\log r_{2}}=o_{r_{2}\to\infty;r_{1}}(1).

Statistical sampling

The point here is that the error term is uniform in the choice of ff and VV.

We now record some variants of this standard “random local averages approximate global averages” fact, in which we perform more exotic empirical averages. We begin with averages along random subspaces of VV.

Let VV be a finite-dimensional vector space, and let f:V→Df:V\to\mathcal{D} be a function. Let v1,…,vm∈Vv_{1},\ldots,v_{m}\in V be chosen independently at random. Then with probability 1−om→∞(1)1-o_{m\to\infty}(1), we have

where v⃗:=(v1,…,vm)\vec{v}:=(v_{1},\ldots,v_{m}), and a⃗⋅v⃗:=a1v1+…+amvm\vec{a}\cdot\vec{v}:=a_{1}v_{1}+\ldots+a_{m}v_{m}.

One can easily make the om→∞(1)o_{m\to\infty}(1) terms more explicit, but we will not need to do so here.

We use the second moment method. Note that

(the om→∞(1)o_{m\to\infty}(1) error arising from the a=0a=0 contribution) so by Chebyshev’s inequality it suffices to show that

In the above lemma, ff was deterministic and thus independent of the viv_{i}. But we can easily extend the result to the case where ff depends on a bounded number of the viv_{i}:

Let VV be a finite-dimensional vector space, let m⩾m0⩾0m\geqslant m_{0}\geqslant 0, let v1,…,vm∈Vv_{1},\ldots,v_{m}\in V be chosen independently at random, and let fv1,…,vm0:V→Df_{v_{1},\ldots,v_{m_{0}}}:V\to\mathcal{D} be a function that depends on v1,…,vm0v_{1},\ldots,v_{m_{0}} but is independent of vm0+1,…,vmv_{m_{0}+1},\ldots,v_{m}. Then with probability 1−om→∞;m0(1)1-o_{m\to\infty;m_{0}}(1), we have

with probability 1−om−m0→∞(1)1-o_{m-m_{0}\to\infty}(1) conditioning on v⃗0\vec{v}_{0}; integrating this we see that the same is true without the conditioning. We can shift hh by a⃗0⋅v⃗0\vec{a}_{0}\cdot\vec{v}_{0}, move the hh average onto the other side, and take expectations to conclude that

for each a⃗0\vec{a}_{0}; averaging over a⃗0\vec{a}_{0} by the triangle inequality we obtain the claim. ∎

We will need to generalise these results further by considering more exotic averages along cubes. A typical result we will need can be stated informally as

when m1m_{1} is large, m2m_{2} is large compared with m1m_{1}, and v⃗\vec{v} is random (see Lemma 3.9 for the formal version of this type of estimate). Such results follow (heuristically, at least), by iterating the previous results. For instance, from Corollary 3.3 we heuristically have

when m2m_{2} is large compared to m1m_{1} and then interchanging the expectations and applying Lemma 3.1 heuristically yields

when m1m_{1} is large, thus giving (3.1).

We will formalise the precise statement along these lines that we need later in this section. We begin with some key definitions.

Let k⩾1k\geqslant 1, let VV be a finite-dimensional vector space, let f:V→Df:V\to\mathcal{D} be a bounded function, and let

be a sequence of integers (or “scales”). We define an accurate sampling sequence for ff of degree kk and at scales H1,H2,…H_{1},H_{2},\ldots to be an infinite sequence of vectors

The denominator r1r_{1} in (3.2) could be replaced by any other fixed function of r1r_{1} that went to infinity as r1→∞r_{1}\to\infty if desired here.

Roughly speaking, an accurate sampling sequence will allow us to estimate all the global averages that we need for the combinatorial inverse conjecture for the Gowers norm by local averages which are suitable for lifting to the ergodic setting via the correspondence principle. We illustrate the use of such sequences by describing the three special cases of (3.2) that we will actually need in our arguments.

Let d⩾1d\geqslant 1, let VV be a finite-dimensional vector space, let f:V→Df:V\to\mathcal{D} be a bounded function, and let v1,v2,…∈Vv_{1},v_{2},\ldots\in V be an accurate sampling sequence for ff of degree dd and at scales H1,H2,…H_{1},H_{2},\ldots. Then for every sequence of scales

As with all other estimates in this section, the point is that the error term is uniform over all choices of ff and VV. Note that the d=2d=2 case of this lemma is a formalisation of (3.1).

where C:z↦z‾{\mathcal{C}}:z\mapsto\overline{z} is the complex conjugation operator. A routine computation gives the identities

Also, it is easy to see that the Lipschitz norm ∥G∥Lip⁡\|G\|_{\operatorname{Lip}} is Od(1)O_{d}(1). The claim now follows immediately from (3.2) and the triangle inequality. ∎

A routine computation gives the identities

Also, it is clear that GG is Lipschitz with norm OF,r0(1)O_{F,r_{0}}(1). The claim then follows from (3.2). ∎

where C{\mathcal{C}} is again the complex conjugation operator. A routine computation gives the identities

for any r0<r1<…<rkr_{0}<r_{1}<\ldots<r_{k}. Also it is clear that GG is Lipschitz with norm OF,r0,k(1)O_{F,r_{0},k}(1). The claim then follows from (3.2) and the triangle inequality. ∎

Of course, in order to utilise the above lemmas we need to know that such accurate sampling sequences in fact exist. This is the purpose of the following proposition.

Let d⩾1d\geqslant 1. Then there exists a sequence

of integers such that for every finite-dimensional vector space VV and any function f:V→Df:V\to\mathcal{D}, there exists an accurate sampling sequence v1,v2,v3,…∈Vv_{1},v_{2},v_{3},\ldots\in V for ff of degree dd at scales H1,H2,…H_{1},H_{2},\ldots.

The key point here is that the scales H1,H2,H3,…H_{1},H_{2},H_{3},\ldots are universal; they depend on dd, but otherwise and work for all vector spaces VV and functions ff.

We use the probabilistic method, choosing v1,v2,…∈Vv_{1},v_{2},\ldots\in V uniformly at random, and showing that (if FF was sufficiently rapid) the resulting sequence will be an accurate sampling sequence with positive probability.

We begin with observing that in order to verify the condition (3.2), it suffices by the triangle inequality to show that with positive probability, one has

of (3.4) holds with probability 1−oHrd′→∞;d,Hr0,…,Hrd′−1,r1(1)1-o_{H_{r_{d^{\prime}}}\to\infty;d,H_{r_{0}},\ldots,H_{r_{d^{\prime}-1}},r_{1}}(1).

Fix d′,r0,…,rd′,Gd^{\prime},r_{0},\ldots,r_{d^{\prime}},G. By Markov’s inequality, it suffices to show that

by linearity of expectation it thus suffices to show that

where fv1,…,vHrd′−1:V→Df_{v_{1},\ldots,v_{H_{r_{d^{\prime}-1}}}}:V\to\mathcal{D} is the function

As the notation suggests, the function fv1,…,vHrd′−1f_{v_{1},\ldots,v_{H_{r_{d^{\prime}-1}}}} depends on the values of v1,…,vHrd′−1v_{1},\ldots,v_{H_{r_{d^{\prime}-1}}} but not on higher elements of the sequence. Also, as GG has Lipschitz norm 11, ff takes values in D\mathcal{D}. The claim now follows from Corollary 3.3. ∎

Proof of main theorems

We are now ready to prove the main theorems. We shall just prove Theorem 1.10 using Theorem 1.20; the deduction of Theorem 1.9 using Theorem 1.19 is exactly analogous (see the brief remarks at the end of this section).

be the sequence in Proposition 3.13; it is important to note that this sequence does not depend on nn. From that proposition, we can find an accurate sampling sequence

for f(n)f^{(n)} of degree kk at these scales. We fix such a sequence for each nn.

Because XX is compact metrisable, and the action of TT is continuous it is a well-known fact that Pr⁡(X)T\Pr(X)^{T} is sequentially compact; thus every sequence of measures in Pr⁡(X)T\Pr(X)^{T} has a vaguely convergent subsequence whose limit is also in Pr⁡(X)T\Pr(X)^{T}.

For each nn, we define a measure μ(n)∈Pr⁡(X)T\mu^{(n)}\in\Pr(X)^{T} on XX by the formula

where δ\delta denotes the Dirac mass and for each x∈V(n)x\in V^{(n)}, ζn,x∈X\zeta_{n,x}\in X is the function

Let f:X→Df:X\to\mathcal{D} be the indicator function f(ζ):=ζ(0)f(\zeta):=\zeta(0). We observe the key correspondence

For continuous ϕ\phi, the claim follows easily from the Stone-Weierstrass theorem (and in this case we can upgrade the L1L^{1} approximation to L∞L^{\infty} approximation). As XX is compact metrisable, the Borel measure μ\mu is in fact a Radon measure, and so (by Urysohn’s lemma) the continuous functions are dense in L∞(X)L^{\infty}(\mathbf{X}) in the L1(X)L^{1}(\mathbf{X}) topology, and the claim follows. ∎

We can now use the machinery of the previous section to deduce various important facts about X\mathbf{X} and ff. For instance, Lemma 3.11 now implies

By the mean ergodic theorem, it suffices to show that

for all g∈L∞(X)g\in L^{\infty}(X). By Lemma 4.2 and a standard limiting argument it suffices to show this for gg which are functions of finitely many shifts of ff, say g=G(Tb⃗1f,…,Tb⃗kf)g=G(T_{\vec{b}_{1}}f,\ldots,T_{\vec{b}_{k}}f). We will then show that

By vague convergence it suffices to show that

for all nn. By (4.3), we can rewrite the left-hand side as

But the claim now follows from Lemma 3.11 (and Remark 3.8). ∎

We have ∥f∥Ud(X)⩾δ\|f\|_{U^{d}(\mathbf{X})}\geqslant\delta.

By reversing the order of averages, it suffices to show that

Fix r1,…,rdr_{1},\ldots,r_{d}. By weak convergence, it suffices to show that

for all nn. By (4.1), it suffices to show that

By (4.3), left-hand side can be rephrased as

and the claim now follows from Lemma 3.9 (and Remark 3.8). ∎

We have now verified all the hypotheses of Theorem 1.19. Applying that theorem, we conclude that ∥f∥uk(X)>c\|f\|_{u^{k}(\mathbf{X})}>c for some c>0c>0 (which could be very small, but positive). Thus we can find a phase polynomial ϕ∈Pk−1(X)\phi\in{\mathcal{P}}_{k-1}(\mathbf{X}) of degree k−1k-1 such that

Since ϕ\phi takes values in D\mathcal{D}, we may assume without loss of generality that GG does also. If ε\varepsilon is small enough depending on cc, we thus have

for all sufficiently large nn (depending on G,m,cG,m,c). Using (4.3), we rearrange this as

Now let r1r_{1} be a large integer depending on the b⃗1,…,b⃗m,ε\vec{b}_{1},\ldots,\vec{b}_{m},\varepsilon, and let rj:=r1+(j−1)r_{j}:=r_{1}+(j-1) for j=2,…,dj=2,\ldots,d. Since ϕ\phi is a phase polynomial of degree k−1k-1, we have

for all sufficiently large nn (depending on ε,Hr1,…,Hrk\varepsilon,H_{r_{1}},\ldots,H_{r_{k}}). Using (4.3), we can rearrange the left-hand side as

Applying Lemma 3.12 we conclude (if r1r_{1} is sufficiently large depending on b⃗1,…,b⃗m,ε\vec{b}_{1},\ldots,\vec{b}_{m},\varepsilon) that

Let VV be a finite-dimensional vector space, let k⩾1k\geqslant 1, let g:V→Dg:V\to\mathcal{D} be a bounded function, and suppose that

for some ε>0\varepsilon>0. Then there exists a phase polynomial ϕ∈Pk−1(V)\phi\in{\mathcal{P}}_{k-1}(V) such that

Applying this lemma, we conclude that there exists ϕ(n)∈Pk−1(V(n))\phi^{(n)}\in{\mathcal{P}}_{k-1}(V^{(n)}) such that

Inserting this into (4.5) we conclude that

if ε\varepsilon is sufficiently small depending on c,kc,k. But this contradicts (4.2). The proof of Theorem 1.10 is complete.

The proof of Theorem 1.9 is identical, but with kk now set equal to dd, and Theorem 1.19 used instead of Theorem 1.20. We leave the details to the reader.

Appendix A Proof of Lemma 4.5

In this appendix we give a proof of Lemma 4.5, following the arguments in and [20, Proposition 4.6]. We begin with a variant of Lemma 1.2:

We induct on kk. For k=0k=0 the claim is obvious, and for k=1k=1 ϕ\phi is a linear character (times a phase) and the claim can be worked out by hand. Now suppose k⩾2k\geqslant 2 and the claim has already been shown for smaller values of kk. Since ϕ\phi is a phase polynomial, we have Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣0…Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣0ϕ=1{\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{0}\ldots{\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{0}\phi=1, and thus ϕ\phi has unit magnitude. Observe that if ∫V∣ϕ−1∣⩽ε\int_{V}|\phi-1|\leqslant\varepsilon, then ∫V∣Thϕ−1∣⩽ε\int_{V}|T_{h}\phi-1|\leqslant\varepsilon for every h∈Vh\in V. Using the elementary estimate

(using the fact that ϕ\phi has unit magnitude) we conclude that

for every h∈Vh\in V. On the other hand, Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣hϕ∈Pk−1(V){\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{h}\phi\in{\mathcal{P}}_{k-1}(V), so by induction hypothesis (if ε\varepsilon is small enough) we conclude that Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣hϕ{\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{h}\phi is constant for all h∈Vh\in V. Thus ϕ∈P1(V)\phi\in{\mathcal{P}}_{1}(V), but then the claim follows from the k−1k-1 case. ∎

We now prove Lemma 4.5. The case k=1k=1 is easy, so suppose that k⩾2k\geqslant 2 and the claim has already been established for k−1k-1. To abbreviate the notation we shall write o(1)o(1) for oε→0;k(1)o_{\varepsilon\to 0;k}(1). We say that a statement P(x)P(x) holds for most x∈Vx\in V if it holds for (1−o(1))∣V∣(1-o(1))|V| elements of vv.

We fix k,V,fk,V,f. We may assume that ε\varepsilon is small depending on dd, as the claim is trivial otherwise. From (4.6) and Markov’s inequality we see that

for most h∈Vh\in V. Let us call hh good if (A.1) holds. Applying the induction hypothesis, we conclude that for any good hh there exists This quantity plays the same role that cocycles do in ergodic theory. ϕh∈Pk−2(V)\phi_{h}\in{\mathcal{P}}_{k-2}(V) such that

In particular, this implies (by Markov’s inequality) that for all good hh, we have

for most VV. Since ff is bounded in magnitude by 11, this implies that

for most xx, and for all good hh we have

for at least one xx. On the other hand, from Lemma A.1 ϕh\phi_{h} takes values in e2πiθe^{2\pi i\theta} times Kth⁡K^{\operatorname{th}} roots of unity for some fixed KK depending only on d,pd,p. Thus e2πipθe^{2\pi ip\theta} times a Kth⁡K^{\operatorname{th}} root of unity is within o(1)o(1) of 11, and so e2πiθe^{2\pi i\theta} lies within o(1)o(1) of a pKth⁡pK^{\operatorname{th}} root of unity. Rotating ϕh\phi_{h} by o(1)o(1) if necessary we may assume that e2πiθe^{2\pi i\theta} is exactly a pKth⁡pK^{\operatorname{th}} root of unity, and in particular we have

Now suppose that h1,h2,h3,h4h_{1},h_{2},h_{3},h_{4} are good and form an additive quadruple in the sense that h1+h2=h3+h4h_{1}+h_{2}=h_{3}+h_{4}. Then from (A.2) we see that

for most xx. Since ∣f(x)∣=1+o(1)|f(x)|=1+o(1) for most xx, we conclude the approximate cocycle relationship

for most xx. In particular, the average of the left-hand side in xx is 1−o(1)1-o(1). Applying Lemma A.2 (and assuming ε\varepsilon small enough), we conclude that the left-hand side is constant in xx; using the discretisation (A.3), we conclude (again for ε\varepsilon small enough) that it is in fact 11. Thus

for all xx and any good additive quadruple h1,h2,h3,h4h_{1},h_{2},h_{3},h_{4}.

whenever h1,h2,h1+h2h_{1},h_{2},h_{1}+h_{2} are simultaenously good. Note that the existence of such an h1,h2h_{1},h_{2} is guaranteed since most hh are good, and (A.5) ensures that the right-hand side of (A.6) does not depend on the exact choice of h1,h2h_{1},h_{2} and so ψ\psi is well-defined. From (A.3) we see that ψ\psi takes values in the pKth⁡pK^{\operatorname{th}} roots of unity, and in particular only has O(1)O(1) possible values.

Now let x∈Vx\in V and hh be good. Then, since most elements of VV are good, we can find good r1,r2,s1,s2r_{1},r_{2},s_{1},s_{2} such that r1+r2=xr_{1}+r_{2}=x and s1+s2=x+hs_{1}+s_{2}=x+h. From (A.4) we see that

for most yy. Combining these (and the fact that ∣f(y)∣=1+o(1)|f(y)|=1+o(1) for most yy) we see that

for most yy. Taking expectations and applying Lemma A.2 and (A.3) as before, we conclude that

for all yy. Specialising to y=0y=0 and applying (A.6) we conclude that

for all x∈Vx\in V and good hh; thus we have succesfully “integrated” ϕh\phi_{h}. We can then extend ϕh(x)\phi_{h}(x) to all h∈Vh\in V (not just good hh) by viewing (A.7) as a definition. Observe that if h∈Vh\in V, then h=h1+h2h=h_{1}+h_{2} for some good h1,h2h_{1},h_{2}, and from (A.7) we have

In particular, since the right-hand side lies in Pk−2(V){\mathcal{P}}_{k-2}(V), the left-hand side does also. Thus we see that Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣hψ∈Pk−2(V){\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{h}\psi\in{\mathcal{P}}_{k-2}(V) for all h∈Vh\in V, and thus Q∈Pk−1(V)Q\in{\mathcal{P}}_{k-1}(V). If we then set g(x):=f(x)ψ‾(x)g(x):=f(x)\overline{\psi}(x), then from (A.2), (A.7) we see that for every h∈Hh\in H we have

for most xx, and Lemma 4.5 then follows.

References