An inverse theorem for the uniformity seminorms associated with the action of $F^ω$
Vitaly Bergelson, Terence Tao, Tamar Ziegler
Introduction
This paper is concerned with the structural theory of measure preserving actions of abelian groups. We begin with some general definitions.
Most of our analysis will take place in the setting of ergodic systems, but for various technical reasons we will sometimes have to work with non-ergodic systems. In some (but not all) cases, results on ergodic systems can be extended satisfactorily to the non-ergodic case using the ergodic decomposition. The hypothesis of separability is a technical one (used in particular in Appendix C to obtain a certain measurability property), but can often be removed in applications by restricting the -algebra to the sub-algebra generated by the functions one is interested in studying, together with all of their shifts.
This paper is concerned with the following seminorms for -systems:
2. Universal characteristic factors
A fundamental concept in the study of the Gowers-Host-Kra uniformity seminorms is that of the universal characteristic factor for such norms. To describe this concept we need some notation.
Observe that any factor of an ergodic -system is also ergodic. The converse, of course, is not true.
The uniqueness of is clear; the existence follows immediately from Lemma A.32. ∎
The following observations are immediate:
3. Main result
From Theorems 1.20, 1.19 and Proposition 1.10 we have the following immediate corollaries:
In a companion paper , we will combine Corollary 1.22 with a version of the Furstenberg correspondence principle, as well as the equidistribution theory in , to obtain a finitary counterpart to this theorem:
Theorem 1.19 should also allow one (assuming sufficiently high characteristic) to obtain a formula for the limit of multiple ergodic averages of quantities such as (as in ), and to be able to show that can be approximated by a function of polynomials in , in the spirit of the results in . We hope to report on these and other applications in a subsequent paper.
4. The Heisenberg example
To illustrate the above results we now pause to describe the model case of a Heisenberg system. (The discussion in this section is not directly used in the remainder of the paper.) To simplify the discussion we restrict attention to the case.
5. Overview of the structure of the paper
for some functions . A key technical point is that while the function is a priori only measurable in , it can be made to be measurable in the parameters also (see Lemma C.4). This will be rather important for us as we will be relying quite heavily on the measurability property For instance, we will need a variant of the classical Steinhaus theorem that asserts that if is a measurable subset of a compact abelian group with positive measure, then the difference set contains a neighborhood of the origin (cf. Lemma D.1). Curiously, analogous results are exploited in the additive-combinatorial approach to the Gowers inverse problem (see e.g. ), where they go by the name of “Bogolyubov-type lemmas”. in our arguments.
6. Acknowledgements
The authors would like to thank Tim Austin, Ben Green, Bernard Host, Bryna Kra, and Trevor Wooley for many enlightening conversations and suggestions. The first author is supported by NSF grant DMS-0600042. The second author is supported by a grant from the MacArthur Foundation, and by NSF grant DMS-0649473. The first and third authors are supported by BSF grant No. 2006094. The third author is supported by a Landau fellowship of the Taub foundations, and by an Alon fellowship. The authors also thank the anonymous referees for many useful suggestions and corrections.
Abelian cohomology
Throughout the paper we will be relying heavily on the language of abelian cohomology of dynamical systems. We record the key definitions here; further discussion of these concepts can be found in Appendix B.
We caution that abelian extensions of ergodic -systems are not necessarily ergodic.
We will need several technical results concerning abelian cohomology groups, which we have collected in Appendix B, and which we will refer to as necessary in the main text of the paper.
Reduction to abelian extensions of order <k+1<k+1
By Theorem A.35, every Abramov system of order is also a -system of order . This and (A.9), will allow us to immediately derive Theorem 1.20 from the following claim:
The following basic fact was established by Host and Kra:
Using Proposition 3.4 and Fourier analysis, we can reduce matters to studying projections of abelian cocycles to the unit circle. More precisely, in future sections we will show the following result.
It remains to prove Theorem 3.8. This will be the objective of the next few sections.
Functions of type <k<k
We make a further reduction, introducing the useful notion This concept is essentially that of a cocycle of type from , but generalized to non-cocycles and to more general group actions. We have replaced “” by “” as such cocycles will have “degree” strictly less than in some sense. of a function of type .
where .
We make some easy observations (see Figure 2):
From (i) we see in particular that coboundaries, being of type , are of type . The claims in (ii) are then easily verified.
To show (iii), we induct on . The claim is easy for (using ergodicity), so suppose that and the claim has already been shown for .
The relevance of the type concept to us lies in the important observation that abelian extensions of order arise from functions of type :
We will prove Theorem 4.5 in later sections. For now, let us observe that we can use Theorem 1.20 to obtain some structural control on systems of order . We say that a group is -torsion for some if we have for all .
Our remaining task is to prove Theorem 4.5.
Reduction to solving a Conze-Lesigne type equation
To prove Theorem 4.5 we will use two lemmas to reduce matters to solving a certain equation of Conze-Lesigne type. The first lemma allows one to descend a type condition on an extension to a type condition on a base, worsening the type if necessary:
For cocycles, a more general (and stronger) statement appears in [18, Corollary 7.8]. However, for technical reasons, it is necessary for us to work with more general functions than just cocycles. (But see Corollary 8.11 below.)
If we expand into a Fourier series and compare Fourier coefficients, we conclude
almost everywhere on any ergodic component. In particular, is invariant and thus constant a.e. on any ergodic component. We extend arbitrarily to a character .
We now claim that is also a -coboundary for every positive side transformation . But a computation shows that
Now we crucially use the fact that and are cocycles to write this as
The second lemma allows us to reduce the type of a function by differentiation in the vertical direction.
Because of these two lemmas, Theorem 4.5 will follow from
We claim for each that
Theorem 4.5 now follows by specializing (5.1) to the case . ∎
Reduction to a finite UU
The purpose of this section is to obtain the following reduction.
In order to prove Theorem 5.4, it suffices to do so in the case when is finite.
By Lemma C.4 we can take to be measurable with respect to .
The next step is to linearize on an open subgroup of , by arguing as follows. Let . Then the cocycle identity and (6.1) give
Now let . From (6.2) we have the second-order cocycle identity for all . If is small enough, we can find so that . We conclude that
for all , while from (6.2), (6.3) we have and thus (by (6.4))
Now we make a crucial use of the finite characteristic hypothesis. By Lemma 4.7, is -torsion for some . By Lemma D.1, we conclude that contains an open subgroup of ; by reducing if necessary, we may assume that is in fact equal to an open subgroup.
The finite group case
for some and some finite (but unbounded) . When is large enough, depending on , we can take .
The idea here is to express as times a polynomial error, for some independent of ; we will then “integrate” this to express as times a polynomial error, times a function invariant under ; this is basically what we need to establish Theorem 5.4.
We turn to the details. The first task is to measure two potential obstructions to being expressible as (modulo polynomial errors), namely the obstruction coming from the torsion of , and the obstruction coming from the multi-dimensionality of .
Observe that we have the telescoping identity
Note, conversely, that needed to be polynomial in order to have any chance to express as times a polynomial.
Suppose now that is sufficiently large depending on , so that . By Lemma D.3 we have , and hence by (7.2), (7.1) we may now strengthen (7.3) to
In a similar fashion, from the commutation identity for any , we see from Lemma B.5(i) and (7.1) that
Again, observe that had to be polynomial in order to have a chance to express , as , modulo polynomials.
Now we compute a derivative of . We clearly have
for any . On the other hand, we have the telescoping identity
When , this claim follows from Corollary D.6, so suppose now that is sufficiently large depending on . Then is a phase polynomial of degree in that takes values in . By Taylor expansion we may thus write
Inserting the above claim into (7.8) we conclude that
Thus if we set , then is cohomologous to and
Now we need to work on the term. From the telescoping identity and (7.9), we have
pushing forward by and then using (3.1), we conclude
In the case that is sufficiently large depending on , we see from Lemma D.3 that we have the improvement , and so is constant in this case.
Inserting this claim back into (7.9) we obtain
From Proposition 7.1 and Proposition 6.1, we obtain Theorem 5.4, and thus Theorem 3.3.
The high characteristic case
We now develop high characteristic analogues of the above theory, establishing the sharp Theorem 1.19 instead of Theorem 1.20 in this setting. The arguments here will be similar to those used to prove Theorem 1.20; the main new difficulty is to be careful to not lose anything in the degree of various functions beyond what is absolutely necessary.
Just as Theorem 1.20 follows from Theorem 4.5, Theorem 1.19 will follow from
It is clear that Theorem 8.1 then follows from the case of
2. Vertical differentiation
is a line cocycle, is of type , and is a quasi-cocycle of order .
3. Reduction to the finite UU case
We now argue (as in Proposition 6.1) that in order to conclude the proof of Theorem 8.6, it suffices to do so in the case when the vertical structure group is finite.
We now invoke the following variant of Lemma B.6 which is more efficient with the degree.
In the converse direction, one can show (by using the properties of the nilpotent group studied in ) that if has degree , then has degree . This may explain the terminology “exact”.
The remaining tasks (in the low order case ) are to verify Theorem 8.6 in the case of finite , and to also verify Proposition 8.9 and Proposition 8.11.
4. The finite group case
We now establish Theorem 8.6 in the case when is finite. This is the analogue of Proposition 7.1, but our arguments here are somewhat simpler thanks to the high characteristic (which allows us to use the full power of Lemma D.3).
By repeated application of Lemma 8.8, we know that has degree with respect to differentiation in the direction. By Lemma D.3, we conclude that , and thus the form a cocycle in the variable, in the sense that
5. The high order case
As in previous arguments, we first reduce to the case when is finite, and then establish the finite case.
Our remaining tasks are to prove Proposition 8.9, Proposition 8.11, and Lemma 8.13.
6. Exact integration
In this subsection we establish Proposition 8.9. We begin with the analogue of Lemma B.5.
We prove (i) and (ii) simultaneously by induction on . If , then is a constant (by ergodicity), and so the map is a homomorphism, and the claim (i) then easily follows. Also in this case (ii) is clearly equivalent to (i).
Now suppose inductively that , and that the claim has already been proven for . We now induct on . When then is constant, and (i) is clear.
and thus (since the action of commutes with that of )
7. Exact descent
We now prove Proposition 8.11 and Lemma 8.13. As noted already in Remark 5.2, an exact descent result for cocycles already appears in [18, Corollary 7.8]; Proposition 8.11 can be viewed as an extension of that result to quasi-cocycles.
Our main tool for both of these tasks is the following equivalent characterization of the finite type condition.
where each is either equal to or its complex conjugate.
On the other hand, has magnitude . We conclude (for small enough) that
and thus by the pigeonhole principle that
for all . Averaging this over a Følner set , we conclude in particular that
for all . From the monotone convergence theorem we conclude that
We can now prove Proposition 8.11 and Lemma 8.13.
The proof of Theorem 1.19 is now complete.
Appendix A General theory of Gowers-Host-Kra seminorms
More generally, for any , define an -dimensional face or -face to be any set formed by intersecting distinct non-parallel sides. Thus has one -face, faces of dimension (i.e. the sides ), faces of dimension , and so forth down to faces of dimension (which are the vertices of the discrete cube).
Let be an -face. Enumerating the elements of in lexicographic order gives a natural bijection , which we call the coordinate map of , which maps the faces of , which are subsets of , to the faces of .
Let be an arbitrary set. We write for the set of functions . For each and each -face , we let be the pushforward map given by the formula for all . For any and sign , we abbreviate the cubic boundary map as .
For future reference we make the trivial observation that the map is a bijection between and .
Let be a (possibly non-abelian) group with identity , and let be a face of . For every , we let denote the group element whose components for are defined to equal when , and to equal otherwise. The map is a bijection from to the face group . When is a side (resp. a positive side) we refer to as a side group (resp. a positive side group); when is the entire cube we refer to as the diagonal group and denote it as , and abbreviate as . We also write (resp. ) for the subgroup of generated by all the side groups (resp. all the positive side groups).
For , the group is generated by
while the group is generated by
For future reference we observe that the side group is the group generated by the positive side group and the diagonal group .
Let be a group acting on a space by transformations for . Then acts on in the obvious manner, with the action of a group element mapping each point to . If is a face, we abbreviate the face transformation as , thus is an action of on . If is a side (resp. a positive side), we refer to as a side transformation (resp. positive side transformation), and if is the entire cube, we refer to as a diagonal transformation and abbreviate it further as .
For future reference we observe that all positive side maps preserve the coordinate of .
We recall the notion of a relative product:
Now we can introduce the cubic measure spaces of Host and Kra.
is the ergodic decomposition of with respect to the diagonal action of .
The cubic measures have a useful symmetry property:
The cubic measures also behave well with respect to passage to subcubes.
and is the -algebra consisting of subsets of which are invariant under the diagonal translations for . Thus the probability space is measure isomorphic to the space with the discrete -algebra and normalized counting measure, whilst is measure isomorphic to the space with the discrete -algebra and normalized counting measure.
A.2. Existence of the seminorms
The objective of this section is to establish that the Gowers-Host-Kra seminorms from Definition 1.3 are in fact well-defined, and to relate them to the cubic measures just constructed.
We prove this by induction on . For this follows from the mean ergodic theorem. Assume the induction hypothesis holds for . Then
Since is invariant with respect to the action of , by the ergodic theorem the above averages converge to
We record some basic properties of the Gowers-Host-Kra seminorms:
[18, Lemma 3.9] (See also [11, Lemmas 3.8, 3.9])
The measure is ergodic with respect to the action of .
For any measure-preserving transformation that commutes with the -action, and any side , the side transformation preserves .
For (i), see [18, Corollary 3.5]; for (ii), see [18, Corollary 3.6]. For (iii) and (iv), see [18, Lemma 5.5]. ∎
A.3. Dual functions
The limits above exist, as in Lemma A.18, by repeated applications of the ergodic theorem. Indeed, we easily verify that
where and the factor map is given by (i.e. the pushforward map ). As a consequence we have
and similarly (by repeated applications of the Cauchy-Schwarz inequality) that
The limit above is a repeated limit, but a posteriori, using Theorem 1.20, one can show that the double (simultaneous) limit exists as well, and both limits coincide, by modifying the proof of [18, Theorem 1.2], and similarly for higher values of . We omit the details.
where is the complex conjugation operator. Dual functions in this setting play an important role in the finitary theory of arithmetic progressions and similar patterns; see .
Note that Lemma A.32 immediately implies Proposition 1.10 in the introduction. From this lemma and (A.5) we also have
for (cf. [18, Corollary 4.4]).
From Lemma A.32 and Lemma A.22 one can show that universal characteristic factors are functorial:
Appendix B Abelian cohomology
The reader may wish to review the definitiosn in Definition 2.1 before proceeding with the rest of this section.
We begin with the following trivial but useful lemma:
Next, we recall that cohomology is trivial for free actions:
If acts freely on , then so does any compact abelian subgroup of .
There is an analogue of Lemma B.4 in the polynomial category. To state it, we first need a useful algebraic lemma.
We prove (i) by induction on . Indeed, the claim is trivial for by ergodicity, and for we have by induction that is a phase polynomial of degree for all , and thus by (3.1) is a phase polynomial of degree as claimed.
Finally, we prove claim (iv). The case follows from the previous claim, so suppose that and the claim has already been proven for the smaller values of . For , we take a derivative . One obtains essentially the same terms that appeared in the previous claim, plus (thanks to the cocycle equation (B.1)) some additional terms involving . But such terms can be dealt with by the induction hypothesis. ∎
thanks to (B.1). Thus we have for all .
We will also need another result in a similar spirit.
We will also take advantage of a useful splitting lemma.
Now we see how cohomology on an abelian extension relates to cohomology on the base space.
for all and almost every , . We rearrange this as
We perform a Fourier expansion in , obtaining
This completes the proof in the case when is trivial.
Appendix C A measurable selection lemma
By dividing by we may assume .
The claim is vacuous when . When we argue as follows. For any we have
If is a phase polynomial of degree , then is constant; if is non-constant, then (by ergodicity) is not identically for at least one . Thus and so , and the claim follows.
We are now ready to establish the measurable selection lemma.
One can also establish this result using a general measure selection result of Dixmier (see e.g. [2, Theorem 1.2.4]) together with Lusin’s theorem and Corollary C.3; we omit the details. One can also appeal to the descriptive set theory of Polish groups, see e.g. [18, Appendix A].
Appendix D Finite characteristic algebra
Recall that a group is -torsion if we have for all .
Let be a compact abelian -torsion group for some . Let be an open neighborhood of the identity in . Then contains an open subgroup of .
We will use a Fourier-analytic method. As is an open neighborhood of the origin, one can find another open neighborhood of the origin such that .
Let be the Haar measure on , then . Let be a small number (depending on ) to be chosen later. By Fourier analysis, we can approximate the indicator function to within in -norm by some linear combination of finitely many characters , where is finite but potentially unbounded. Since is -torsion, each character takes on at most values, with each level set of being a coset of an open subgroup of . If we let be the intersection of the kernels of all the , then is also an open subgroup of , and is constant on every coset of . Since approximates to within , we conclude (if is sufficiently small depending on ) that there exists a coset of on which has density greater than . But then this forces and hence , as desired. ∎
Let be a compact abelian -torsion group for some . Let be an open subgroup of . Then there exists a splitting , where is an open subgroup of , and is a finite abelian -torsion group.
It is known (see e.g. [23, Chapter 5, Theorem 18]) that a compact abelian -torsion group is topologically isomorphic to the direct product of cyclic -torsion groups In particular, the bounded torsion allows us to avoid having to deal with procyclic groups which are not direct products of cyclic groups.. Thus must contain a cylinder neighbourhood of the origin, i.e. a cofinite sub-product of these cyclic groups. Since one clearly has the desired splitting , the claim follows. ∎
D.2. Polynomials are discretely valued
Let . Since and , we conclude using the binomial formula that Since has degree , . We conclude that
Inverting the expression in brackets using Neumann series (and using the fact that annihilates ) we conclude that for any , thus by ergodicity is constant as claimed.
To prove (ii), we first observe that it suffices to prove the claim for of the form for integer . But the claim is trivial for , and from (i), we see that the claim for implies the claim for , and so (ii) follows by induction.
D.3. Roots of phase polynomials
Recall that a map into an additive group is a polynomial of degree if we have for all .
We may assume inductively that the claim is already proven for smaller values of ; for the same value of and smaller values of ; or the same value of and and smaller values of . We abbreviate as .
From primary school arithmetic we know that we have a formula of the form
for some “carry bit” functions . Applying this with and for some group element we conclude
Finally, by the induction hypothesis on , we know that
Putting all this together we see that is a polynomial of degree for all , and hence is a polynomial of degree , thus closing the induction. ∎
We isolate one special case of Corollary D.6:
By rotating by a constant and using Lemma D.3, we may assume that takes values in for some . If is not divisible by , then is invertible in and the claim is immediate, so it suffices to check the case when is a power of . But then the claim follows immediately from Corollary D.6. ∎
Another interesting consequence of Corollary D.6 (or Proposition D.5) is that phase polynomials can always be expressed in terms of -valued polynomials of higher degree.
Appendix E Connection with cubic complexes
In this appendix we point out some connections between the notions of polynomiality and type in this paper with the theory of cubic complexes as used in topology, as set out in , in analogy with the more well-known simplicial complexes used in that field. This material is not used elsewhere in this paper.