An inverse theorem for the uniformity seminorms associated with the action of $F^ω$

Vitaly Bergelson, Terence Tao, Tamar Ziegler

Introduction

This paper is concerned with the structural theory of measure preserving actions of abelian groups. We begin with some general definitions.

Most of our analysis will take place in the setting of ergodic systems, but for various technical reasons we will sometimes have to work with non-ergodic systems. In some (but not all) cases, results on ergodic systems can be extended satisfactorily to the non-ergodic case using the ergodic decomposition. The hypothesis of separability is a technical one (used in particular in Appendix C to obtain a certain measurability property), but can often be removed in applications by restricting the σ\sigma-algebra BX\mathcal{B}_{X} to the sub-algebra generated by the functions one is interested in studying, together with all of their shifts.

This paper is concerned with the following seminorms for GG-systems:

2. Universal characteristic factors

A fundamental concept in the study of the Gowers-Host-Kra uniformity seminorms is that of the universal characteristic factor for such norms. To describe this concept we need some notation.

Observe that any factor of an ergodic GG-system is also ergodic. The converse, of course, is not true.

The uniqueness of Z<k\mathcal{Z}_{<k} is clear; the existence follows immediately from Lemma A.32. ∎

The following observations are immediate:

3. Main result

From Theorems 1.20, 1.19 and Proposition 1.10 we have the following immediate corollaries:

In a companion paper , we will combine Corollary 1.22 with a version of the Furstenberg correspondence principle, as well as the equidistribution theory in , to obtain a finitary counterpart to this theorem:

Theorem 1.19 should also allow one (assuming sufficiently high characteristic) to obtain a formula for the limit of multiple ergodic averages of quantities such as c(g):=μ(A∩TgA∩…∩T(k−1)gA)c(g):=\mu(A\cap T^{g}A\cap\ldots\cap T^{(k-1)g}A) (as in ), and to be able to show that c(g)c(g) can be approximated by a function of polynomials in gg, in the spirit of the results in . We hope to report on these and other applications in a subsequent paper.

4. The Heisenberg example

To illustrate the above results we now pause to describe the model case of a Heisenberg system. (The discussion in this section is not directly used in the remainder of the paper.) To simplify the discussion we restrict attention to the k=3k=3 case.

5. Overview of the structure of the paper

for some functions Ft1,…,tm(x)F_{t_{1},\ldots,t_{m}}(x). A key technical point is that while the function Ft1,…,tm(x)F_{t_{1},\ldots,t_{m}}(x) is a priori only measurable in xx, it can be made to be measurable in the parameters t1,…,tmt_{1},\ldots,t_{m} also (see Lemma C.4). This will be rather important for us as we will be relying quite heavily on the measurability property For instance, we will need a variant of the classical Steinhaus theorem that asserts that if AA is a measurable subset of a compact abelian group UU with positive measure, then the difference set A−AA-A contains a neighborhood of the origin (cf. Lemma D.1). Curiously, analogous results are exploited in the additive-combinatorial approach to the Gowers inverse problem (see e.g. ), where they go by the name of “Bogolyubov-type lemmas”. in our arguments.

6. Acknowledgements

The authors would like to thank Tim Austin, Ben Green, Bernard Host, Bryna Kra, and Trevor Wooley for many enlightening conversations and suggestions. The first author is supported by NSF grant DMS-0600042. The second author is supported by a grant from the MacArthur Foundation, and by NSF grant DMS-0649473. The first and third authors are supported by BSF grant No. 2006094. The third author is supported by a Landau fellowship of the Taub foundations, and by an Alon fellowship. The authors also thank the anonymous referees for many useful suggestions and corrections.

Abelian cohomology

Throughout the paper we will be relying heavily on the language of abelian cohomology of dynamical systems. We record the key definitions here; further discussion of these concepts can be found in Appendix B.

We caution that abelian extensions of ergodic GG-systems are not necessarily ergodic.

We will need several technical results concerning abelian cohomology groups, which we have collected in Appendix B, and which we will refer to as necessary in the main text of the paper.

Reduction to abelian extensions of order <k+1<k+1

By Theorem A.35, every Abramov system of order <k<k is also a GG-system of order <k<k. This and (A.9), will allow us to immediately derive Theorem 1.20 from the following claim:

The following basic fact was established by Host and Kra:

Using Proposition 3.4 and Fourier analysis, we can reduce matters to studying projections of abelian cocycles to the unit circle. More precisely, in future sections we will show the following result.

It remains to prove Theorem 3.8. This will be the objective of the next few sections.

Functions of type <k<k

We make a further reduction, introducing the useful notion This concept is essentially that of a cocycle of type kk from , but generalized to non-cocycles and to more general group actions. We have replaced “kk” by “<k<k” as such cocycles will have “degree” strictly less than kk in some sense. of a function of type <k<k.

where sgn⁡(w1,…,wk):=w1…wk∈{−1,+1}\operatorname{sgn}(w_{1},\ldots,w_{k}):=w_{1}\ldots w_{k}\in\{-1,+1\}.

We make some easy observations (see Figure 2):

From (i) we see in particular that coboundaries, being of type <0<0, are of type <k<k. The claims in (ii) are then easily verified.

To show (iii), we induct on kk. The claim is easy for k=0,1k=0,1 (using ergodicity), so suppose that k⩾2k\geqslant 2 and the claim has already been shown for k−1k-1.

The relevance of the type <k<k concept to us lies in the important observation that abelian extensions of order <k+1<k+1 arise from functions of type <k<k:

We will prove Theorem 4.5 in later sections. For now, let us observe that we can use Theorem 1.20 to obtain some structural control on systems of order <k<k. We say that a group UU is mm-torsion for some m⩾1m\geqslant 1 if we have um=1u^{m}=1 for all u∈Uu\in U.

Our remaining task is to prove Theorem 4.5.

Reduction to solving a Conze-Lesigne type equation

To prove Theorem 4.5 we will use two lemmas to reduce matters to solving a certain equation of Conze-Lesigne type. The first lemma allows one to descend a type condition on an extension to a type condition on a base, worsening the type if necessary:

For cocycles, a more general (and stronger) statement appears in [18, Corollary 7.8]. However, for technical reasons, it is necessary for us to work with more general functions than just cocycles. (But see Corollary 8.11 below.)

If we expand into a Fourier series K(y,m)=∑χ∈M^aχ(y)χ(m)K({\bf y},{\bf m})=\sum_{\chi\in\hat{M}}a_{\chi}({\bf y})\chi({\bf m}) and compare Fourier coefficients, we conclude

almost everywhere on any ergodic component. In particular, ∣aχ∣|a_{\chi}| is invariant and thus constant a.e. on any ergodic component. We extend χ\chi arbitrarily to a character χ∈U[k+1]^≡U^[k+1]\chi\in\widehat{U^{[k+1]}}\equiv{\hat{U}}^{[k+1]}.

We now claim that (Th)α[k+1]H(T_{h})^{[k+1]}_{\alpha}H is also a (G,A,S1)(G,A,S^{1})-coboundary for every positive side transformation (Th)α[k+1]H(T_{h})^{[k+1]}_{\alpha}H. But a computation shows that

Now we crucially use the fact that d[k]fd^{[k]}f and ρ\rho are cocycles to write this as

The second lemma allows us to reduce the type of a function by differentiation in the vertical direction.

Because of these two lemmas, Theorem 4.5 will follow from

We claim for each 0⩽j⩽m0\leqslant j\leqslant m that

Theorem 4.5 now follows by specializing (5.1) to the case j=0j=0. ∎

Reduction to a finite UU

The purpose of this section is to obtain the following reduction.

In order to prove Theorem 5.4, it suffices to do so in the case when UU is finite.

By Lemma C.4 we can take qt,Ftq_{t},F_{t} to be measurable with respect to tt.

The next step is to linearize qtq_{t} on an open subgroup of UU, by arguing as follows. Let t,u∈Ut,u\in U. Then the cocycle identity Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣tuf=(Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣t(Vuf))Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣uf{\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{tu}f=({\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{t}(V_{u}f)){\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{u}f and (6.1) give

Now let u,v∈U′u,v\in U^{\prime}. From (6.2) we have the second-order cocycle identity ϕuv,wϕu,v=ϕu,vwVuϕv,w\phi_{uv,w}\phi_{u,v}=\phi_{u,vw}V_{u}\phi_{v,w} for all w∈Uw\in U. If ε\varepsilon is small enough, we can find ww so that w,vw,uvw∉Ew,vw,uvw\not\in E. We conclude that

for all u∈Uu\in U, while from (6.2), (6.3) we have Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣Fuv′(VuFv′)Fu′=1{\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}\frac{F^{\prime}_{uv}}{(V_{u}F^{\prime}_{v})F^{\prime}_{u}}=1 and thus (by (6.4))

Now we make a crucial use of the finite characteristic hypothesis. By Lemma 4.7, UU is pnp^{n}-torsion for some n=Ok(1)n=O_{k}(1). By Lemma D.1, we conclude that U′U^{\prime} contains an open subgroup of UU; by reducing U′U^{\prime} if necessary, we may assume that U′U^{\prime} is in fact equal to an open subgroup.

The finite group case

for some 1⩽n1,…,nN⩽Ok(1)1\leqslant n_{1},\ldots,n_{N}\leqslant O_{k}(1) and some finite (but unbounded) NN. When pp is large enough, depending on kk, we can take n1=…=nN=1n_{1}=\ldots=n_{N}=1.

The idea here is to express FjF_{j} as Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣ejF{\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{e_{j}}F times a polynomial error, for some FF independent of jj; we will then “integrate” this to express ff as Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣F{\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}F times a polynomial error, times a function invariant under e1,…,eNe_{1},\ldots,e_{N}; this is basically what we need to establish Theorem 5.4.

We turn to the details. The first task is to measure two potential obstructions to FjF_{j} being expressible as Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣ejF{\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{e_{j}}F (modulo polynomial errors), namely the obstruction coming from the torsion of UU, and the obstruction coming from the multi-dimensionality of UU.

Observe that we have the telescoping identity

Note, conversely, that ∏t=0pnj−1VejtFj\prod_{t=0}^{p^{n_{j}}-1}V_{e_{j}^{t}}F_{j} needed to be polynomial in order to have any chance to express FjF_{j} as Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣ejF{\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{e_{j}}F times a polynomial.

Suppose now that pp is sufficiently large depending on kk, so that nj=1n_{j}=1. By Lemma D.3 we have ∏t=0pnj−1Vejtpj=1\prod_{t=0}^{p^{n_{j}}-1}V_{e_{j}^{t}}p_{j}=1, and hence by (7.2), (7.1) we may now strengthen (7.3) to

In a similar fashion, from the commutation identity Δ ⁣ ⁣ ⁣ ⁣⋅  ⁣eiΔ ⁣ ⁣ ⁣ ⁣⋅  ⁣ejfΔ ⁣ ⁣ ⁣ ⁣⋅  ⁣ejΔ ⁣ ⁣ ⁣ ⁣⋅  ⁣eif=1\frac{{\Delta\!\!\!\!\cdot\ \!}_{e_{i}}{\Delta\!\!\!\!\cdot\ \!}_{e_{j}}f}{{\Delta\!\!\!\!\cdot\ \!}_{e_{j}}{\Delta\!\!\!\!\cdot\ \!}_{e_{i}}f}=1 for any 1⩽i,j⩽N1\leqslant i,j\leqslant N, we see from Lemma B.5(i) and (7.1) that

Again, observe that Δ ⁣ ⁣ ⁣ ⁣⋅  ⁣eiFjΔ ⁣ ⁣ ⁣ ⁣⋅  ⁣ejFi\frac{{\Delta\!\!\!\!\cdot\ \!}_{e_{i}}F_{j}}{{\Delta\!\!\!\!\cdot\ \!}_{e_{j}}F_{i}} had to be polynomial in order to have a chance to express FiF_{i}, FjF_{j} as Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣eiF{\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{e_{i}}F, Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣ejF{\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{e_{j}}F modulo polynomials.

Now we compute a derivative of FF. We clearly have

for any 1⩽j⩽N1\leqslant j\leqslant N. On the other hand, we have the telescoping identity

When p=Ok,m(1)p=O_{k,m}(1), this claim follows from Corollary D.6, so suppose now that pp is sufficiently large depending on k,mk,m. Then ωij(y,[t1,…,ti−1,ti′,0,…,0])\omega_{ij}(y,[t_{1},\ldots,t_{i-1},t^{\prime}_{i},0,\ldots,0]) is a phase polynomial of degree Ok,m(1)O_{k,m}(1) in ti′t^{\prime}_{i} that takes values in CpC_{p}. By Taylor expansion we may thus write

Inserting the above claim into (7.8) we conclude that

Thus if we set f′:=f/Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣Ff^{\prime}:=f/{\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}F, then ff is cohomologous to f′f^{\prime} and

Now we need to work on the Fj′F^{\prime}_{j} term. From the telescoping identity ∏0⩽tj<pnjVejtjΔ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣ejf′=1\prod_{0\leqslant t_{j}<p^{n_{j}}}V_{e_{j}^{t_{j}}}{\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{e_{j}}f^{\prime}=1 and (7.9), we have

pushing forward by π\pi and then using (3.1), we conclude

In the case that pp is sufficiently large depending on k,mk,m, we see from Lemma D.3 that we have the improvement π∗Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣(Fj′)pnj=1\pi^{*}{\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}(F^{\prime}_{j})^{p^{n_{j}}}=1, and so (Fj′)pnj(F^{\prime}_{j})^{p^{n_{j}}} is constant in this case.

Inserting this claim back into (7.9) we obtain

From Proposition 7.1 and Proposition 6.1, we obtain Theorem 5.4, and thus Theorem 3.3.

The high characteristic case

We now develop high characteristic analogues of the above theory, establishing the sharp Theorem 1.19 instead of Theorem 1.20 in this setting. The arguments here will be similar to those used to prove Theorem 1.20; the main new difficulty is to be careful to not lose anything in the degree of various functions beyond what is absolutely necessary.

Just as Theorem 1.20 follows from Theorem 4.5, Theorem 1.19 will follow from

It is clear that Theorem 8.1 then follows from the j=kj=k case of

2. Vertical differentiation

Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣tf{\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{t}f is a line cocycle, is of type <k−j+1<k-j+1, and is a quasi-cocycle of order <k−j<k-j.

3. Reduction to the finite UU case

We now argue (as in Proposition 6.1) that in order to conclude the proof of Theorem 8.6, it suffices to do so in the case when the vertical structure group Uj−1U_{j-1} is finite.

We now invoke the following variant of Lemma B.6 which is more efficient with the degree.

In the converse direction, one can show (by using the properties of the nilpotent group G[k]{\mathcal{G}}^{[k]} studied in ) that if QQ has degree <l+j<l+j, then Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣tQ{\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{t}Q has degree <l<l. This may explain the terminology “exact”.

The remaining tasks (in the low order case j⩽kj\leqslant k) are to verify Theorem 8.6 in the case of finite Uj−1U_{j-1}, and to also verify Proposition 8.9 and Proposition 8.11.

4. The finite group case

We now establish Theorem 8.6 in the case when Uj−1U_{j-1} is finite. This is the analogue of Proposition 7.1, but our arguments here are somewhat simpler thanks to the high characteristic (which allows us to use the full power of Lemma D.3).

By repeated application of Lemma 8.8, we know that qeq_{e} has degree <p<p with respect to differentiation Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣e{\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{e} in the ee direction. By Lemma D.3, we conclude that ∏i=0p−1qe(g,Veix)=1\prod_{i=0}^{p-1}q_{e}(g,V_{e}^{i}x)=1, and thus the qesq_{e^{s}} form a cocycle in the ese^{s} variable, in the sense that

5. The high order case

As in previous arguments, we first reduce to the case when Uj−1U_{j-1} is finite, and then establish the finite case.

Our remaining tasks are to prove Proposition 8.9, Proposition 8.11, and Lemma 8.13.

6. Exact integration

In this subsection we establish Proposition 8.9. We begin with the analogue of Lemma B.5.

We prove (i) and (ii) simultaneously by induction on ll. If l=1l=1, then ptp_{t} is a constant (by ergodicity), and so the map t↦ptt\mapsto p_{t} is a homomorphism, and the claim (i) then easily follows. Also in this case (ii) is clearly equivalent to (i).

Now suppose inductively that l⩾2l\geqslant 2, and that the claim has already been proven for l−1l-1. We now induct on mm. When m=1m=1 then rr is constant, and (i) is clear.

and thus (since the action of UU commutes with that of GG)

7. Exact descent

We now prove Proposition 8.11 and Lemma 8.13. As noted already in Remark 5.2, an exact descent result for cocycles already appears in [18, Corollary 7.8]; Proposition 8.11 can be viewed as an extension of that result to quasi-cocycles.

Our main tool for both of these tasks is the following equivalent characterization of the finite type condition.

where each fwf_{\bf w} is either equal to ff or its complex conjugate.

On the other hand, ff has magnitude 11. We conclude (for ε\varepsilon small enough) that

and thus by the pigeonhole principle that

for all g∈Gg\in G. Averaging this over a Følner set Φn\Phi_{n}, we conclude in particular that

for all nn. From the monotone convergence theorem we conclude that

We can now prove Proposition 8.11 and Lemma 8.13.

The proof of Theorem 1.19 is now complete.

Appendix A General theory of Gowers-Host-Kra seminorms

More generally, for any 0⩽l⩽k0\leqslant l\leqslant k, define an ll-dimensional face or ll-face to be any set formed by intersecting k−lk-l distinct non-parallel sides. Thus 2k{\bf 2}^{k} has one kk-face, 2k2k faces of dimension k−1k-1 (i.e. the sides βj±\beta_{j}^{\pm}), k2k−lk2^{k-l} faces of dimension (k−l)(k-l), and so forth down to 2k2^{k} faces of dimension 00 (which are the vertices of the discrete cube).

Let α\alpha be an ll-face. Enumerating the elements of α\alpha in lexicographic order gives a natural bijection ∂(α):α→2l\partial(\alpha):\alpha\to{\bf 2}^{l}, which we call the coordinate map of α\alpha, which maps the faces of 2k{\bf 2}^{k}, which are subsets of α\alpha, to the faces of 2l{\bf 2}^{l}.

Let SS be an arbitrary set. We write S[k]:=S2kS^{[k]}:=S^{{\bf 2}^{k}} for the set of functions s:2k→S{\bf s}:{\bf 2}^{k}\to S. For each 0⩽l⩽k0\leqslant l\leqslant k and each ll-face α\alpha, we let ∂(α)∗:S[k]→S[l]\partial(\alpha)_{*}:S^{[k]}\to S^{[l]} be the pushforward map given by the formula ∂(α)∗(s)(w):=s(∂(α)−1(w))\partial(\alpha)_{*}({\bf s})({\bf w}):={\bf s}(\partial(\alpha)^{-1}({\bf w})) for all w∈2l{\bf w}\in{\bf 2}^{l}. For any 1⩽j⩽k1\leqslant j\leqslant k and sign ±\pm, we abbreviate the cubic boundary map ∂(βj±)∗:S[k]→S[k−1]\partial(\beta_{j}^{\pm})_{*}:S^{[k]}\to S^{[k-1]} as ∂j±=∂j,k±\partial_{j}^{\pm}=\partial_{j,k}^{\pm}.

For future reference we make the trivial observation that the map s↦(∂k+1−s,∂k+1+s){\bf s}\mapsto(\partial_{k+1}^{-}{\bf s},\partial_{k+1}^{+}{\bf s}) is a bijection between S[k+1]S^{[k+1]} and S[k]×S[k]S^{[k]}\times S^{[k]}.

Let GG be a (possibly non-abelian) group with identity id⁡G\operatorname{id}_{G}, and let α\alpha be a face of 2k{\bf 2}^{k}. For every g∈Gg\in G, we let gα[k]∈G[k]g^{[k]}_{\alpha}\in G^{[k]} denote the group element whose components (gα[k])w(g^{[k]}_{\alpha})_{\bf w} for w∈2k{\bf w}\in{\bf 2}^{k} are defined to equal gg when w∈α{\bf w}\in\alpha, and to equal id⁡G\operatorname{id}_{G} otherwise. The map g↦gα[k]g\mapsto g^{[k]}_{\alpha} is a bijection from GG to the face group Gα[k]:={gα[k]:g∈G}⩽G[k]G^{[k]}_{\alpha}:=\{g^{[k]}_{\alpha}:g\in G\}\leqslant G^{[k]}. When α\alpha is a side (resp. a positive side) we refer to Gα[k]G^{[k]}_{\alpha} as a side group (resp. a positive side group); when α=2k\alpha={\bf 2}^{k} is the entire cube we refer to G2k[k]G^{[k]}_{{\bf 2}^{k}} as the diagonal group and denote it as diag⁡(G[k])\operatorname{diag}(G^{[k]}), and abbreviate g2k[k]g^{[k]}_{{\bf 2}^{k}} as g[k]g^{[k]}. We also write ∂[k]G\partial^{[k]}G (resp. ∂+[k]\partial^{[k]}_{+}) for the subgroup of G[k]G^{[k]} generated by all the side groups (resp. all the positive side groups).

For k=2k=2, the group ∂G\partial^{}G is generated by

while the group ∂+G\partial_{+}^{}G is generated by ⋃g∈G{(id⁡G,id⁡G,g,g),(id⁡G,g,id⁡G,g)}g∈G.\bigcup_{g\in G}\{(\operatorname{id}_{G},\operatorname{id}_{G},g,g),(\operatorname{id}_{G},g,\operatorname{id}_{G},g)\}_{g\in G}.

For future reference we observe that the side group ∂[k]G\partial^{[k]}G is the group generated by the positive side group ∂+[k]G\partial^{[k]}_{+}G and the diagonal group diag⁡(G[k])\operatorname{diag}(G^{[k]}).

Let GG be a group acting on a space XX by transformations Tg:X→XT_{g}:X\to X for g∈Gg\in G. Then G[k]G^{[k]} acts on X[k]X^{[k]} in the obvious manner, with the action Tg[k]T^{[k]}_{\bf g} of a group element g=(gw)w∈2k∈G[k]{\bf g}=(g_{\bf w})_{{\bf w}\in{\bf 2}^{k}}\in G^{[k]} mapping each point (xw)w∈2k∈X[k](x_{\bf w})_{{\bf w}\in{\bf 2}^{k}}\in X^{[k]} to (Tgw(xw))w∈2k(T_{g_{\bf w}}(x_{\bf w}))_{{\bf w}\in{\bf 2}^{k}}. If α\alpha is a face, we abbreviate the face transformation Tgα[k][k]T^{[k]}_{g^{[k]}_{\alpha}} as (Tg)α[k](T_{g})^{[k]}_{\alpha}, thus ((Tg)α[k])g∈G((T_{g})^{[k]}_{\alpha})_{g\in G} is an action of GG on X[k]X^{[k]}. If α\alpha is a side (resp. a positive side), we refer to (Tg)α[k](T_{g})^{[k]}_{\alpha} as a side transformation (resp. positive side transformation), and if α=2k\alpha={\bf 2}^{k} is the entire cube, we refer to (Tg)2k[k](T_{g})^{[k]}_{{\bf 2}^{k}} as a diagonal transformation and abbreviate it further as (Tg)[k](T_{g})^{[k]}.

For future reference we observe that all positive side maps preserve the (−1,…,−1)(-1,\ldots,-1) coordinate of X[k]X^{[k]}.

We recall the notion of a relative product:

Now we can introduce the cubic measure spaces of Host and Kra.

is the ergodic decomposition of μ[k]\mu^{[k]} with respect to the diagonal action (Tg[k])g∈G(T_{g}^{[k]})_{g\in G} of GG.

The cubic measures have a useful symmetry property:

The cubic measures also behave well with respect to passage to subcubes.

and Ik\mathcal{I}_{k} is the σ\sigma-algebra consisting of subsets of G[k]G^{[k]} which are invariant under the diagonal translations (xw)w∈2k↦(xw+h)w∈2k(x_{\bf w})_{{\bf w}\in{\bf 2}^{k}}\mapsto(x_{\bf w}+h)_{{\bf w}\in{\bf 2}^{k}} for h∈Gh\in G. Thus the probability space (X[k],B[k],μ[k])(X^{[k]},\mathcal{B}^{[k]},\mu^{[k]}) is measure isomorphic to the space Gk+1={(x,h1,…,hk):x,h1,…,hk∈G}G^{k+1}=\{(x,h_{1},\ldots,h_{k}):x,h_{1},\ldots,h_{k}\in G\} with the discrete σ\sigma-algebra and normalized counting measure, whilst (X[k],Ik,μ[k])(X^{[k]},\mathcal{I}_{k},\mu^{[k]}) is measure isomorphic to the space Gk={(h1,…,hk):h1,…,hk∈G}G^{k}=\{(h_{1},\ldots,h_{k}):h_{1},\ldots,h_{k}\in G\} with the discrete σ\sigma-algebra and normalized counting measure.

A.2. Existence of the seminorms

The objective of this section is to establish that the Gowers-Host-Kra seminorms from Definition 1.3 are in fact well-defined, and to relate them to the cubic measures just constructed.

We prove this by induction on kk. For k=1k=1 this follows from the mean ergodic theorem. Assume the induction hypothesis holds for k−1k-1. Then

Since μ[k−1]\mu^{[k-1]} is invariant with respect to the action of (Tg[k])g∈G(T_{g}^{[k]})_{g\in G}, by the ergodic theorem the above averages converge to

We record some basic properties of the Gowers-Host-Kra seminorms:

[18, Lemma 3.9] (See also [11, Lemmas 3.8, 3.9])

The measure PkP_{k} is ergodic with respect to the action of ∂+[k]G\partial^{[k]}_{+}G.

For any measure-preserving transformation u:X→Xu:X\to X that commutes with the GG-action, and any side α\alpha, the side transformation uα[k]u^{[k]}_{\alpha} preserves μ[k]\mu^{[k]}.

For (i), see [18, Corollary 3.5]; for (ii), see [18, Corollary 3.6]. For (iii) and (iv), see [18, Lemma 5.5]. ∎

A.3. Dual functions

The limits above exist, as in Lemma A.18, by repeated applications of the ergodic theorem. Indeed, we easily verify that

where −1:=(−1,…,−1)-{\bf 1}:=(-1,\ldots,-1) and the factor map is given by (xw)w∈2k↦x−1(x_{\bf w})_{{\bf w}\in{\bf 2}^{k}}\mapsto x_{-{\bf 1}} (i.e. the pushforward map ∂({−1})∗\partial(\{-{\bf 1}\})_{*}). As a consequence we have

and similarly (by repeated applications of the Cauchy-Schwarz inequality) that

The limit above is a repeated limit, but a posteriori, using Theorem 1.20, one can show that the double (simultaneous) limit exists as well, and both limits coincide, by modifying the proof of [18, Theorem 1.2], and similarly for higher values of kk. We omit the details.

where C:z↦z‾{\mathcal{C}}:z\mapsto\overline{z} is the complex conjugation operator. Dual functions in this setting play an important role in the finitary theory of arithmetic progressions and similar patterns; see .

Note that Lemma A.32 immediately implies Proposition 1.10 in the introduction. From this lemma and (A.5) we also have

for 0<j⩽k0<j\leqslant k (cf. [18, Corollary 4.4]).

From Lemma A.32 and Lemma A.22 one can show that universal characteristic factors are functorial:

Appendix B Abelian cohomology

The reader may wish to review the definitiosn in Definition 2.1 before proceeding with the rest of this section.

We begin with the following trivial but useful lemma:

Next, we recall that cohomology is trivial for free actions:

If GG acts freely on XX, then so does any compact abelian subgroup of GG.

There is an analogue of Lemma B.4 in the polynomial category. To state it, we first need a useful algebraic lemma.

We prove (i) by induction on dd. Indeed, the claim is trivial for d=1d=1 by ergodicity, and for d>1d>1 we have by induction that Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣gΔ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣up=Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣uΔ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣gp{\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{g}{\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{u}p={\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{u}{\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{g}p is a phase polynomial of degree <d−2<d-2 for all g∈Gg\in G, and thus by (3.1) Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣up{\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{u}p is a phase polynomial of degree <d−1<d-1 as claimed.

Finally, we prove claim (iv). The case d′=0d^{\prime}=0 follows from the previous claim, so suppose that d′>0d^{\prime}>0 and the claim has already been proven for the smaller values of d′d^{\prime}. For g∈Gg\in G, we take a derivative Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣gP{\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{g}P. One obtains essentially the same terms that appeared in the previous claim, plus (thanks to the cocycle equation (B.1)) some additional terms involving qΔ ⁣ ⁣ ⁣ ⁣⋅  ⁣gr(y,u)q_{{\Delta\!\!\!\!\cdot\ \!}_{g}r(y,u)}. But such terms can be dealt with by the induction hypothesis. ∎

thanks to (B.1). Thus we have qu=Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣uQq_{u}={\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{u}Q for all u∈Uu\in U.

We will also need another result in a similar spirit.

We will also take advantage of a useful splitting lemma.

Now we see how cohomology on an abelian extension relates to cohomology on the base space.

for all g∈Gg\in G and almost every x∈Xx\in X, u∈Uu\in U. We rearrange this as

We perform a Fourier expansion in UU, obtaining

This completes the proof in the case when VV is trivial.

Appendix C A measurable selection lemma

By dividing ϕ,ψ\phi,\psi by ψ\psi we may assume ψ=1\psi=1.

The claim is vacuous when k=1k=1. When k=2k=2 we argue as follows. For any h∈Gh\in G we have

If ϕ\phi is a phase polynomial of degree <2<2, then Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣hϕ{\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{h}\phi is constant; if ϕ\phi is non-constant, then (by ergodicity) Δ ⁣ ⁣ ⁣ ⁣ ⁣ \textbullet  ⁣hϕ{\Delta\!\!\!\!\!\hbox{\raisebox{0.86108pt}{\tiny\ \textbullet}}\ \!}_{h}\phi is not identically 11 for at least one hh. Thus ∫Xϕ dμX=0\int_{X}\phi\ d\mu_{X}=0 and so ∥ϕ−1∥L2(X)=2\|\phi-1\|_{L^{2}(X)}=\sqrt{2}, and the claim follows.

We are now ready to establish the measurable selection lemma.

One can also establish this result using a general measure selection result of Dixmier (see e.g. [2, Theorem 1.2.4]) together with Lusin’s theorem and Corollary C.3; we omit the details. One can also appeal to the descriptive set theory of Polish groups, see e.g. [18, Appendix A].

Appendix D Finite characteristic algebra

Recall that a group UU is mm-torsion if we have um=1u^{m}=1 for all u∈Uu\in U.

Let UU be a compact abelian mm-torsion group for some m⩾1m\geqslant 1. Let VV be an open neighborhood of the identity in UU. Then VV contains an open subgroup WW of UU.

We will use a Fourier-analytic method. As VV is an open neighborhood of the origin, one can find another open neighborhood V′V^{\prime} of the origin such that V′−V′⊂VV^{\prime}-V^{\prime}\subset V.

Let μ\mu be the Haar measure on UU, then μ(V′)>0\mu(V^{\prime})>0. Let ε>0\varepsilon>0 be a small number (depending on μ(V′)\mu(V^{\prime})) to be chosen later. By Fourier analysis, we can approximate the indicator function 1V′1_{V^{\prime}} to within ε\varepsilon in L2(U)L^{2}(U)-norm by some linear combination FF of finitely many characters χ1,…,χn∈U^\chi_{1},\ldots,\chi_{n}\in\hat{U}, where nn is finite but potentially unbounded. Since UU is mm-torsion, each character χj\chi_{j} takes on at most mm values, with each level set of χj\chi_{j} being a coset of an open subgroup of UU. If we let WW be the intersection of the kernels of all the χj\chi_{j}, then WW is also an open subgroup of UU, and FF is constant on every coset of WW. Since FF approximates 1V′1_{V^{\prime}} to within ε\varepsilon, we conclude (if ε\varepsilon is sufficiently small depending on μ(V′)\mu(V^{\prime})) that there exists a coset of WW on which V′V^{\prime} has density greater than 1/21/2. But then this forces W⊂V′−V′W\subset V^{\prime}-V^{\prime} and hence W⊂UW\subset U, as desired. ∎

Let UU be a compact abelian mm-torsion group for some m⩾1m\geqslant 1. Let WW be an open subgroup of UU. Then there exists a splitting U=W′×YU=W^{\prime}\times Y, where W′W^{\prime} is an open subgroup of WW, and YY is a finite abelian mm-torsion group.

It is known (see e.g. [23, Chapter 5, Theorem 18]) that a compact abelian mm-torsion group UU is topologically isomorphic to the direct product of cyclic mm-torsion groups In particular, the bounded torsion allows us to avoid having to deal with procyclic groups which are not direct products of cyclic groups.. Thus WW must contain a cylinder neighbourhood W′W^{\prime} of the origin, i.e. a cofinite sub-product of these cyclic groups. Since one clearly has the desired splitting U=W′×YU=W^{\prime}\times Y, the claim follows. ∎

D.2. Polynomials are discretely valued

Let g∈Gg\in G. Since Tgpf=fT_{g}^{p}f=f and Tg=1+ΔgT_{g}=1+\Delta_{g}, we conclude using the binomial formula that ∑i=0p(pi)Δgif=f.\sum_{i=0}^{p}\binom{p}{i}\Delta_{g}^{i}f=f. Since ff has degree <p<p, Δgpf=0\Delta_{g}^{p}f=0. We conclude that

Inverting the expression in brackets using Neumann series (and using the fact that Δgp−1\Delta_{g}^{p-1} annihilates Δgpf\Delta_{g}pf) we conclude that Δgpf=0\Delta_{g}pf=0 for any gg, thus by ergodicity pfpf is constant as claimed.

To prove (ii), we first observe that it suffices to prove the claim for kk of the form k=pm+1k=pm+1 for integer mm. But the claim is trivial for m=1m=1, and from (i), we see that the claim for mm implies the claim for m+1m+1, and so (ii) follows by induction.

D.3. Roots of phase polynomials

Recall that a map P:X→HP:X\to H into an additive group HH is a polynomial of degree <d<d if we have Δg1…ΔgdP=0\Delta_{g_{1}}\ldots\Delta_{g_{d}}P=0 for all g1,…,gd∈Gg_{1},\ldots,g_{d}\in G.

We may assume inductively that the claim is already proven for smaller values of dd; for the same value of dd and smaller values of ll; or the same value of dd and ll and smaller values of jj. We abbreviate Ol,d,p,j(1)O_{l,d,p,j}(1) as O(1)O(1).

From primary school arithmetic we know that we have a formula of the form

for some “carry bit” functions cj:{0,1,…,p−1}2j→{0,1}c_{j}:\{0,1,\ldots,p-1\}^{2j}\to\{0,1\}. Applying this with PP and ΔgP\Delta_{g}P for some group element gg we conclude

Finally, by the induction hypothesis on ll, we know that

Putting all this together we see that Δg(bj(P))\Delta_{g}(b_{j}(P)) is a polynomial of degree O(1)O(1) for all gg, and hence bj(P)b_{j}(P) is a polynomial of degree O(1)O(1), thus closing the induction. ∎

We isolate one special case of Corollary D.6:

By rotating ϕ\phi by a constant and using Lemma D.3, we may assume that ϕ\phi takes values in CpmC_{p^{m}} for some m=Od,p(1)m=O_{d,p}(1). If nn is not divisible by pp, then nn is invertible in CpmC_{p^{m}} and the claim is immediate, so it suffices to check the case when nn is a power of pp. But then the claim follows immediately from Corollary D.6. ∎

Another interesting consequence of Corollary D.6 (or Proposition D.5) is that phase polynomials can always be expressed in terms of CpC_{p}-valued polynomials of higher degree.

Appendix E Connection with cubic complexes

In this appendix we point out some connections between the notions of polynomiality and type in this paper with the theory of cubic complexes as used in topology, as set out in , in analogy with the more well-known simplicial complexes used in that field. This material is not used elsewhere in this paper.

References