The rate of convergence in the method of alternating projections
Catalin Badea, Sophie Grivaux, Vladimir Muller
Introduction
Throughout the paper is a complex Hilbert space. For a closed linear subspace of we denote by its orthogonal complement in , and by the orthogonal projection of onto . In this paper denotes a fixed positive integer greater or equal than .
It was proved by J. von Neumann [27, p. 475] that for two closed subspaces and of , with intersection , the following convergence result holds:
Using the notation , von Neumann’s result says that the iterates of are strongly convergent to . The method of constructing the iterates of by alternately projecting onto one subspace and then the other is called the method of alternating projections. This algorithm, and its variations, occur in several fields, pure or applied. We refer to [10, Chapter 9] as a source for more information.
A generalization of von Neumann’s result to closed subspaces with intersection was proved by Halperin : for each we have
The algorithm provided by Halperin’s result will be called in this paper the method of cyclic alternating projections.
A Banach space extension of Halperin’s result was proved by Bruck and Reich : if is a uniformly convex Banach space and , , are norm one projections in , then the iterates of are strongly convergent. The strong limit is a projection of norm one onto the intersection of the ranges of . The same result holds if is uniformly smooth and each projection is of norm one. It also holds if is a reflexive (complex) Banach space and each projection is hermitian (that is, with real numerical range). We refer to and the references therein for other Banach space results of this type.
An interesting extension of the method of cyclic alternating projections is the method of random alternating projections. Let , , be orthogonal projections in , , and let be a sequence from (random samples). The method of random alternating projections asks about the convergence of the sequence given by , . It is an open problem to know whether is always convergent in the topology of . The convergence of in the weak topology has been proved by Amemiya and Ando . If each between and occurs infinitely many times in the sequence of random samples, then the weak limit of is . We refer to for results related to this problem.
B. The rate of convergence
It is important for applications to know how fast the algorithm given by the method of alternating projections, or its variations, converge. For a quite complete description of the rate of convergence is known, it terms of the notion of angle of subspaces.
(Friedrichs angle) Let and be two closed subspaces of the Hilbert space with intersection . The Friedrichs angle between the subspaces and is defined to be the angle in whose cosine is given by
where is the unit ball of . The minimal angle (or Dixmier angle) between the subspaces and is defined to be the angle in whose cosine is given by
We note that , and that if . We also have . We refer to the survey paper for more information about different notions of angle between subspaces of infinite dimensional Hilbert spaces and their properties, and to [28, Lecture VIII] for different occurences of the Friedrichs angle in functional-theoretical problems.
It was proved by Aronszajn (upper bound) and by Kayalar and Weinert (equality) that
This formula shows that the sequence of iterates of converges uniformly to if and only if , i.e., if the Friedrichs angle between and is positive. When this happens, the iterates of converge “quickly” (i.e. at the rate of a geometrical progression) to , in the following sense:
(quick uniform convergence) there exist and such that
It is also known that if and only if is closed, if and only if is closed, if and only if is closed.
When is not closed, we have strong, but not uniform convergence. It was recently proved by Bauschke, Deutsch and Hundal (see for the history of this result) that, given any sequence of reals decreasing to zero, there exists a point in the space with the property that the convergence in the method of alternating projections (von Neumann’s theorem) is at least as slow as this sequence of reals. Thus the iterates of the product of two orthogonal projections converge quickly, or arbitrarily slowly. We call this alternative the (QUC)/(ASC) dichotomy : one has quick uniform convergence or arbitrarily slow convergence. We shall consider several meanings of (ASC) in this paper.
The results concerning the rate of convergence in Halperin’s theorem for are not as complete as the results described above for . We refer to , [10, Chapter 9] and their references for several results concerning the rate of convergence in the method of cyclic alternating projections. For instance, [13, Example 3.7] shows that for the error bound for the method of cyclic alternating projections is not a function of the various Friedrichs angles between pairs of subspaces.
C. What this paper is about
The main goal of the present paper is to discuss the rate of convergence in Halperin’s theorem and to generalize some of the previous known results () to the case of several subspaces (). We show by operator-theoretical methods that the (QUC)/(ASC) dichotomy always holds as soon as the iterates of are strongly convergent. Several interpretations of (ASC) are proposed, and general dichotomy theorems are obtained in the Hilbert or Banach space situation, depending on several spectral properties imposed upon the operator . This implies at once the dichotomy (QUC)/(ASC) in all above-mentioned generalizations of the method of alternating projections. We also give a generalization of the Friedrichs angle to several subspaces, , and prove that condition (QUC) holds in Halperin’s theorem if and only if . Estimates for the error are given in this case and several statements equivalent to the condition are obtained. Some of them are expressed in terms of random products of projections. More specific descriptions of these results, and information about how the paper is organized, are given below.
D. Conditions for arbitrarily slow convergence
Several dichotomy theorems of the type quick uniform convergence versus arbitrarily slow convergence are proved in this paper. The quick uniform condition is the condition (QUC) presented above. We shall consider in Section 2 the following conditions for (ASC):
(arbitrarily slow convergence, variant 1) for every and every sequence of positive numbers such that , there exists a vector such that and for all .
(arbitrarily slow convergence, variant 2) for every sequence of positive numbers such that , there exists a dense subset of points such that for all but a finite number of ’s.
(arbitrarily slow convergence, variant 3) for every sequence of positive numbers such that , there exist two vectors and (the dual of ) such that for all . Furthermore, if there is a Banach space such that is a (isometrical) subspace of , then the vector can be chosen in ;
(arbitrarily slow convergence, Hilbertian version) for every and every sequence of positive numbers such that , there exists a vector such that and for all . Here is supposed to be a complex Hilbert space.
The dichotomy results of Section 2 are based upon general results about the existence of large (weak) orbits of operators (see ).
Let us recall here the main result of concerning large weak orbits in the Banach space setting:
Let be a Banach space which does not contain , and a bounded operator on such that belongs to the spectrum of and tends to zero as tends to infinity for every . Then for any sequence such that tends to zero as tends to infinity, there exists a vector and a functional such that for every .
E. A generalization of the Friedrichs angle
In order to quantify the rate of convergence in the method of alternating projections, an extension of the cosine of Friedrichs angle to several subspaces will be given in Section 3. It is a parameter which lies between and , defined as follows:
In Section we characterize in several ways when the dichotomy (QUC)/(ASC) arises. The characterizations are in terms of geometric properties of , of spectral properties of , or of random products . We give an estimate for the geometric convergence of to zero when .
General dichotomy theorems and applications
Let be a Banach space and let be such that the sequence of iterates is strongly convergent to . Then the following dichotomy holds : either (QUC), or (ASC1). The quick uniform convergence (condition (QUC)) holds if and only if
In these statements, the condition (ASC1) can be replaced by (ASC2).
Suppose that the sequence of iterates is strongly convergent to . Then is mean ergodic, i.e., the Cesàro means are strongly convergent. Therefore ([21, page 73]) the space can be decomposed as the direct sum of the kernel of and the closure of the range of the same operator, . Moreover, is the projection onto along . Notice also that acts on the space as the identity. With respect to the decomposition we can write
Case (1). We have . Notice that we have
Since , there exist and such that
This estimate and (2.2) gives the quick uniform convergence condition (QUC).
Case (2). We have . Recall that as , for each . The conditions (ASC1) and (ASC2) follow now from [26, Thm 14, p. 333].
The following is a different argument for the last part of the proof, without the use of Fredholm theory. As and is closed, the operator is lower bounded, and thus is not in the approximate point spectrum of . As every point in the boundary of the spectrum is in the approximate point spectrum, we obtain the desired contradiction. We refer the reader to as a basic reference for the spectral theory of linear operators we are using in the present paper.
In all these statements, the quick uniform convergence condition (QUC) holds if and only if
B. Applications to the method of alternating projections
We introduce first some notation, and recall for the convenience of the reader some Banach space terminology. Let . Let be a Banach space and let be fixed projections () acting on . We denote by the convex multiplicative semigroup generated by . Recall that this is the convex hull of the set of all products with factors from , and that the convex hull of every multiplicative semigroup of operators is a semigroup.
The space is said to be uniformly convex if for every there exists such that for any two vectors, and , with and , implies . An (equivalent) definition of a uniformly smooth Banach space is the following: is uniformly smooth if its dual, , is uniformly convex. We refer to for more information.
We call a norm one projection (non-zero orthoprojection) if and . A self-adjoint projection in a Hilbert space is called, as usual, an orthogonal projection. Recall that an operator on a Banach space is called hermitian if its numerical range is real. This is equivalent to ask that for every real . Hermitian operators on Hilbert spaces coincide with the self-adjoint ones; see for instance and the references therein.
Let . Let be a complex Banach space, and let be projections on . Let be an operator in . If one of the following conditions below holds true, then the sequence of iterates of converges strongly and every dichotomy (QUC)/(ASC1), (QUC)/(ASC2), (QUC)/(ASC3) and (QUC)/(ASCH) (if is a Hilbert space) applies:
the space is uniformly convex and each , , is a norm one projection;
the space is uniformly smooth, and each , , is a norm one projection;
the space is reflexive and for each there exists with such that . In particular, this holds if each is hermitian, .
A generalization of Friedrichs angle for NN subspaces
As mentioned in the introduction, the rate of convergence in the method of alternating projections for two closed subspaces and is controlled by the Friedrichs angle . We introduce and study in this section a generalization of Friedrichs angle for subspaces.
In order to introduce our generalization of the cosine of the Friedrichs angle to several closed subspaces, we start by giving an equivalent definition of the Friedrichs angle .
(a) Let and be two closed subspaces of . Then
(b) Let and be two closed subspaces in . Then
We give the proof only for the first equality of the second part. Denote by the first supremum from the statement of part (b). For every admissible pair with we have
As is arbitrary, we obtain . ∎
Let . Let be closed subspaces of with intersection . The Dixmier number associated to is defined as
The Friedrichs number associated to is defined as
B. Other parameters and properties of the Friedrichs number.
We found convenient to introduce the following parameters, called the (reduced or not) configuration constants, although they can be expressed in terms of the Dixmier and Friedrichs numbers (see Proposition 3.6, (f)).
Let . Let be closed subspaces of with intersection . The number
is called the non-reduced configuration constant of . The number
is called the configuration constant of .
The configuration constant is related to the maximal possible norms of Gramian matrices. Recall that the Gramian matrix of an -tuple of vectors is the matrix
Let . Let be closed subspaces of with intersection . Then
The conclusion follows by taking the supremum and noting that the Gramian matrix is a Hermitian matrix. ∎
Consider the product Hilbert space which is the Hilbertian direct sum of copies of , with scalar product
We denote by the Cartesian product , and by the diagonal subset . Recall that .
The projections onto , and are given by
The formulae for and were proved in . For the third one, we note first that
The infimum is realized when the gradient is zero, , that is when . ∎
Let . Let be closed subspaces of with intersection . Then
if , while if and only if the subspaces are pairwise orthogonal;
, and thus if ;
and ;
and ;
and ;
and similar statements hold for .
We start by giving the proof of part (e). We have
The proof of the equality is similar.
We prove now that . Indeed, we have
The proof of the equality for is similar.
The upper bound in (d) follows from the Cauchy-Schwarz inequality:
For the lower bound , notice that we have, for ,
The inequalities for follow from
Now (c) is a consequence of (f) and (d), while (b) and (a) are easy to prove. For the first equality in (b) notice that . ∎
Let . Let be closed subspaces of with intersection . Then
Using Lemma 3.5, can be written as
Let be the matrix having all entries equal to . One way to compute the norm of is to note that, like every circulant matrix, is unitarily equivalent to a diagonal matrix. Indeed, denote by the unitary matrix representing the discrete Fourier transform , where is a primitive th root of unity. Then
Since , this can be written as
The following definition is related to the minimum gap between two subspaces (see [18, p. 219 and Lemma 4.4]). See also the regularity (or boundedly linearly regularity) condition from , and the references therein.
Let . Let be closed subspaces of with intersection . The number
is called the inclination of .
Let . Let be closed subspaces of with intersection . Then
Therefore . We also have and . Thus
As this inequality is true for every we get
Set and . Then and
Let . Then , and we have
Characterising (ASC) for products of projections
When is the product of orthogonal projections, we know from Theorem 2.4 that the dichotomy (QUC)/(ASC) holds, and that we have quick uniform convergence if and only if the range of is closed. The following qualitative result gives a characterization of the (ASC) condition in terms of several parameters associated to , or spectral properties of , or random products. We denote by the essential norm and by the essential spectrum.
Let . Let be closed subspaces of with intersection . Denote the orthogonal projection onto , , and by the orthogonal projection onto . Let . The following assertions are equivalent:
for every and every sequence of indices such that , is not closed;
one of the conditions (ASC1), (ASC2), (ASC3), (ASCH) holds for ;
(ASCH) for random products: for every , every sequence of positive reals with , and every sequence of indices in , there exists with such that
for every , every closed subspace of finite codimension (in ), there exists such that and ;
for every and every we have ;
for every and every sequence of indices , , with we have ;
for every and every we have ;
for every , every we have ;
for every , every closed subspace of finite codimension (in ), there exists such that ;
for every , every closed subspace of finite codimension (in ), there exists such that for every , every ;
the sum of and is not closed in (and equivalent statements for , );
is not closed in .
The conditions and , most of them of spectral nature, are conditions about , while the corresponding conditions denoted with primes are analog conditions about random products . The conditions and are about the geometry of subspaces .
Notice that we have the dichotomy (QUC)/(ASC) in all possible senses, and that (QUC) holds if and only if . A quantitative estimate reflecting the geometric convergence of to zero, in terms of the Friedrichs number, will be given after the proof of the theorem.
”“ The equivalence of (1) and (2) follows from Theorem 2.4.
”“ The equivalence of (1) and (5) follows from the proof of Theorem 2.1 (see also Remark 2.2). Notice that, with respect to the decomposition , we have , where .
”“ We prove this implication in a quantitative form. Denote
the reduced minimum modulus of . Then is closed if and only if . Clearly for . If , then
We successively obtain , , …, , and finally . Thus .
Let . There exists with such that . We obtain
Thus .
Let ; then . For a fixed between and we can write
for every . Hence and, as is arbitrary,
Set and for . Suppose that
Therefore (4.1) holds for every , and we obtain
The implication ”(1′) “ is clear. Note also that the above proof for and implies that
”(1′) (6′)“ Note that always. Suppose now that . We want to show that the range of is closed. Notice first that since . Let be such that . We have
Therefore the reduced minimum modulus of verifies . In particular, is closed if .
The implication ”(6′) “ is easy.
Let , and set for , . For every with we have
”“ Let between and . Using the Cauchy-Schwarz inequality and Lemma 4.2 we obtain
”(1) (9)“ Let . Let be a closed subspace of finite codimension in . With respect to the decomposition , the operator has the following matrix decomposition
Since is not closed, the range of the operator , acting on , is not closed. This means that is not an upper semi-Fredholm operator, and therefore there exists such that and . It follows that .
”(9) (4)“ Let be as in (9). Then , , and . We have
Set for and . Then for each and is orthogonal to . Hence
and , for each . We obtain
For we have ; hence
Therefore As is arbitrary, the proof of this implication is over.
”(4) (9′)“ Suppose that (4) holds. Let and let be a closed subspace of finite codimension in . Then there exists such that and . Let . Set , for . Then and for .
We shall prove by induction the following two claims :
Both claims are clearly true for . Suppose that both inequalities are true for some . Then, using several times the induction hypothesis, we have
Thus both (* ‣ 4A) and (** ‣ 4A) are true ; in particular we have
As was arbitrary, we obtain (9′).
”(9′) (8′)“ We have , so the range of this operator is in . The assertion (9′) implies that belongs to the essential spectrum of the restriction of to . Therefore .
The implication ”(8′) (8)“ is clear.
”(8) (7) (6)“ The statement (8) implies the following sequence of inequalities for the essential spectral radius and the essential norm of :
The proofs of implications ”(8′) (7′) (6′)“ are similar. The implications ”(8′) (5′) (5)“ are clear.
”(9′) (2′)“ Let be the operator restricted to . The condition (9′) implies that is in the boundary of the essential spectrum of the operator . According to , on the space the operators converge weakly to . The assertion (2′) can be proved exactly as in [4, Theorem 1] by replacing there by .
The implication ”(2′) (2)“ is clear.
The implication ”(11) (3)“ follows from . The proof is complete. ∎
B. Quantitative statements
Some remarks concerning the proof of Theorem 4.1 are in order.
Let . There exists with such that . Denote and
Since is orthogonal to , and , we have
Since , for each we have . Therefore
Let . Let be a complex Hilbert space. Let be closed subspaces of with intersection . Let , , and be the corresponding orthogonal projections. Denote .
(i) Suppose that . Then is uniformly convergent to , with
(ii) Suppose that . Then is strongly convergent to and we have (ASC), in all possible meanings of this paper.
C. Comparison with other estimates
Let be closed subspaces of , with intersection . Denote for .
In particular, we have quick uniform convergence whenever one of the cosine of the Dixmier angles is strictly less than one.
Moreover, for any sequence of integers such that , Theorem 4.1 shows that we have (QUC) for if and only if we have (QUC) for . Hence we have (QUC) for as soon as there exist integers such that . The following example shows that this sufficient condition for (QUC) is by far stronger than the condition .
Let be an orthonormal basis of , and let , and be the following closed subspaces of : , and . Then , , and . We obtain for any and . But : indeed if , then a straightforward computation shows that , so that .
The following proposition shows, even in a quantitative way, that the sufficient condition , for each , reminiscent of [31, Theorem 2.2] and [13, Theorem 2.7], implies that .
Let . Let be closed subspaces of with intersection . Denote for between and . Then
In particular, if each , .