Minimization of the Probabilistic p-frame Potential
Martin Ehler, Kasso A. Okoudjou
Introduction
Frames are overcomplete (or redundant) sets of vectors that serve to faithfully represent signals. They were introduced in by Duffin and Schaeffer , and reemerged with the advent of wavelets . Though the overcompleteness of frames precludes signals from having unique representation in the frame expansions, it is, in fact, the driving force behind the use of frames in signal processing .
In the finite dimensional setting, frames are exactly spanning sets. However, many applications require “custom-built” frames that possess additional properties which are dictated by these applications. As a result, the construction of frames with prescribed structures has been actively pursued. For instance, a special class called finite unit norm tight frames (FUNTFs) that provide a Parseval-type representation very similar to orthonormal bases, has been customized to model data transmissions . Since then the characterization and construction of FUNTFs and some of their generalizations have received a lot of attention . Beyond their use in applications, FUNTFs are also related to some deep open problems in pure mathematics such as the Kadison-Singer conjecture . FUNTFs appear also in statistics where, for instance, Tyler used them to construct -estimators of multivariate scatter . We elaborate more on the connection between the -estimators and FUNTFs in Remark 2.3. These -estimators were subsequently used to construct maximum likelihood estimators for the the wrapped Cauchy distribution on the circle in and for the angular central Gaussian distribution on the sphere in .
FUNTFs are exactly the minimizers of a functional called the frame potential . This was extended to characterize all finite tight frames in . Furthermore, in , finite tight frames with a convolutional structure, which can be used to model filter banks, have been characterized as minimizers of an appropriate potential. All these potentials are connected to other functionals whose extremals have long been investigated in various settings. We refer to for details and related results.
In Section 3, we give lower estimates on the -frame potentials, and prove that in certain cases their minimizers are FUNTFs, which possess additional properties and structure. In particular, if , we completely characterize the minimizers of the -frame potentials when for some positive integer . Moreover, when and , we characterize the minimizers of the -frame potentials, under a technical condition, which, we have only been able to establish when . We conjecture that this technical condition holds when . Finally in Section 4, we introduce probabilistic -frames that generalize the concepts of frames and -frames. We characterize the minimizers of probabilistic -frame potentials in terms of probabilistic -frames. The latter problem is solved completely for , and for all even integers . In particular, these last results generalize as well as the recently introduced notion of the probabilistic frame potential in .
Further relations to statistics: Besides the results on FUNTFs used in , and mentioned above, frame theory has essentially evolved independently of statistical fields such as statistical shape analysis and directional statistics . Nevertheless, there still exist several overlaps, and to the best of our knowledge, these overlaps have not yet been fully explored. Recently, frame theory has been used in directional statistics , where FUNTFs are utilized to investigate on statistical tests for directional uniformity and to model and analyze patterns found in granular rod experiments. We must point out that similar results were obtained earlier by Tyler in .
Probabilistic tight frames are multivariate probability distributions whose second moments’ matrix is a multiple of the identity, and they are used in to obtain approximate FUNTFs. The latter approximation procedure is connected to a classical problem in multivariate statistics, namely estimating the population covariance from a sample, which is closely related to the -estimators addressed in . The -frame potentials and their probabilistic counterparts that we consider in the sequel, are linked to the notion of shape measure, shape space, and mean shape used in statistical shape analysis. In Section 2.2, we establish a precise connection between the full Procrustes estimate of mean shape [23, Definition 3.3] which is the eigenvector corresponding to the largest eigenvalue of the frame operator. Moreover, this eigenvalue coincides with the upper frame redundancy as introduced in . The full Procrustes estimate of mean shape also saturates the upper frame inequality. Moreover, the -frame potentials form size measures as required in statistical shape analysis, and their minimizers among all collections of points on the sphere define a shape space modulo rotations.
We hope that the present paper will renew interests in more investigation on the role of frames and the -th frame potential in directional statistics and statistical shape analysis.
The p𝑝p-frame potential
To introduce frames and their elementary properties, we follow the textbook .
A collection of unit vectors is called equiangular if there exists a nonnegative constant such that , for .
Its adjoint operator is called the synthesis operator and given by
Using these operators, it is easy to see that is a frame if and only if the frame operator defined by
is positive, self-adjoint, and invertible. In this case, the following reconstruction formula holds
and , in fact, is a frame too, called the canonical dual frame. If is a frame, then is a finite tight frame. Moreover, note that is a FUNTF if and only if its frame operator is times the identity.
As mentioned in the introduction, the question of the existence and characterization of FUNTFs was settled in , where the frame potential, defined by
was introduced and used to give a characterization of its minimizers in terms of FUNTFs. More specifically, they prove the following result:
[2, Theorem 7.1] Let be fixed and consider the minimization of the frame potential among all collections of points on the sphere .
We shall prove in the sequel that the frame potential is just an example in a family of functionals defined on points on the sphere, and whose minimizers have approximation properties similar to those of the frame potential. But first, we briefly comment on the relation between FUNTFs and -estimators of multivariate scatter:
is the identity matrix. Whenever this is possible, the estimate of the population scatter matrix is then given by . We refer to for details. Note that implies that
forms a FUNTF. Moreover, is a FUNTF if and only if .
2 Definition of the p𝑝p-frame potential
Let be a positive integer, and . Given a collection of unit vectors , the -frame potential is the functional
When, , the definition reduces to
It is immediate that the following family of functionals can be seen as size measures:
cf. [14, Definition 3.3]. The average axis from the complex Bingham maximum likelihood estimator is the same as the full Procrustes estimate of mean shape for two-dimensional shapes, and the same holds for the complex Watson distribution, cf. [14, Sections 6.2 6.3]. If we assume that are centered around zero, then is given by the eigenvector corresponding to the largest eigenvalue of the frame operator of the normalized collection . Furthermore, one observes that this eigenvalue is the upper frame redundancy of as introduced in , which also coincides with the optimal upper frame bound in (1). Therefore, the full Procrustes estimate of mean shape satisfies the upper frame inequality with equality. Moreover, we have observed that the root mean square of the full Procrustes estimate of mean shape is .
For each , the -frame potential is induced by the conservative force , for , where
is a central force between the ‘particles’ and that we call the -frame force.
is differentiable and satisfies . This is sufficient to verify that the potential , defined for , satisfies , where is held fixed. Thus, is a conservative vector field. The physical meaningful potential is in fact given by
where we used that . Consequently, the -frame potential is induced by the conservative central force . ∎
As a consequence of the above lemma, are in equilibrium under the -frame force if they minimize the -frame potential among all collections of points on the sphere. Note that such a collection of equilibria modulo rotations form a shape space.
We will use repeatedly the fact that for a fixed , the -frame potential is a decreasing and continuous function of .
Lower estimates for the p𝑝p-frame potential
We start this section with a few elementary results about the minimizers of the -frame potential as well as their connection to -designs. In fact, potentials on the sphere, and -designs have been well investigated . However, one of the key differences between -designs and our -frame potential is that the former is considered only for positive integers while the latter is investigated for .
If is an even integer, one can use Welch’s results to conclude that, for ,
We shall verify that this estimate is not optimal for small , by proving an estimate for when . The following Proposition first appeared in :
Let , , and , then
and equality holds if and only if is an equiangular FUNTF.
For , Hölder’s inequality yields
Raising to the -th power and applying leads to
Using the fact that (see Theorem 2.2) implies that
see, , for details. Consequently, if is an equiangular FUNTF, then (7) holds with equality.
On the other hand, if equality holds in (7), then and is a FUNTF due to Theorem 2.2. Moreover, the Hölder estimate (8) must have been an equality which means that for , and some constant . Thus, the FUNTF must be equiangular. ∎
By comparing (6) with (7), it is easily seen that the Welch bound is not optimal for small :
Let and be an even integer. If , then
The condition on implies , and adding (N-1)\big{(}\frac{N-d}{d(N-1)}\big{)}^{k}>0 to the right hand side leads to
Multiplication by and Proposition 3.1 then yield (11). ∎
The estimate in Proposition 3.1 is sharp if and only if an equiangular FUNTF exists. In [31, Sections 4 6], construction (and hence existence) of equiangular FUNTFs was established when . For general and , a necessary condition for existence of equiangular FUNTFs is given, and it is conjectured that the conditions are sufficient as well. The authors essentially provide on upper bound on that depends on the dimension . Therefore, Proposition 3.1 might not be optimal when the redundancy is much larger than .
2 Relations to spherical t𝑡t-designs
for all homogeneous polynomials of total degree equals or less than in variables and where denotes the uniform surface measure on normalized to have mass one. The following result is due to [34, Theorem 8.1] (see , for similar results).
[34, Theorem 8.1] Let be an even integer and , then
and equality holds if and only if is a spherical -design.
3 Optimal configurations for the p𝑝p-frame potential
We first use Theorem 2.2 to characterize the minimizers of the -frame potential for provided that the number of points is a multiple of the dimension :
Let and assume that for some positive integer . Then the minimizers of the -frame potential are exactly the copies of any orthonormal basis modulo multiplications by . The minimum of (5), over all sets of unit norm vectors, is .
If we fix a collection of vectors , then the frame potential is a decreasing function in . Therefore,
We now consider the -frame potential for . The case is covered by Theorem 2.2. Note that any FUNTF with vectors is equiangular . Hence, the case is settled by Proposition 3.1, so we focus on .
One easily verifies that, for , an orthonormal basis plus one repeated vector and an equiangular FUNTF have the same -frame potential . Under the assumption that those two systems are exactly the minimizers of , the next result will give a complete characterization of the minimizers of , for . However, we have only been able to establish the validity of this assumption when , cf. Corollary 3.7.
Let and set . Let , and assume that with equality holds if and only if is an orthonormal basis plus one repeated vector or an equiangular FUNTF. Then,
for , then for any , we have and equality holds if and only if is an orthonormal basis plus one repeated vector,
for , then for any , we have and equality holds if and only if is an equiangular FUNTF.
Under the assumptions of the Theorem, let , then
Consequently, . Since an orthonormal basis plus one repeated vector minimizes the -frame potential for , it must also minimize for , which proves .
Assume now that . Choose such that , Hölder’s inequality yields
Since we assume that holds, we have , which leads to
This concludes the proof of . By applying (10), one then checks that an equiangular FUNTF satisfies with equality.
The “only” part comes from the fact that the Hölder inequality becomes an equality only if the sequences are linearly dependent. This means that the are equiangular. They must then satisfy . Thus, by (10) they form an equiangular FUNTF . ∎
When we can in fact verify the main hypothesis of Theorem 3.6, which leads to the following result:
Let , and set . Then,
and equality holds if and only if is an orthonormal basis plus one repeated vector or an equiangular FUNTF.
for , then for any , we have and equality holds if and only if is an orthonormal basis plus one repeated vector,
for , then for any , we have and equality holds if and only if is an equiangular FUNTF.
The minimum of , for is plotted in Figure 1.
Clearly and follows from Theorem 3.6 once the minimizers of are characterized.
Without loss of generality, let be the smallest angle between , , and and let be the second smallest angle between them. This yields, of course, .
Case 1: For , we have
Since , the -frame potential is differentiable in and , and its critical points are
This implies that either or and . In the first case, we have a maximum since it implies . The latter case means that two points are identical and the third one is perpendicular which is a potential minimum of the -frame potential.
Case 2: We can assume that (otherwise we are in Case 1). We can further assume that (if . Otherwise, substitute with . We now have
By subtracting one equation from the other and raising to the second power, we obtain
where and . Since for all , we have and because .
We consider the function on . achieves its maximum at and is convex, cf. Figure 2.
Therefore, for and , we have if and only if or and . For , we have and yields . The case leads to which implies . This is equivalent to . Since we assume , we obtain .
on , cf. Figure 2. Its derivative
By substituting , we obtain, for
where . To show that has only one extremal point on , we differentiate
The term vanishes if and only if
Hence, and does not have any other zeros on . This means that has only one extremal point and can then only have two zeros on . The zero at corresponds to a minimum. This means that the zero of at is a maximum of . Hence the other zero of g is between and . However, this other zero corresponds to a maximum of . The minimum of can thus be at or . This implies or . It is easy to verify that would lead to , and yields . Thus, the minimum of the -frame potential corresponds to either an orthonormal basis plus one repeated element (, ) or an equiangular FUNTF (). One easily checks that both situations lead to the same global minimum.
In view of Theorem 3.6 and Corollary 3.7, we have the following conjecture:
Let and . Then
and equality holds if and only if is an orthonormal basis plus one repeated vector or an equiangular FUNTF.
One can check that , for . According to Proposition 3.1, the minimizers of the -frame potential for are exactly the equiangular FUNTFs. Thus, our conjecture essentially addresses the range .
The probabilistic p𝑝p-frame potential
The present section is dedicated to introducing a probabilistic version of the previous section. We shall consider probability distributions on the sphere rather than finite point sets. Let denote the collection of probability distributions on the sphere with respect to the Borel sigma algebra .
We begin by introducing the probabilistic -frame which generalizes the notion of probabilistic frames introduced in .
We call a tight probabilistic -frame if and only if we can choose .
Due to Cauchy-Schwartz, the upper bound always exists. Consequently, in order to check that is a probabilistic -frame one only needs to focus on the lower bound .
Since the uniform surface measure on is invariant under orthogonal transformations, one can easily check that it constitutes a tight probabilistic -frame, for any .
Given a probability measure , we call
the analysis operator. It is trivially seen that
for all . The dual of is called synthesis operator and is given by
where , and . In fact, is well-defined and bounded operator on all where . Indeed, for we have
Given , the second moments matrix of is the matrix defined by
If , then is a probabilistic -frame if and only if is onto.
As mentioned earlier the upper bound in the probabilistic -frame definition always holds. So we only need to show the equivalence between the lower bound and the surjectivity of .
for all . A contradiction argument leads to for all which implies that . Thus is surjective. ∎
The second moments matrix can also be used to show that a probabilistic -frame gives rise to a reconstruction formula that extends the finite frame expansion in (3). In addition, the next result generalizes the reconstruction formula for tight probabilistic frames obtained in [16, Lemma 3.7].
The result follows by noticing that . ∎
The above result motivates the following definition:
a) If is probabilistic frame, then it is a probabilistic -frame for all . Conversely, if is a probabilistic -frame for some , then it is a probabilistic frame.
where we have used the fact that for , . This conclude the proof of a).
We are particularly interested in tight probabilistic -frame potentials, which we seek to characterize in terms of minimizers of appropriate potentials. This motivates the following definition:
For and , the probabilistic -frame potential is defined by
From the weak-star-compactness of the collection of all probability distributions on the sphere, we can deduce that admits a minimizer which satisfies
We now turn to the minimizers of the probabilistic frame potential . In the process, we extend some ideas developed in to the probabilistic frame potential.
Let and let be a minimizer of (15), then
, for all ,
, for all .
The proof will use the following observation. Let be a probability measure on and choose a measure , such that and , for all . Let us also introduce the notation
We thus have , for all , which implies .
We now prove using a contradiction argument. In particular, assume that does not hold. This implies that there are such that
Set . Let be an open ball around in and so small that and that the oscillation of on is smaller than . Let . One can check that the measure defined by
satisfies , and . Hence, . On the other hand, we can estimate
This is a contradiction to and implies that there is a constant such that , for all . We still have to verify that the constant is in fact :
The proof of is similar to the one above, and so we omit it. ∎
The following result is an immediate consequence of Proposition 4.7.
Let and let be a minimizer of (15), then
, for all ,
directly follows from in Proposition 4.7.
We can now characterize the minimizers of the probabilistic -frame potential when . In fact, we shall show that these minimizers are discrete probability measures, and the following theorem is the analogue of Proposition 3.5:
Let , then the minimizers of (15) are exactly those probability distributions that satisfy both,
The measure in Theorem 4.9 denotes the counting measure of the set .
Since , for , we have . In [16, Theorem 3.10] it was shown that the normalized counting measure of an orthonormal basis minimizes . Due to , we obtain that also minimizes and hence .
In the following, we prove that all minimizers of are essentially induced by an orthonormal basis. Let be a minimizer and let . We first show that . The implications if and only if and if and only if are trivial.
Suppose now that and , then there exist and such that
and .
for all and , .
By using , this implies
For even integers , we can give the minimum of and characterize its minimizers. The following theorem generalizes Theorem 3.4. Moreover, note that the bounds are now sharp, i.e., for any even integer , there is a probabilistic tight -frame:
Let be an even integer. For any probability distribution on ,
and equality holds if and only if is a probabilistic tight -frame.
Let and consider the Gegenbauer polynomials defined by
is an orthogonal basis for the collection of polynomials of degree less or equal to on the interval $$ with respect to the weight
The polynomials , an even integer, can be represented by means of
It is known (see, e.g., ) that , , and is given by
see . Note that the probability measures with finite support are weak star dense in . Since is continuous, we obtain, for all ,
From the results in , one can deduce that
We still have to address the “if and only if” part. Equality holds if and only if satisfies
is a polynomial in . In fact, the integral resolves in the polynomial’s coefficients. These two observations enable us to follow the lines in , and we can conclude the proof. ∎
One may speculate that Theorem 4.10 could be extended to that are not even integers. This is not true in general. For and , for instance, the equiangular FUNTF with elements induces a smaller potential than the uniform distribution. The uniform distribution is a probabilistic tight -frame, but the equiangular FUNTF is not.
Acknowledgements
The authors would like to thank C. Bachoc, W. Czaja, C. Wickman, and W. S. Yu for discussions leading to some of the results presented here. M. Ehler was supported by the Intramural Research Program of the National Institute of Child Health and Human Development and by NIH/DFG Research Career Transition Awards Program (EH 405/1-1/575910). K. A. Okoudjou was partially supported by ONR grant N000140910324, by RASA from the Graduate School of UMCP, and by the Alexander von Humboldt foundation.