Rounding Sum-of-Squares Relaxations
Boaz Barak, Jonathan Kelner, David Steurer
Introduction
One of the reasons for this paucity of positive results is that we have relatively few tools to round such convex hierarchies. A rounding algorithm maps a solution to the relaxation to a solution to the original program. While the name derives from the prototypical case of relaxing an integer program to a linear program by allowing the variables to take non-integer values, we use “rounding algorithm” for any mapping from relaxation solutions to actual solutions, even in cases where the actual solutions are themselves non-integer. In the case of a hierarchy, the relaxation solution satisfies more constraints, but we do not always know how to take advantage of this when rounding. For example, [ARV04] used a very sophisticated analysis to get better rounding when the solution to a Sparsest Cut relaxation satisfies a constraint known as triangle inequalities, but we have no general tools to use the additional constraints that come from higher levels of the hierarchies, nor do we know if these can help in rounding or not. This lack of rounding techniques is particularly true for the Sum of Squares (SOS, also known as Lasserre) hierarchy [Par00, Las01]. While it is common in the TCS community to use Lasserre to describe the primal version of this SDP, and Sum of Squares (SOS) to describe the dual, in this paper we use the more descriptive SOS name for both programs. We note that in all the applications we consider, strong duality holds, and so these programs are equivalent. This is the strongest variant of the canonical semidefinite programming hierarchies, and has recently shown promise to achieve tasks beyond the reach of weaker hierarchies [BBH+12]. But there are essentially no general rounding tools that take full advantage of its power. The closest general tool we are aware of is the repeated conditioning methods of [BRS11, GS11], though these can be implemented in weaker hierarchies too and so do not seem to use the full power of the SOS hierarchy. However, this technique does play a role in this work as well.
In this work we propose a general approach to rounding SOS hierarchies, and instantiate this approach in two cases, giving new algorithms making progress on natural variants of two longstanding problems. Our approach is based on the intimate connection between the SOS hierarchy and the “Positivstellensatz”/“Sum of Squares” proof system. This connection was used in previous work for either negative results [Gri01b, Gri01a, Sch08], or positive results for specific instances [BBH+12, OZ13, KOTZ14], translating proofs of a bound on the actual value of these instances into proofs of bounds on the relaxation value. In contrast, we use this connection to give explicit rounding algorithms for general instances of certain computational problems.
Our work uses the Sum of Squares (SOS) semidefinite programming hierarchy and in particular its relationship with the Sum of Squares (or Positivstellensatz) proof system. We now briefly review both the hierarchy and proof system. See the introduction of [OZ13] and the monograph [Lau09] for a more in depth discussion of these concepts and their history. Underlying both the SDP and proof system is the natural approach to prove that a real polynomial is nonnegative via showing that it equals a sum of squares: for some polynomials . The question of when a nonnegative polynomial has such a “certificate of non-negativity” was studied by Hilbert who realized this doesn’t always hold and asked (as his problem) whether a nonnegative polynomial is always a sum of squares of rational functions. This was proven to be the case by Artin, and also follows from the more general Positivstellensatz (or “Positive Locus Theorem”) [Kri64, Ste74].
The Positivstellensatz/SOS proof system of Grigoriev and Vorobjov [GV01] is based on the Positivstellensatz as a way to refute the assertion that a certain set of polynomial equations
can be satisfied by showing that there exists some polynomials and a sum of squares polynomial such that
([GV01] considered inequalities as well, although in our context one can always restrict to equalities without loss of generality.) One natural measure for the complexity of such proof is the degree of the polynomials and .
The sum of squares semidefinite program was proposed independently by several authors [Sho87, Par00, Nes00, Las01] One way to describe it is as follows. If the set of equalities (1.1) is satisfiable then in particular there exists some random variable over such that
That is, is some distribution over the non-empty set of solutions to (1.1).
If is the constant polynomial then
Until recently, this relation was mostly used for negative results, translating proof complexity lower bounds into integrality gap results for the SOS hierarchy [BBH+12, OZ13, KOTZ14]. However, in 2012 Barak, Brandão, Harrow, Kelner, Steurer and Zhou [BBH+12] used this relation for positive results, showing that the SOS hierarchy can in fact solve some interesting instances of the Unique Games maximization problem that fool weaker hierarchies. Their idea was to use the analysis of the previous works that proved these integrality gaps for weaker hierarchies. Such proofs work by showing that (a) the weaker hierarchy outputs a large value on this particular instance but (b) the true value is actually small. [BBH+12]’s insight was that oftentimes the proof of (b) only uses arguments that can be captured by the SOS/Positivstellensatz proof system, and hence inadvertently shows that the SOS SDP value is actually small as well. Some follow up works [OZ13, KOTZ14] extended this to other instances, but all these results held for very specific instances which have been proven before to have small objective value.
While the SOS hierarchy is relevant to many algorithmic applications, some recent work focused on its relation to Khot’s Unique Games Conjecture (UGC) [Kho02]. On a high level, the UGC implies that the basic semidefinite program is an optimal efficient algorithm for many problems, and hence in particular using additional constant or polylogarithmic levels of the SOS hierarchy will not help. More concretely, as discussed in Section 1.3 below, the UGC is closely related to the question of how hard it is to find sparse (or “analytically sparse”) vectors in a given subspace. Our work shows how the SOS hierarchy can be useful in general, and in particular gives strong average-case results and nontrivial worst-case results for finding sparse vectors in subspaces. Therefore, it can be considered as giving some (far from conclusive) evidence that the UGC might be false.
2 Optimizing polynomials with nonnegative coefficients bounded spectral norm
Our first result yields an additive approximation to this optimization problem for polynomials with nonnegative coefficients, when the value is scaled by the spectral norm of an associated matrix. If is an -variate degree- homogeneous polynomial with nonnegative coefficient, then it can be represented by a tensor such that for every . It is convenient to state our result in terms of this tensor representation:
There is an algorithm , based on levels of the SOS hierarchy, such that for every even The algorithm easily generalizes to polynomials of odd degree and to non-homogenous polynomials, see Remark 3.5. and nonnegative ,
where denotes the standard dot product, and denotes the spectral norm of , when considered as an matrix.
Note that the algorithm of Theorem 1.2 only uses a logarithmic number of levels, and thus it shows that this fairly natural polynomial optimization problem can be solved in quasipolynomial time, as opposed to the exponential time needed for optimizing over general polynomials of degree . Indeed, previous work on the convergence of the Lasserre hierarchy for general polynomials [DW12] can be described in our language here as trying to isolate a solution in the support of the distribution, and this generally requires a linear number of levels. Obtaining the logarithmic bound here relies crucially on constructing a “combined” solution that is not necessarily in the support. The algorithm is also relatively simple, and so serves as a good demonstration of our general approach.
or to simply distinguish between the case that this quantity is at least and the case that is separable. The best result known in this area is the paper [BCY11] mentioned above, which solved the distinguishing variant of quantum separability problem in the case that measurements are restricted to so-called Local Operations and one-way classical communication (one-way LOCC) operators. However, they did not have an rounding algorithm, and in particular did not solve the problem of actually finding a separable state that maximizes the probability of acceptance of a given one-way LOCC measurement. The techniques of this work were used by Brandão and Harrow [BH13] to solve the latter problem, and also greatly simplify the proof of [BCY11]’s result, which originally involved relations between several measures of entanglement proved in several papers.The paper [BH13] was based on a previous version of this work [BKS12] that contained only the results for nonnegative tensors.
Relation to small set expansion. Nonnegative tensors also arise naturally in some applications, and in particular in the setting of small set expansion for Cayley graphs over the cube, which was our original motivation to study them. In particular, one corollary of our result is:
We discuss the derivation and the meaning of this corollary in Section 6 but note that the condition of having small value seems reasonable. Having implies that the graph is a small set expander, and in particular the known natural examples of Cayley graphs that are small set expanders, such as the noisy Boolean hypercube and the “short code” graph of [BGH+12] have . Thus a priori one might have thought that a graph that is hard to distinguish from small set expanders would have a small value of .
3 Optimizing hypercontractive norms and finding analytically sparse vectors
Finding a sparse nonzero vector inside a dimensional linear subspace is a natural task arising in many applications in machine learning and optimization (e.g., see [DH13] and the references therein). Related problems are known under many names including the “sparse null space”, “dictionary learning”, “blind source separation”, “min unsatisfy”, and “certifying restricted isometry property” problems. (These problems all have the same general flavor but differ on various details such as worst-case vs. average case, affine vs. linear subspaces, finding a single vector vs. a basis, and more.) Problems of this type are often NP-hard, with some hardness of approximation results known, and conjectured average-case hardness (e.g., see [ABSS97, KZ12, GN10] and the references therein).
We consider a natural relaxation of this problem, which we call the analytically sparse vector problem (ASVP), which assumes the input subspace (almost) contains an actually sparse vector, but allows the algorithm to find a vector that is only “analytically sparse” in the sense that is large. More formally, for and , we say that a vector is -sparse if . That is, a vector is -sparse if it has the same -norm vs -norm ratio as a vector of measure at most .
This is a natural relaxation, and similar conditions have been considered in the past. For example, Spielman, Wang, and Wright [SWW12] used in their work on dictionary learning a subroutine finds a vector in a subspace that maximizes the ratio (which can be done efficiently via linear programs). However, because any subspace of dimension contains an -sparse vector, this relaxation can only detect the existence of vectors that are supported on less than coordinates. Some works have observed that the ratio is a much better proxy for sparsity [ZP01, DH13], but computing it is a non-convex optimization problem for which no efficient algorithm is known. Similarly, the ratio is a good proxy for sparsity for subspaces of small dimension (say ) but it is non-convex, and it is not known how to efficiently optimize it. It seems that what makes our relaxation different from the original problem is not so much the qualitative issue of considering analytically sparse vectors as opposed to actually sparse vectors, but the particular choice of the ratio, which on one hand seems easier (even if not truly easy) to optimize over than the ratio, but provides better guarantees than the ratio. However, this choice does force us to restrict our attention to subspaces of low dimension, while in some applications such as certifying the restricted isometry property, the subspace in question is often the kernel of a “short and fat” matrix, and hence is almost full dimensional. Nonetheless, we believe it should be possible to extend our results to handle subspaces of higher dimension, perhaps at the some mild cost in the number of rounds.
Nevertheless, because is a degree polynomial, the problem of maximizing it for of unit norm amounts to a polynomial maximization problem over the sphere, that has a natural SOS program. Indeed, [BBH+12] showed that this program does in fact yield a good approximation of this ratio for random subspaces. As we show in Section 5, we can use this to improve upon the results of [DH13] and find planted sparse vectors in random subspaces that are of not too large a dimension:
In particular, we note that this recovers a planted vector with up to nonzero coordinates when , and it can recover vectors with more than the nonzero coordinates that are necessary for existing techniques whenever .
Perhaps more significantly, we prove the following nontrivial worst-case bound for this problem:
There is a polynomial-time algorithm , based on levels of the SOS hierarchy, that on input a -dimensional subspace such that there is a -vector with at most nonzero coordinates, outputs an -sparse vector in .
Moreover, this holds even if is not completely inside but only satisfies , for some absolute constant , where is the projector to .
The condition that the vector is can be significantly relaxed, see Remark 4.12. Theorem 4.1 is also motivated by the Small Set Expansion problem. The current best known algorithms for Small Set Expansion and Unique Games [ABS10] reduce these problems into the task of finding a sparse vector in a subspace, and then find this vector using brute force enumeration. This enumeration is the main bottleneck in improving the algorithms’ performance. This is the only step that takes super-polynomial time in [ABS10]’s algorithm for Small Set Expansion. Their algorithm for Unique Games has an additional divide and conquer step that takes subexponential time, but, in our opinion, seems less inherently necessary. Thus we conjecture that if the sparse-vector finding step could be sped up then it would be possible to speed up the algorithm for both problems. [BBH+12] showed that, at least for the Small Set Expansion question, finding an analytically sparse vector would be good enough. Using their work we obtain the following corollary of Theorem 1.5:
There is an algorithm that given an -vertex graph that contains a set of size with expansion at most , outputs a set of measure with expansion bounded away from , i.e., , where is the dimension of the eigenspace of ’s random walk matrix corresponding to eigenvalues larger than .
The derivation and meaning of this result is discussed in Section 6. We note that this is the first result that gives an approximation of this type to the small set expansion in terms of the dimension of the top eigenspace, as opposed to an approximation that is polynomial in the number of vertices.
4 Related work
Our paper follows the work of [BBH+12], that used the language of pseudoexpectation to argue that the SOS hierarchy can solve specific interesting instances of Unique Games, and perhaps more importantly, how it is often possible to almost mechanically “lift” arguments about actual distributions to the more general setting of pseudodistribution. In this work we show how the same general approach be used to obtain positive results for general instances.
The fact that LP/SDP solutions can be viewed as expectations of distributions is well known, and several rounding algorithms can be considered as trying to “reverse engineer” a relaxation solution to get a good distribution over actual solutions.
Techniques such as randomized rounding, the hyperplane rounding of [GW95], and the rounding for TSP [GSS11, AKS12] can all be viewed in this way. One way to summarize the conceptual difference between our techniques and those approaches is that these previous algorithms often considered the relaxation solution as giving moments of an actual distribution on “fake” solutions. For example, in [GW95]’s Max Cut algorithm, where actual solutions are modeled as vectors in , the SDP solution is treated as the moment matrix of a Gaussian distribution over real vectors that are not necessarily -valued. Similarly in the TSP setting one often considers the LP solution to yield moments of a distribution over spanning trees that are not necessarily TSP tours. In contrast, in our setting we view the solution as providing moments of a “fake” distribution on actual solutions.
Treating solutions explicitly as “fake distributions” is prevalent in the literature on negative results (i.e., integrality gaps) for LP/SDP hierarchies. For hierarchies weaker than SOS, the notion of “fake” is different, and means that there is a collection of local distributions, one for every small subset of the variables, that are consistent with one another but do not necessarily correspond to any global distribution. Fake distributions are also used in some positive results for hierarchies, such as [BRS11, GS11], but we make this more explicit, and, crucially, make much heavier use of the tools afforded by the Sum of Squares relaxation.
The notion of a “combining algorithm” is related to the notion of polymorphisms [BJK05] in the study of constraint satisfaction problems. A polymorphism is a way to combine a number of satisfying assignments of a CSP into a different satisfying assignments, and some relations between polymorphism, their generalization to approximation problems, rounding SDP’s are known (e.g., see the talk [Rag10]). The main difference is polymorphisms operate on each bit of the assignment independently, while we consider here combining algorithms that can be very global.
In a follow up (yet unpublished) work, we used the techniques of this paper to obtain improved results for the sparse dictionary learning problem, recovering a set of vectors from random samples of -sparse linear combinations of them for any , improving upon previous results that required [SWW12, AGM13, AAJ+13].
5 Organization of this paper
In Section 2 we give a high level overview of our general approach, as well as proof sketches for (special cases of) our main results. Section 3 contains the proof of Theorem 1.2— a quasipolynomial time algorithm to optimize polynomials with nonnegative coefficients over the sphere. Section 4 contains the proof of Theorem 1.5— a polynomial time algorithm for an -approximation of the “analytical sparsest vector in a subspace” problem. In Section 5 we show how to use the notion of analytical sparsity to solve the question of finding a “planted” sparse vector in a random subspace. Section 6 contains the proofs of Corollaries 1.3 and 1.6 of our results to the small set expansion problem. Appendix A contains certain technical lemmas showing that pseudoexpectation operators obey certain inequalities that are true for actual expectations. Appendix C contains a short proof (written in classical notation, and specialized to the real symmetric setting) of [BCY11, BH13]’s result that the SOS hierarchy yields a good approximation to the acceptance probability of QMA(2) verifiers / measurement operators that have bounded one-way LOCC norm. Appendix B shows a simpler algorithm for the case that the verifier satisfies the stronger condition of a bounded (Frobenius) norm. For the sake of completeness, Appendix D reproduces the proof from [BBH+12] of the relation between hypercontractive norms and small set expansion. Our papers raises many more questions than it answers, and some discussion of those appears in Section 7.
6 Notation
We will use linear subspaces of the form where is a finite set with an associated measure . The -norm of a vector is defined as . Similarly, the inner product of is defined as . We will only use two measures in this work: the counting measure, where for every , and the uniform measure, where for all . (The norms corresponding to this measure are often known as the expectation norms.)
We will use vector notation (i.e., letters such as , and indexing of the form ) for elements of subspaces with the counting measure, and function notation (i.e., letters such as and indexing of the form ) for elements of subspaces with the uniform measure. The dot product notation will be used exclusively for the inner product with the counting measure.
Pseudoexpectations
(In the context of optimization, to enforce the inequality constraint , it is always possible to add an auxiliary variable and then enforce the equality constraint .) Appendix A contains several useful facts about pseudoexpectations.
Overview of our techniques
Traditionally to design a mathematical-programming based approximation algorithm for some optimization problem , one first decides what the relaxation is— i.e., whether it is a linear program, semidefinite program, or some other convex program, and what constraints to put in. Then, to demonstrate that the value of the program is not too far from the actual value, one designs a rounding algorithm that maps a solution of the convex program into a solution of the original problem of approximately the same value. Our approach is conceptually different— we design the rounding algorithm first, analyze it, and only then come up with the relaxation.
Initially, this does not seem to make much sense— how can you design an algorithm to round solutions of a relaxation when you don’t know what the relaxation is? We do this by considering an idealized version of a rounding algorithm which we call a combining algorithm. Below we discuss this in more detail but roughly speaking, a combining algorithm maps a distribution over actual solutions of into a single solution (that may or may not be part of the support of this distribution). This is a potentially much easier task than rounding relaxation solutions, and every rounding algorithm yields a combining algorithm. In the other direction, every combining algorithm yields a rounding algorithm for some convex programming relaxation, but in general that relaxation could be of exponential size. Nevertheless, we show that in several interesting cases, it is possible to transform a combining algorithm into a rounding algorithm for a not too large relaxation that we can efficiently optimize over, thus obtaining a feasible approximation algorithm. The main tool we use for that is the Sum of Squares proof system, which allows to lift certain arguments from the realm of combining algorithms to the realm of rounding algorithms.
We now explain more precisely the general approach, and then give an overview of how we use this approach for our two applications— finding “analytically sparse” vectors in subspaces, and optimizing polynomials with nonnegative coefficients over the sphere.
Consider a general optimization problem of minimizing some objective function in some set , such as the dimensional Boolean hypercube or the unit sphere. A convex relaxation for this problem consists of an embedding that maps elements in into elements in some convex domain, and a suitable way to generalize the objective function to a convex function on this domain. For example, in linear programming relaxations we typically embed into the set , while in semidefinite programming relaxations we might embed into the set of positive semidefinite matrices using the map where . Given this embedding, we can use convex programming to find the element in the convex domain that maximizes the objective, and then use a rounding algorithm to map this element back into the domain in a way that approximately preserves the objective value.
A combining algorithm takes as input a distribution over solutions in and maps it into a single element of , such that the objective value of is approximately close to the expected objective value of a random element in . Every rounding algorithm yields a combining algorithm . The reason is that if there is some embedding mapping elements in into some convex domain , then for every distribution over , we can define to be . By convexity, will be in and its objective value will be at most the average objective value of an element in . Thus if we define to output then will be a combining algorithm with approximation guarantees at least as good as ’s.
We now turn to giving a high level overview of our results. For the sake of presentations, we focus on certain special cases of these two applications, and even for these cases omit many of the proof details and only provide rough sketches of the proofs. The full details can be found in Sections 5, 4 and 3.
We consider the following natural problem, which was also studied by Demanet and Hand [DH13]. Let be a sparse function over some universe of size . That is, is supported on at most coordinates for some . Let be the subspace spanned by and random (say Gaussian) functions . Can we recover from any basis for ?
Demanet and Hand showed that if is very small, specifically , then would be the most -sparse function in , and hence (as mentioned above) can be recovered efficiently by running linear programs. The SOS framework yields a natural and easy to describe algorithm for recovering as long as is a sufficiently small constant and the dimension is at most . The algorithm uses the SOS program for finding the most -sparse function in , which, as mentioned above, is simply the polynomial optimization problem of maximizing over in the intersection of and the unit Euclidean sphere.
for some constant . Therefore using triangle inequality, and using the fact that , it must hold that
In particular this implies that if we apply a singular value decomposition (SVD) to the second moment matrix of (i.e., ) then the top eigenvector will have correlation with , and hence we can simply output it as our solution.
To make this combining algorithm into a rounding algorithm we use the result of [BBH+12] that showed that (2.1) can actually be proven via a sum of squares argument. Namely they showed that there is a degree sum of squares polynomial such that
(2.4) implies that even if is merely a pseudodistribution then it must satisfy (2.1). (When the latter is raised to the fourth power to make it a polynomial inequality.) We can then essentially follow the argument, proving a version of (2.2) raised to the power by appealing to the fact that pseudodistributions satisfy Hölder’s inequality, (Corollary A.11) and hence deriving that will satisfy (2.3), with possibly slightly worse constants, even when it is only a pseudodistribution.
In Section 5, we make this precise and extend the argument to obtain nontrivial (but weaker) guarantees when . We then show how to use an additional correction step to recover the original function up to arbitrary accuracy, thus boosting our approximation of into an essentially exact one.
2 Finding “analytically sparse” vectors in general subspaces
We now outline the ideas behind the proof of Theorem 4.1— finding analytically sparse vectors in general (as opposed to random) subspaces. This is a much more challenging setting than random subspaces, and indeed our algorithm and its analysis is more complicated (though still only uses a constant number of SOS levels), and at the moment, the approximation guarantee we can prove is quantitatively weaker. This is the most technically involved result in this paper, and so the reader may want to skip ahead to Section 2.3 where we give an overview of the simpler result of optimizing over polynomials with nonnegative coefficients.
We consider the special case of Theorem 4.1 where we try to distinguish between a YES case where there is a valued -sparse function that is completely contained in the input subspace, and a NO case where every function in the subspace has its four norm bounded by a constant times its two norm. That is, we suppose that we are given some subspace of dimension and a distribution over functions in such that for every in the support of , and . The goal of our combining algorithm to output some function such that . (Once again, we use the expectation inner product and norms, with uniform measure over .)
Since the ’s correspond to sets of measure , we would expect the inner product of a typical pair (which equals the measure of the intersection of the corresponding sets) to be roughly . Indeed, one can show that if the average inner product is then it’s easy to find such a desired function . Intuitively, this is because in this case the distribution of sets does not have an equal chance to contain all the elements in , but rather there is some set of coordinates which is favored by . Roughly speaking, that would mean that a random linear combination of these functions would have most of its mass concentrated inside this small set , and hence satisfy . But it turns out that letting be a random gaussian function matching the first two moments of is equivalent to taking such a random linear combination, and so our combining algorithm can obtain this using moment information alone.
Our combining algorithm will also try all coordinate projection functions. That is, let be the function such that equals if and equals otherwise, (and hence under our expectation inner product ). The algorithms will try all functions of the form where is the projector to the subspace . Fairly straightforward calculations show that -norm squared of such a function is expected to be , and it turns out in our setting we can assume that the norm is well concentrated around this expectation (or else we’d be able to find a good solution in some other way). Thus, if coordinate projection fails then it must hold that
It turns out that (2.5) implies some nontrivial constraints on the distribution . Specifically we know that
But since and is symmetric, the RHS is equal to
where the last inequality uses Cauchy–Schwarz. If we square this inequality we get that
But since is a projector satisfying , we can use (2.5) and obtain that
Equation (2.6), contrasted with the fact that , means that the inner product of two random functions in is somewhat “surprisingly unconcentrated”, which seems to be a nontrivial piece of information about . Interestingly, this part of the argument does not require to be , and some analogous “non-concentration” property of can be shown to hold for a hard to round for any . However, we currently know how to take advantage of this property to obtain a combining algorithm only in the case that . Indeed, because the ’s are nonnegative functions, if we pick a random and consider the distribution where the probability of every function is reweighed proportionally to , then intuitively that should increase the probability of pairs with large inner products. Indeed, as we show in Lemma A.4, one can use Hölder’s inequality to prove that there exist such that under the distribution where every element is reweighed proportionally to , it holds that
(2.7) and (2.6) together imply that , which, as mentioned above, means that we can find a function satisfying by taking a gaussian function matching the first two moments of .
Once again, this combining algorithm can be turned into an algorithm that uses levels of the SOS hierarchy. The main technical obstacle (which is still not very hard) is to prove another appropriate generalization of Hölder’s inequality for pseudoexpectations (see Lemma A.4). Generalizing to the setting that in the YES case the function is only approximately in the vector space is a bit more cumbersome. We need to consider apart from the function that is obtained by first projecting to the subspace and then “truncating” it by rounding each coordinate where is too small to zero. Because this truncation operation is not a low degree polynomial, we include the variables corresponding to as part of the relaxation, and so our pseudoexpectation operator also contains the moments of these functions as well.
3 Optimizing polynomials with nonnegative coefficients
We now consider the task of maximizing a polynomial with nonnegative coefficients over the sphere, namely proving Theorem 3.1. We consider the special case of Theorem 3.1 where the polynomial is of degree . That is, we are given a parameter and an nonnegative matrix with spectral norm at most and want to find an additive approximation to the maximum of
over all with , where in this section we let be the standard (counting) Euclidean norm .
One can get some intuition for this problem by considering the case where is valued and is valued for some . In this case one can think of is a -uniform hypergraph on vertices and as a subset that maximizes the number of edges inside divided by , and so this problem is related to some type of a densest subgraph problem on a hypergraph. The condition of maximizing is related to the log density condition used by [BCC+10] in their work on the densest subgraph problem, since, assuming that the set of all vertices is not the best solution, the set satisfies that . However, we do not know how to use their algorithm to solve this problem. Beyond the fact that we consider the hypergraph setting, their algorithm manages to find a set of nontrivial density under the assumption that there is a “log dense” subset, but it is not guaranteed to find the “log dense” subset itself.
Let’s assume that we are given a distribution over unit vectors that achieve some value in (2.8). This is a non convex problem, and so generally the average of these vectors would not be a good solution. However, it turns out that the vector defined such that can sometimes be a good solution for this problem. Specifically, we will show that if it fails to give a solution of value at least , then we can find a new distribution obtained by reweighing elements that is in some sense “simpler” than . More precisely, we will define some nonnegative potential function such that for all and under the above conditions. This will show that we will need to use this reweighing step at most logarithmically many times.
where is the -dimensional vector defined by . Indeed, (2.10) follows from the non-negativity of and the Cauchy–Schwarz inequality since
Note that since is a distribution over unit vectors, both and are unit vectors, and hence (2.9) and (2.10) together with the fact that has bounded spectral norm imply that
However, it turns out that equals times the Hellinger distance of the two distributions over defined as follows: while (see Section 3). At this point we can use standard information theoretic inequalities to derive from (2.11) that there is mutual information between the two parts of . Another way to say this is that the entropy of the second part of drops on average by if we condition on the value of the first part. To say the same thing mathematically, if we define to be the distribution over and to be the distribution then
But one can verify that where is the distribution over ’s such that , which means that if we define then we get that
and hence is exactly the potential function we were looking for.
Approximation for nonnegative tensor maximization
Let be a degree- homogeneous polynomial in with nonnegative coefficients. Then, there is an algorithm, based on levels of the SOS hierarchy, that finds a unit vector such that
To prove Theorem 3.1 we first come up with a combining algorithm, namely an algorithm that takes (the moment matrix of) a distribution over unit vectors such that and find a unit vector such that . We then show that the algorithm will succeed even if is merely a level pseudo distribution; that is, the moment matrix is a pseudoexpectation operator. The combining algorithm is very simple:
Combining algorithm for polynomials with nonnegative coefficients:
Input: distribution over unit such that .
Operation: Do the following for steps:
For , let . If then output and quit.
Try to find such that the distribution satisfies , and set , where:
is defined by letting be proportional to for every .
is defined to be where is the Shannon entropy function and is the distribution over obtained by letting for every .
Clearly is always in , and hence if we can show that we always succeed in at least one of the steps, then eventually the algorithm will output a good . We now show that if the direct rounding step fails, then the conditioning step must succeed. We do the proof under the assumption that is an actual distribution. Almost of all of this analysis holds verbatim when is a pseudodistribution of level at least , and we note the one step where the extension requires using a nontrivial (though easy to prove) property of pseudoexpectations, namely that they satisfy the Cauchy–Schwarz inequality.
For any two jointly-distributed random variables and ,
1 Direct Rounding
Given , we define the following correlated random variables over : the probability that is equal to . Note that for every , the random variable is distributed according to . (Note that even if is only a pseudodistribution, are actual random variables.) The following lemma gives a sufficient condition for our direct rounding step to succeed:
Next, we bound the difference between and
Since both and are unit vectors, . By construction, the vector corresponds to the distribution and corresponds to the distribution . In particular, . Together with the bounds (3.1) and (3.2),
To verify this carries over when is a pseudodistribution, we just need to use the fact that Cauchy–Schwarz holds for pseudoexpectations (Lemma A.2).
2 Making Progress
The following lemma shows that if the sufficient condition above is violated, then on expectation we can always make progress. (Because are actual random variables, it automatically holds regardless of whether is an actual distribution or a pseudodistribution.)
If , then
The bound follows by combining a hybrid argument with Lemma 3.2.
Let be independent copies of so that
We consider the sequence of distributions with
By assumption, . Therefore, there exists an index such that . Let and . Then, and . By Lemma 3.2,
Since are independent of ,
By symmetry and the monotonicity of entropy under conditioning, we conclude
Lemma 3.4 implies that if our direct rounding fails then the expectation of conditioned on is at most , but in particular this means there exist so that . The probability of under this distribution is proportional to , which means that it exactly equals the distribution . Thus we see that . This concludes the proof of Theorem 3.1. ∎
If the polynomial is not homogenous but only has monomials of even degree, we can homogenize it by multiplying every monomial with an appropriate power of which is identically equal to on the sphere. To handle odd degree monomials we can introduce a new variable and set a constraint that it must be identically equal to . This way we can represent all odd degree monomials by even degree monomials with a blowup of in the coefficients. Note that if the pseudoexpectation operator is consistent with this constraint then our rounding algorithm will in fact output a vector that satisfies it.
Finding an “analytically sparse” vector in a subspace
In this section we prove Theorem 1.5. We let be a universe of size and be the vector space of real-valued functions . The measure on the set is the uniform probability distribution and hence we will use the inner product and norm for and .
There is a constant and a polynomial-time algorithm , based on levels of the SOS hierarchy, that on input a projector operator such that there exists a -sparse Boolean function satisfying , outputs a function such that
We will prove Theorem 4.1 by first showing a combining algorithm and then transforming it into a rounding algorithm. Note that the description of the combining algorithm is independent of the actual relaxation used, since it assumes a true distribution on the solutions, and so we first describe the algorithm before specifying the relaxation. In our actual relaxation we will use some auxiliary variables that will make the analysis of the algorithm simpler.
Combining algorithm for finding an analytically sparse vector:
Input: Distribution over Boolean (i.e., valued) functions that satisfy:
.
.
For , let be the function that satisfies for all . Go over all vectors of the form for and if there is one that satisfies (4.1) then output it. Note that the output of this procedure is independent of the distribution .
Choose a random gaussian vector and output if it satisfies (4.1). (Note that this is also independent of the distribution .)
Go over all choices for and modify the distribution to the distribution defined such that is proportional to for every .
For every one of these choices, let to be a random Gaussian that matches the first two moments of the distribution , and output if it satisfies (4.1).
Because we will make use of this fact later, we will note when certain properties hold not just for expectations of actual probability distributions but for pseudoexpectations as well. The extension to pseudoexpectations is typically not deep, but can be cumbersome, and so the reader might want to initially restrict attention to the case of a combining algorithm, where we only deal with actual expectations. We show the consequences for each of the steps failing, and then combine them together to get a contradiction.
We start by analyzing the random function rounding step. Let be an orthonormal basis for the space of functions . Let be a standard Gaussian function in , i.e., for independent standard normal variable (each with mean and variance ). The following lemmas combined show what are the consequences if is not much bigger than .
For any ,
In the basis, and . Then, and , which has expectation . Hence, the left-hand side is the same as the right-hand side. ∎
The moment of satisfies
By the previous lemma, the Gaussian variable has variance . Therefore,
since . ∎
The moment of satisfies
The random variable has a -distribution with degrees of freedom. The mean of this distribution is and the variance is . It follows that . ∎
2 Coordinate projection rounding
We now turn to showing the implications of the failure of projection rounding. We start by noting the following technical lemma, that holds for both the expectation and counting inner products:
Let and be two independent, vector-valued random variables. Then,
where the last equality holds by independence. But this is simply equal to
The following lemma shows a nontrivial consequence for being small:
For any distribution over ,
3 Gaussian Rounding
In this subsection we analyze the gaussian rounding step. Let be a random function with the Gaussian distribution that matches the first two moments of a distribution over .
The moment of satisfies
If have Gaussian distribution, then
The fourth moment of satisfies
4 Conditioning
We now show the sense in which conditioning can make progress. Let be a distribution over . For , let be the distribution reweighed by for . That is, , or in other words, for every function , . Similarly, we write for the distribution reweighed by .
For every even , there are points such that the reweighed distribution satisfies
but using and , the RHS is lower bounded by
Now, if was an actual expectation, then we could use Hölder’s inequality to lower bound the numerator of the RHS by which would lower bound the RHS by . For pseudoexpectations this follows by appealing to Lemma A.4. ∎
5 Truncating functions
The following observation would be useful for us for analyzing the case that the distribution is over functions that are not completely inside the subspace. Note that if the function is inside the subspace, we can just take in Lemma 4.11, and so the reader may want to skip this section in a first reading and just pretend that below.
Let , be a projector on and suppose that satisfies that and . Then there exists a function such that:
.
For every , .
Fix to be some sufficiently small constant (e.g., will do). Let . We define (i.e., if and otherwise) and define . Clearly for every .
Since if and only if , clearly and hence . Using , we see that . Now since is in the subspace, and hence for , . Therefore the probability that is at most . This means that with probability at least it holds that and , in which case . In particular, we get that . ∎
The proof of Lemma 4.11 establishes much more than its statement. In particular note that we did not make use of the fact that is nonnegative, and a function into with would work just the same. We also did not need the nonzero values to have magnitude exactly one, since the proof would easily extend to the case where they are in for some constant . One can also allow some nonzero values of the function to be outside that range, as long as their total contribution to the 2-norm squared is much smaller than .
6 Putting things together
We now show how the above analysis yields a combining algorithm, and we then discuss the changes needed to extend this argument to pseudodistributions, and hence obtain a rounding algorithm.
Let be a distribution over Boolean functions with and . The goal is to compute a function with , given the low-degree moments of .
Suppose that random-function rounding and coordinate-projection rounding fail to produce a function with . Then, (from failure of random-function rounding and Lemmas 4.3 and 4.4). By the failure of coordinate-projection rounding (and using Lemma 4.6 applied to the distribution over ) we get that
Since (by Lemma 4.11), for every and in the support of , we have for all in the support. Thus,
By the reweighing lemma, there exists such that the reweighted distribution satisfies
The failure of Gaussian rounding (applied to ) implies
By the properties of and Lemma 4.11, the left-hand side is and the right-hand side is . Therefore, we get
Finding planted sparse vectors
The goal here should be thought of as recovering to arbitrarily high precision (“exactly”), and thus the running time of an algorithm should be logarithmic in . We note that is not required to be random, and it may be chosen adversarially based on the choice of . We will prove the following theorem, which is this section’s main result:
(Theorem 1.4, restated) For some absolute constant , there is an algorithm that solves PlantedRecovery with high probability in time for any , where
Our algorithm will work in two stages. It will first solve a constant-degree sum-of-squares relaxation to find a somewhat noisy approximate solution. It will then solve an auxiliary linear program that converts any sufficiently good approximate solution into an exact one.
The first stage is based on the following theorem (proven in Section 5.1), which shows that we can approximately recover a vector when it is planted in a subspace consisting of vectors with substantially smaller ratio, provided that we can certify this property of the subspace using a low-degree sum-of-squares proof. To avoid unnecessary notation, we will use a degree 4 certificate in the statement and proof of the theorem; the proof goes through in greater generality, but this suffices for our application.
Furthermore, assume that (5.1) has a degree 4 sum-of-squares proof, i.e., that
where is the orthogonal projection onto , and is a degree 4 sum of squares.
There is a polynomial-time algorithm based on a constant-degree sum-of-squares relaxation that returns a vector with .
Since is -sparse, we know that , so we can take . We can thus solve a constant-degree sum-of-squares program to obtain a vector with whenever , i.e., when
For the second stage, we will consider the following linear program, which can be thought of as searching for a sparse vector in with a large inner product with :
In Section 5.2, we will prove the following theorem, which provides conditions under which the linear program will exactly recover from any that is reasonably correlated to it:
, [ is a -sparse vector] a
where for all ) [ doesn’t contain any -sparse vectors] a
[ is correlated with ] a
for all [ is not very correlated with anything in ]. a
then is the unique optimal solution to (5.4).
Because we believe the result might be useful elsewhere, we state the theorem in much more generality than needed for our application. In particular in our application we only need the trivial bound . Also, a bound on can also be derived using the relations between the norm and the norm on vectors in .
Fix , let , and let be a random -dimensional subspace given by the span of independent standard Gaussians. There exists a constant and absolute constants such that
for all with probability .
By Cauchy-Schwarz, we have , so Theorem 5.3 implies that we recover exactly as long as We note that the last two bounds were somewhat weak: Lemma 5.5 holds for subspaces of linear dimension, but we only applied it to a subspace with ; and the application of Cauchy-Schwarz could have been tightened using a better analysis. However, these were sufficient to prove Theorem 5.1.
is a constant for any fixed , so, by taking sufficiently small in the statement of the theorem, we may assume that , and thus that the right-hand side of (5.5) is at least 2. In this case, we can recover as long as . Combining this with Equation (5.3) and choosing appropriately thus completes the proof of Theorem 5.1. ∎
It thus suffices to prove Theorems 5.2 and 5.3, which we will do in sections 5.1 and 5.2, respectively. We note that Theorems 5.2 and 5.3 hold for any that meets certain norm requirements, and they do not require to be a uniformly random subspace. As such, the results of this section hold in a broader context. (For example, they immediately generalize to other distributions of subspaces that meet the norm bounds.) We hope that the technical results of this section will find other uses, so we have stated them in a somewhat general way to facilitate their application in other settings.
In this section, we prove Theorem 5.2, which allows us to recover a vector that is reasonably well-correlated with . The basic idea is that has a much larger ratio than anything in , so maximizing the the ratio should give a vector near .
The key ingredient of the theorem is the following lemma about (pseudo-)distributions supported on -sparse functions in . Note that this lemma does not need the space to be random, but only that it can be certified to have no sparse vectors by the SOS SDP.
Let be a linear subspace such that
We can obtain a pseudodistribution meeting the requirements in Lemma 5.6 by solving a degree 8 sum-of-squares program that maximizes over with . If we sample a random Gaussian consistent with the first two moments of , then we will obtain a vector whose expected -norm squared is and whose expected inner product with is , so Lemma 5.6 therefore implies Theorem 5.2.
Write every vector in the support of in the form where and . We know that
This concludes the proof for actual expectations. To argue about pseudoexpectations, we need to use only constraints involving polynomials, and therefore we use
The existence of a degree 4 sum-of-squares proof of (5.7) implies that must be consistent with the constraint . We can thus use Cauchy–Schwarz and Hölder’s inequality (Lemma A.10 and Corollary A.11), to bound all of the terms except the first one by a constant times , and so we get
Using the fact that the expectation is consistent with the constraint , we obtain
In this section, we prove Theorem 5.3, which allows us to use a vector near to recover exactly (up to the precision used when solving the linear program). Intuitively, this is relies on the same tendency towards sparsity of vectors with minimal -norm that underlies the earlier works that are based on -sparsity. Minimizing the -sparsity amounts to solving the linear program in (5.4) with equal to each of the unit basis vectors, and then taking the best of the solutions. When is sparse enough, it will have at least one fairly large coefficient, and will then be sufficiently correlated with the corresponding unit basis vector for the linear program to find it. This breaks down when , at which point any one basis vector is expected to be more correlated with some vector in than it is with . Here, instead of using the unit basis vectors, we use a vector that shares many coordinates with , which then lets us handle a much broader range of .
To analyze the optimum of (5.4), we decompose as for and . We will show that
for all , with equality only if , which immediately implies Theorem 5.3.
Let and be the vectors obtained from by zeroing out the coordinates outside and , respectively, so that Since is zero outside of , we have
where the second inequality in (5.10) is strict unless the two terms inside the are equal. To prove the inequality asserted in (5.8), and thus Theorem 5.3, it therefore suffices to show that
for all
We can bound the left-hand side of (5.11) using the assumptions that and that is -sparse:
To bound the numerator of the right-hand side of (5.11), we need to show that cannot have too large a fraction of its 1-norm concentrated in the coordinates in . We first note that, if this occurred, it would lead to a large contribution to the -norm:
Combining this with our assumption that , gives
If this implies that
Results for Small Set Expansion
As stated in Corollaries 1.3 and 1.6, our results imply two consequences for the Small Set Expansion problem of [RS10]. This is the problem of deciding, given an input graph and parameters , whether there is a measure- subset of ’s vertices where all but an fraction of ’s edges stay inside it, or that is a small set expander in the sense that every sufficiently small set has almost all its edges leaving it. Beyond being a natural problem in its own right, Small Set Expansion is also closely related to the Unique Games problem whose conjectured hardness is known as Khot’s “Unique Games Conjecture” [Kho02]. [RS10] gave a reduction from Small Set Expansion to Unique Games. While a reduction in the other direction is not known, all currently known algorithmic and integrality gap results apply to both problems equally well (e.g., [ABS10, RST10, BGH+12, BBH+12]), and thus they are likely to be computationally equivalent.
We give an algorithm to solve Small Set Expansion in quasipolynomial time on an interesting family of Cayley graphs, and a new polynomial-time approximation algorithm for this problem on general graphs, with the approximation guarantee depending on the dimension of the input graph’s top eigenspace.
In this section, we describe approximation algorithms with running times that depend on . The algorithms run in quasipolynomial time if is polylogarithmic. We will show interesting families of graphs with . (See Theorem 6.3.)
The following theorem shows that low-degree sum-of-squares relaxations can detect -sparse functions in the subspaces (in the case when is not too large). This result follows from Theorem 3.1 and the fact that the polynomial has nonnegative coefficients in an appropriate basis.
Sum-of-squares relaxations of degree provide an additive -approximation to the maximum of over all non-zero functions .
Using the characterization of small-set expansion in terms of -sparse functions [BBH+12], Theorem 6.1 implies the following approximation algorithm for small-set expansion on Cayley graphs. This theorem implies Corollary 1.3.
For some absolute constant and all small enough, sum-of-squares relaxations of degree can distinguish between the following two cases with .
The Cayley graph contains a vertex set of measure at most and expansion at most .
All vertex sets of measure at most in have expansion at least .
We will show that the maximum of over distinguishes the two cases (by a constant margin). Therefore, Theorem 6.1 implies that we can distinguish between the cases using sum-of-squares relaxations.
Yes-case: Let be the indicator function of a set with measure at most and expansion at most . Then, . It follows that . (See Lemma 4.11.) Therefore, .
No-case: Let . By [BBH+12, Theorem 2.4], graphs with this kind of small-set expansion satisfy for all functions . ∎
The following theorem shows that there are interesting Cayley graphs that satisfy for . We consider constructions based on the long code and the short code [BGH+12]. These constructions are parameterized by the size of the graph and its eigenvalue gap. In the context of the Unique Games Conjecture and the Small-Set Expansion Hypothesis, the most relevant case is that the eigenvalue gap is a constant. (The eigenvalue gap corresponds to the gap to perfect completeness.)
Long-code and short-code based graphs with constant eigenvalue gap satisfy for all .
2 Approximating small-set expansion using ASVP
The approximation algorithm for the analytical sparse vector problem (Theorem 4.1) implies the following approximation algorithm for small-set expansion. An algorithm for the same problem with the factor replaced by a constant would refute the Small-Set Expansion Hypothesis [RS10, RST12]. It’s plausible that, under standard complexity assumptions such as , even a smaller improvement to a factor instead of would refute this hypothesis, though we have no proof of such an implication.
For some absolute constant and all small enough, sum-of-squares relaxations with constant degree can solve the promise problem on regular graphs :
The graph contains a vertex with measure at most and expansion at most , where .
All vertex sets of measure at most have expansion at least .
Suppose satisfies the Yes property. Let be the indicator functions of a set with measure at most and expansion at most . Then, , where we can make as large as we like by making larger. By Theorem 4.1, constant-degree sum-of-squares relaxations allow us to find an -sparse function , so that . By [BBH+12, Theorem 2.4] (see also Appendix D), such a function certifies that we are not in the No case. ∎
Discussion and open questions
A general open question is to find other applications of our approach for rounding sum-of-squares relaxations. Natural candidates would be problems where it seems that they do not display a “dichotomy” behavior, where beating some simple algorithm is likely to be exponentially hard, but rather suggest more of a smooth tradeoff between time and performance. As far as we are aware, all known “robust”For knapsack-like problems, there exist explicit lower bounds [Gri01a], but here low-degree sum-of-squares proofs provide very good approximation (in this sense, the lower bound is not robust). lower-bound results for the sum-of-squares method are non-constructive, i.e., they show that hard instances for the sos method exist but do not give an efficient way of constructing them. More concretely, the results use the probabilistic method and show that with high probability, random instances are hard for sum-of-squares relaxations [Gri01b, Sch08]. Therefore, SOS seems promising for problems where random instances do not seem to be the most difficult, e.g., problems related to the Unique Games Conjecture. A concrete problem of that type to look at is Sparsest Cut. In particular, can we obtain even a small improvementAn approach to obtain constant-factor approximations for sparsest cut in subexponential time is outlined in the dissertation [Ste10a, Chapter 9]. However, this approach also works with weaker hierarchies. An approach tailored to sum-of-squares would be interesting. to [ARV04]’s algorithm using more SOS levels? In fact, we believe that even finding a natural reinterpretation of [ARV04] result in our framework would be interesting. That said, our result for finding a planted sparse vector shows that SOS can be useful for average-case problems as well, and in particular we believe SOS might be a strong tool for solving unsupervised learning problems, especially for nonlinear models.
A relaxation-based approximation algorithm can be thought of as having three components: the relaxation, the rounding algorithm, and its analysis. In our approach there is almost no creativity in choosing the relaxation, which is simply taken to be a sufficiently high level of the SOS hierarchy. (Though there may be some flexibility in how we represent solutions.) Can we similarly show a “universal” rounding algorithm, thus pushing all the creative choices into the analysis? A related question is whether one can formulate a theorem giving a translation from combining algorithms into rounding algorithms under sufficiently general conditions, so that results like ours would follow as special cases, and as mentioned in Section 1.4, have already made some progress in this direction.
In the context of the Small-Set-Expansion Hypothesis / Unique Games Conjecture, the most important question is whether our results of Section 4 can be further improved. We do not know of any candidate hard instances for this problem (in the relevant range of parameters) and so conjecture that our algorithm (or at least our analysis of it) is not optimal and can be improved further.
Related to the question of finding hard instances, our work suggests a different type of negative results for convex relaxations. While integrality gaps are instances that are hard for a particular relaxation, regardless of the rounding algorithm, one can consider the notion of “combining gaps”. These will be instances where there is a distribution of good solutions, but a particular combining algorithm fails to find one. Hence, viewing as a rounding algorithm, such a result shows that will fail regardless of the relaxation used. (Karloff’s work [Kar99] on hard instances for the [GW95] hyperplane cut rounding algorithms can be viewed as such an example.) Studying such gaps can shed more light on our approach and computational difficulty in general. In particular, it might be interesting to consider this question for random satisfiable instances of SAT or other constraint satisfaction problems.
References
Appendix A Pseudoexpectation toolkit
We recall here the definition of pseudoexpectation from [BBH+12] and prove some of its useful properties. Some of these were already proven in [BBH+12] but others are new.
We now record various useful ways in which pseudoexpectations behave close to actual expectations.
For two polynomials and , we write if for some polynomials .
One of the most useful properties of pseudo-expectation is that it satisfies the Cauchy–Schwarz inequality:
In particular this implies the following corollary
In this paper we also need the following variant of Hölder’s inequality:
which using our induction hypothesis on vs , we can lower bound by
We sometime would need to extend a pseudoexpectation of one random variable to a pseudoexpectation of two independent copies of it. The following lemma would be useful there
Write where are monomials, then and so under our definition
We would like to understand how polynomials behave on linear subspaces of . A map is polynomial over a linear subspace if restricted to agrees with a polynomial in the coefficients for some basis of . Concretely, if is an (orthonormal) basis of , then is polynomial over if agrees with a polynomial in . We say that holds over a subspace if , as a polynomial over , is a sum of squares.
The relation holds if and only if . Furthermore, if and , then .
If , then implies . (Multiplying both sides with a sum of squares preserves the order.) On the other hand, suppose . Since , we also have . Since , the relation also implies .
For the second part of the lemma, suppose and . Using the first part of the lemma, we have . It follows that , which in turn implies (using the other direction of the first part of the lemma). ∎
The right-hand side minus the LHS equals the square polynomial ∎
Here is another form of the Cauchy–Schwarz inequality.
If is a level- p.d. over , then
And it implies another form of Hölder’s inequality
If is a level p.d. over , then
Here we note the following alternative characterization of the spectral norm of a polynomial:
On the other hand, suppose that where the ’s are quadratic polynomials. We can let be such that , and then let be the quadratic operator on such that for every . One can easily verify that the spectral norm of is at most and for every . ∎
Appendix B Low-Rank Tensor Optimization
Consider an -variate degree- polynomial of the form for quadratic polynomials .
There exists an algorithm that, given and , computes up to multiplicative error in time .
For , consider the polynomial .
First, we claim that . On the one hand, by Cauchy–Schwarz. On the other hand, if for some vector , then . Therefore, if we choose as a unit vector that maximizes , then .
Next, we claim that we can compute up to error in time . Since is quadratic, we can compute in polynomial time. (The norm of is equal to the largest singular value of the coefficient matrix of .) The idea is to compute for all vectors , where is an -net of the unit ball in . Let be the vector that achieves the maximum, be the corresponding input, and be the vector . Thus and . Therefore, for every
But if then
Thus if then we get a multiplicative approximation to . ∎
If is a symmetric PSD matrix with Frobenius norm at most then we can compute an additive approximation to
in time.
Write in its eigenbasis as for matrices with Frobenius norm at most , and let . Since we know that the rank of is at most , and therefore we can compute a multiplicative approximation to the maximum of over unit (which in particular implies an additive approximation since this value is bounded by ). But this implies an -additive approximation for this maximum over since these quantities can differ by at most . ∎
Appendix C LOCC Polynomial Optimization
Let be a degree- homogeneous polynomial of the form for quadratic polynomials with and quadratic polynomials with and . Note that this corresponds to the tensor corresponding to having a one-way local operations and classical communication (LOCC) norm bounded by [BCY11]. Without loss of generality, we may assume . (We can choose and choose appropriately without changing .) Our goal is to compute the norm of , defined as . In the quantum setting, this corresponds to finding the maximum probability of acceptance by a separable state for the measurement operator .
In this section, we will show that sum-of-squares relaxations provide good approximation for the norm of polynomials of the form above. (For the case that the variables of are disjoint from the variables for , the theorem is due to Brandão, Christandl, and Yard [BCY11]. Up to the Gaussian rounding step, the proof here is essentially the same as the proof by Brandão and Harrow [BH13].)
Sum-of-squares relaxations with degree achieve the following approximations for the norm of degree- polynomials of the form above:
the value of the relaxation is at most for degree .
in the case that the variables in are disjoint from the variables in , the value of the relaxation is at most for degree .
As direct rounding for a distribution , we choose a Gaussian variable with the same first two moments as . To analyze this rounding procedure, the following lemma is useful.
The lemma considers an arbitrary distribution over unit vectors in (intended to maximize ). We express the second moment of this distribution as convex combination for and . By the assumptions on the distribution , the matrices and are positive semidefinite and have trace (density matrices). The quantum entropy is concave so that . The assumption of the lemma is that the inequality is approximately tight. Roughly speaking, this condition means that the matrices are close . For the distribution , this condition means that reweighing by the polynomials does not affect second moments of the distribution. We say that the distribution has low global correlation with respect to the polynomials . (This notion is related to but distinct from the notion of global correlation in [BRS11]). The lemma asserts that if the distribution has low global correlation with respect to the polynomials , then sampling independently for the -part and -part of the polynomial gives roughly the same value as sampling in a correlated way. (The next lemma explains why our direct rounding achieves at least the quantity corresponding to sampling independently for the two parts.)
Let be a distribution over that satisfies the constraint . Suppose for and . Then,
Moreover, the statement holds if is a degree- pseudo-distribution.
Consider the block-diagonal density matrix and the block-diagonal measurement matrix (In this construction, we identify the quadratic polynomial with its representation as a symmetric square matrix.) Furthermore, consider the partial traces and . We can express the two sides of the conclusion of the lemma as follows,
Since has spectral norm at most , we can bound the difference by. (Here, is the trace norm—the dual of the spectral norm.) By Pinsker’s inequality, . By the chain rule, . The assumption of the lemma allows us to bound the trace norm by . At this point, the conclusion of the lemma follows from the bound . ∎
The following lemma shows that Gaussian rounding achieves a value at least as large as the value achieved by sampling independently for the -part and -part of the polynomial .
Let be a distribution over that satisfies the constraint . Suppose is a Gaussian distribution with the same first two moments as . Then, satisfies , , and
Moreover, the statement holds if is a degree- pseudo-distribution.
Using the assumption , the lemma follows from the fact that Gaussian variables satisfy . ∎
The previous two lemmas together yield the following corollary.
Let be a distribution over that satisfies the constraint . Suppose and are as in the previous two lemmas, that is, is a Gaussian distribution with the same first two moments as and for and . Then,
Moreover, the statement holds if is a degree- pseudo-distribution.
Making progress
The following lemma shows that there exists a low-degree polynomial so that reweighing by the polynomial results in a distribution that has low global correlation with respect to the polynomials .
Let be a distribution over that satisfies the constraint . Then, there exists a polynomial of the form with such that for and . Moreover, the statement holds if is a degree- pseudo-distribution.
By contraposition, suppose that holds for all polynomials of the form with . Then, we can greedily construct a sequence of polynomial such that in each step the entropy decreases by at least . In particular, for and . Since and , we have . As desired it follows that there exists a polynomial of the desired form such that . ∎
Putting things together
The following lemma combines the conclusion about direct-rounding and making-progress.
Let be a distribution over that satisfies the constraints and . Then, there exists a polynomial of the form with such that the Gaussian distribution that matches the first two moments of reweighted by satisfies
(Concretely, is the Gaussian distribution that satisfies for quadratic polynomial ). Moreover, the statement holds for degree- pseudo-distributions.
Take the polynomial as in Lemma C.5. Reweigh the distribution by the polynomial . Apply Corollary C.4 to the resulting distribution. ∎
At this time, we have all ingredients for the proof of Theorem C.1.
Let be a degree- pseudo-distribution over that satisfies the constraints and for . By the previous lemma, there exists a distribution over such that . It follows that there exists a vector with . (We can also find such a vector efficiently because we can sample from the distribution efficiently and the random variables and are well-behaved.) By homogeneity, we get .
In the case that the variables in are disjoint from the variables in , we can modify the direct-rounding distribution slightly and sample the variables for the polynomials independently from the variables for the polynomials. By Lemma C.2, we still have . We can assume that (by adding the corresponding constraint to the sos relaxation). Therefore, . It follows that there exists a vector in with . By homogeneity, we can assume that and . In this case, and as desired. ∎
Appendix D The 2-to-q norm and small-set expansion
This appendix reproduces from [BBH+12] the proof that a graph is a small-set expander if and only if the projector to the subspace of its adjacency matrix’s top eigenvalues has a bounded norm for even . We also note that while [BBH+12] stated their result for the decision question, it does yield an efficient algorithm to transform a vector in the top eigenspace with large norm into a small set that does not expand.
Our main theorem of this section is the following:
For every regular graph , and even ,
(Norm bound implies expansion) For all , implies that .
(Expansion implies norm bound) There is a constant such that for all , implies . Moreover there is an efficient algorithm such that given a function such that finds a set of measure less than such that .
If there is a polynomial-time computable relaxation yielding good approximation for the , then the Small-Set Expansion Hypothesis of [RS10] is false.
Using [RST12], to refute the small-set expansion hypothesis it is enough to come up with an efficient algorithm that given an input graph and sufficiently small , can distinguish between the Yes case: and the No case for any and some constant . In particular for all and constant , if is small enough then in the No case . Using Theorem D.1, in the Yes case we know , while in the No case, if we choose to be smaller then in the Theorem, then we know that . Clearly, if we have a good approximation for the norm then, for sufficiently small we can distinguish between these two cases. ∎
The first (easier) part of Theorem D.1 is proven in Section D.1. The second part will follow from the following lemma:
Set , with a constant . Then for every and , if is a graph that satisfies
for all with , then for all . Moreover, there is an efficient algorithm that given a function such that finds a set that violates (D.1).
Proving the second part of Theorem D.1 from Lemma D.3
We use the variant of the local Cheeger bound obtained in [Ste10b, Theorem 2.1], stating that if then for every satisfying , . The proof follows by noting that for every set , if is the characteristic function of then , and . Because this local Cheeger bound is algorithmic (and transforms a function with large ratio into a set by simply using a threshold cut), this part is algorithmic as well. ∎
Fix . We assume that the graph satisfies the condition of the Lemma with , for a constant that we’ll set later. Let be such a graph, and be function in with that maximizes . We write where denote the eigenfunctions of with values that are at least . Assume towards a contradiction that . We’ll prove that satisfies . This is a contradiction since (using ) , and we assumed is a function in with a maximal ratio of . (To prove the “moreover” part, where we don’t assume is the maximal function, we repeat this process with until we get stuck.)
Let be the set of vertices such that for all . Using Markov and the fact that , we know that , meaning that under our assumptions any subset satisfies . On the other hand, because , we know that contributes at least half of the term . That is, if we define to be then . We’ll prove the lemma by showing that .
Let be a sufficiently large constant ( will do). We define to be the set , and let be the maximal such that is non-empty. Thus, the sets form a partition of (where some of these sets may be empty). We let be the contribution of to . That is, , where . Note that . We’ll show that there are some indices such that:
.
For all , there is a nonnegative function such that .
For every , .
Showing these will complete the proof, since it is easy to see that for two nonnegative functions and even , , , and hence (ii) and (iii) imply that
Using (i) we conclude that for , the right-hand side of (D.2) will be larger than .
We find the indices iteratively. We let be initially the set of all indices. For we do the following as long as is not empty:
Let be the largest index in .
Remove from every index such that .
We let denote the step when we stop. Note that our indices are sorted in descending order. For every step , the total of the ’s for all indices we removed is less than and hence we satisfy (i). The crux of our argument will be to show (ii) and (iii). They will follow from the following claim:
Let and be such that and for all . Then there is a set of size at least such that .
The claim will follow from the following lemma:
Let be a distribution with and be some function. Then there is a set of size such that .
Identify the support of with the set for some , we let denote the probability that outputs , and sort the ’s such that . We let denote ; that is, . We separate to two cases. If , we define the distribution as follows: we set to be for , and we let all be equiprobable (that is be output with probability ). Clearly, , but on the other hand, since the maximum probability of any element in is at most , it can be expressed as a convex combination of flat distributions over sets of size , implying that one of these sets satisfies , and hence .
The other case is that . In this case we use Cauchy–Schwarz and argue that
But using our bound on the collision probability, the right-hand side of (D.3) is upper bounded by . ∎
By construction , and hence we know that for every , . This means that if we let be the distribution then
By the expansion property of , and thus by Lemma D.5 there is a set of size satisfying . ∎
We will construct the functions by applying iteratively Claim D.4. We do the following for :
Let be the set of size that is obtained by applying Claim D.4 to the function and the set . Note that , where we let (and hence for every , ).
Let be the function on input that outputs if and otherwise, where is a scaling factor that ensures that equals exactly .
We define .
Note that the second step ensures that , while the third step ensures that for all , and in particular . Hence the only thing left to prove is the following:
Recall that for every , , and hence (using for ):
Now fix . Since is at least (in fact equal) and , we can use (D.4) and , to reduce proving the claim to showing the following:
We know that . We claim that (D.5) will follow by showing that for every ,
where . (Note that since in our construction the indices are sorted in descending order.)
Indeed, (D.6) means that if we let momentarily denote then
The first inequality holds because we can write as , where . Then, on the one hand, , and on the other hand, since . The second inequality holds because . By squaring (D.7) and plugging in the value of we get (D.5).
Proof of (D.6)
since otherwise the index would have been removed from the at the step. Since , we can plug (D.4) in (D.8) to get
Since for all , it follows that . On the other hand, we know that . Thus,
and now we just choose sufficiently large so that . ∎ ∎
D.1 Norm bound implies small-set expansion
In this section, we show that an upper bound on norm of the projector to the top eigenspace of a graph implies that the graph is a small-set expander. This proof appeared elsewhere implicitly [KV05, O’D07] or explicitly [BGH+12, BBH+12] and is presented here only for completeness. Fix a graph (identified with its normalized adjacency matrix), and , letting denote the subspace spanned by eigenfunctions with eigenvalue at least .
If satisfy then . Indeed, by Hölder’s inequality, and by choosing and normalizing one can see this equality is tight. In particular, for every , and . As a consequence
Note that if is a projection operator, . Thus, part 1 of Theorem D.1 follows from the following lemma:
Let be regular graph and . Then, for every ,
Let be the characteristic function of , and write where and is the projection to the eigenvectors with value less than . Let . We know that
And , meaning that . We now write
Plugging this into (D.9) yields the result. ∎