Sparse Covers for Sums of Indicators
Constantinos Daskalakis, Christos Papadimitriou
Introduction
A Poisson Binomial Distribution of order is the discrete probability distribution of the sum of independent indicator random variables. The distribution is parameterized by a vector of probabilities, and is denoted . In this paper we establish that the set of all Poisson Binomial distributions of order admits certain useful covers with respect to the total variation distance between distributions. Namely
For all , there exists a set such that:
is an -cover of in total variation distance; that is, for all , there exists some such that
can be computed in time .
Moreover, all distributions in the cover satisfy at least one of the following properties, for some positive integer
Covers such as the one provided by Theorem 1 are of interest in the design of algorithms, when one is searching a class of distributions to identify an element of the class with some quantitative property, or in optimizing over a class with respect to some objective. If the metric used in the construction of the cover is relevant for the problem at hand, and the cover is discrete, relatively small and easy to construct, then one can provide a useful approximation to the sought distribution by searching the cover, instead of searching all of . For example, it is shown in [DP07, DP09, DP13] that Theorem 1 implies efficient algorithms for computing approximate Nash equilibria in an important class of multiplayer games, called anonymous [Mil96, Blo99].
We proceed with a fairly detailed sketch of the proof of our main cover theorem, Theorem 1, stating two additional results, Theorems 2 and 3. The complete proofs of Theorems 1, 2 and 3 are deferred to Sections 3, 4 and 5 respectively. Section 1.4 discusses related work, while Section 2 provides formal definitions, as well as known approximations to the Poisson Binomial distribution by simpler distributions, which are used in the proof.
At a high level, the proof of Theorem 1 is obtained in two steps. First, we establish the existence of an -cover whose size is polynomial in and , via Theorem 2. We then show that this cover can be pruned to size polynomial in and using Theorem 3, which provides a quantification of how the total variation distance between Poisson Binomial distributions depends on the number of their first moments that are equal.
We proceed to state the two ingredients of the proof, Theorems 2 and 3. We start with Theorem 2 whose detailed sketch is given in Section 1.2, and complete proof in Section 4.
Theorem 2 implies the existence of an -cover of whose size is . This cover can be obtained by enumerating over all Poisson Binomial distributions of order that are in -sparse or -Binomial form as defined in the statement of the theorem, for .
The next step is to sparsify this cover by removing elements to obtain Theorem 1. Note that the term in the size of the cover is due to the enumeration over distributions in sparse form. Using Theorem 3 below, we argue that there is a lot of redundancy in those distributions, and that it suffices to only include of them in the cover. In particular, Theorem 3 establishes that, if two Poisson Binomial distributions have their first moments equal, then their distance is at most . So we only need to include at most one sparse form distribution with the same first moments in our cover. We proceed to state Theorem 3, postponing its proof to Section 5. In Section 1.3 we provide a sketch of the proof.
Condition in the statement of Theorem 3 constrains the first power sums of the expectations of the constituent indicators of two Poisson Binomial distributions. To relate these power sums to the moments of these distributions we can use the theory of symmetric polynomials to arrive at the following equivalent condition to :
We provide a proof that in Proposition 2 of Section 6.
In view of Remark 1, Theorem 3 says the following:
“If two sums of independent indicators with expectations in [0,1/2] have equal first moments, then their total variation distance is .”
We note that the bound (1) does not depend on the number of variables , and in particular does not rely on summing a large number of variables. We also note that, since we impose no constraint on the expectations of the indicators, we also impose no constraint on the variance of the resulting Poisson Binomial distributions. Hence we cannot use Berry-Esséen type bounds to bound the total variation distance of the two Poisson Binomial distributions by approximating them with Normal distributions. Finally, it is easy to see that Theorem 3 holds if we replace with . See Corollary 1 in Section 6.
In Section 3 we show how to use Theorems 2 and 3 to obtain Theorem 1. We continue with the outlines of the proofs of Theorems 2 and 3, postponing their complete proofs to Sections 4 and 5.
2 Outline of Proof of Theorem 2
Given arbitrary indicators we obtain indicators , satisfying the requirements of Theorem 2, in two steps. We first massage the given variables to obtain variables such that
that is, we eliminate from our collection variables that have expectations very close to or , without traveling too much distance from the starting Poisson Binomial distribution.
Variables do not necessarily satisfy Properties 2a or 2b in the statement of Theorem 2, but allow us to define variables which do satisfy one of these properties and, moreover,
(2), (3) and the triangle inequality imply , concluding the proof of Theorem 2.
Stage 2: The definition of depends on the number of ’s which are not or . The case corresponds to Case 2a in the statement of Theorem 2, while the case corresponds to Case 2b.
Case : First, we set , if . We then argue that each , , can be rounded to some , which is an integer multiple of , so that (3) holds. Notice that, if we were allowed to use multiples of , this would be immediate via an application of Lemma 2:
We improve the required accuracy to via a series of Binomial approximations to the Poisson Binomial distribution, using Ehm’s bound [Ehm91] stated as Theorem 5 in Section 2.1. The details involve partitioning the interval into irregularly sized subintervals, whose endpoints are integer multiples of . We then round all but one of the ’s falling in each subinterval to the endpoints of the subinterval so as to maintain their total expectation, and apply Ehm’s approximation to argue that the distribution of their sum is not affected by more than in total variation distance. It is crucial that the total number of subintervals is to get a total hit of at most in variation distance in the overall distribution. The details are given in Section 4.2.1.
Case : We approximate with a Translated Poisson distribution (defined formally in Section 2), using Theorem 6 of Section 2.1 due to Röllin [R0̈7]. The quality of the approximation is inverse proportional to the standard deviation of , which is at least , by the assumption . Hence, we show that is -close to a Translated Poisson distribution. We then argue that the latter is -close to a Binomial distribution , where and is an integer multiple of . In particular, we show that an appropriate choice of and implies (3), if we set of the ’s equal to and the remaining equal to . The details are in Section 4.2.2.
3 Outline of Proof of Theorem 3
4 Related Work
It is believed that Poisson [Poi37] was the first to study the Poisson Binomial distribution, hence its name. Sometimes the distribution is also referred to as “Poisson’s Binomial Distribution.” PBDs have many uses in research areas such as survey sampling, case-control studies, and survival analysis; see e.g. [CL97] for a survey of their uses. They are also very important in the design of randomized algorithms [MR95].
In Probability and Statistics there is a broad literature studying various properties of these distributions; see [Wan93] for an introduction to some of this work. Many results provide approximations to the Poisson Binomial distribution via simpler distributions. In a well-known result, Le Cam [LC60] shows that, for any vector ,
where is the Poisson distribution with parameter . Subsequently many other proofs of this bound and improved ones, such as Theorem 4 of Section 2.1, were given, using a range of different techniques; [HC60, Che74, BH84, DP86] is a sampling of work along these lines, and Steele [Ste94] gives an extensive list of relevant references. Much work has also been done on approximating PBDs by Normal distributions (see e.g. [Ber41, Ess42, Mik93, Vol95, CGS10]) and by Binomial distributions; see e.g. Ehm’s result [Ehm91], given as Theorem 5 of Section 2.1, as well as Soon’s result [Soo96] and Roos’s result [Roo00], given as Theorem 7 of Section 2.1.
These results provide structural information about PBDs that can be well approximated by simpler distributions, but fall short of our goal of approximating a PBD to within arbitrary accuracy. Indeed, the approximations obtained in the probability literature (such as the Poisson, Normal and Binomial approximations) typically depend on the first few moments of the PBD being approximated, while higher moments are crucial for arbitrary approximation [Roo00]. At the same time, algorithmic applications often require that the approximating distribution is of the same kind as the distribution that is being approximated. E.g., in the anonymous game application mentioned earlier, the parameters of the given PBD correspond to mixed strategies of players at Nash equilibrium, and the parameters of the approximating PBD correspond to mixed strategies at approximate Nash equilibrium. Approximating the given PBD via a Poisson or a Normal distribution would not have any meaning in the context of a game.
As outlined above, the proof of our main result, Theorem 1, builds on Theorems 2 and 3. A weaker form of these theorems was announced in [Das08, DP09], while a weaker form of Theorem 1 was announced in [DDS12].
Preliminaries
Covers: Let be a set of probability distributions. A subset is called a (proper) -cover of in total variation distance if, for all , there exists some such that .
Let be mutually independent indicators with expectations respectively. Similarly let be mutually independent indicators with expectations respectively. The distributions of and are different if and only if .
Translated Poisson Distribution: We say that an integer random variable has a translated Poisson distribution with parameters and and write iff
where represents the fractional part of .
Similarly, we write if and only if there exist positive reals and such that
We are casual in our use of the order notation and throughout the paper. Whenever we write or in some bound where ranges over the integers, we mean that there exists a constant such that the bound holds true for sufficiently large if we replace the or in the bound by . On the other hand, whenever we write or in some bound where ranges over the positive reals, we mean that there exists a constant such that the bound holds true for sufficiently small if we replace the or in the bound with .
We conclude with an easy but useful lemma whose proof we defer to Section 6.
Let be mutually independent random variables, and let be mutually independent random variables. Then
We present a collection of known approximations to the Poisson Binomial distribution via simpler distributions. The quality of these approximations can be quantified in terms of the first few moments of the Poisson Binomial distribution that is being approximated. We will make use of these bounds to approximate Poisson Binomial distributions in different regimes of their moments. Theorems 4—6 are obtained via the Stein-Chen method.
where is the Binomial distribution with parameters and .
where and .
The approximation theorems stated above do not always provide tight enough approximations. When these fail, we employ the following theorem of Roos [Roo00], which provides an expansion of the Poisson Binomial distribution as a weighted sum of a finite number of signed measures: the Binomial distribution (for an arbitrary value of ) and its first derivatives with respect to the parameter , at the chosen value of . For the purposes of the following statement we denote by the probability assigned by the Binomial distribution to integer .
Let , be mutually independent indicators with expectations , and . Then, for all and ,
where for the purposes of the above expression:
where for the last definition we interpret as a function of .
If then, for all :
Proof of Theorem 1
We next show that we can remove from a large number of the sparse-form distributions it contains to obtain a -cover of . In particular, we shall only keep sparse-form distributions by appealing to Theorem 3. To explain the pruning we introduce some notation. For a collection of probability values we denote by and by . Theorem 3, Corollary 1, Lemma 2 and Lemma 1 imply that if two collections and of probability values satisfy
then . In particular, for some , this bound becomes at most .
For a collection , we define its moment profile to be the -dimensional vector
By the previous discussion, for two collections , if then .
Given the above we sparsify as follows: for every possible moment profile that can arise from a Poisson Binomial distribution in -sparse form, we keep in our cover a single Poisson Binomial distribution with such moment profile. The cover resulting from this sparsification is a -cover, since the sparsification loses us an additional in total variation distance, as argued above.
We now bound the cardinality of the sparsified cover. The total number of moment profiles of -sparse Poisson Binomial distributions is . Indeed, consider a Poisson Binomial distribution in -sparse form. There are at most choices for , at most choices for , and at most choices for . We also claim that the total number of possible vectors
is . Indeed, if there is just one such vector, namely the all-zero vector. If , then, for all , and it must be an integer multiple of . So the total number of possible values of is at most , and the total number of possible vectors
The same upper bound applies to the total number of possible vectors
The moment profiles we enumerated over are a superset of the moment profiles of -sparse Poisson Binomial distributions. We call them compatible moment profiles. We argued that there are at most compatible moment profiles, so the total number of Poisson Binomial distributions in -sparse form that we keep in the cover is at most . The number of Poisson Binomial distributions in -Binomial form is the same as before, i.e. at most , as we did not eliminate any of them. So the size of the sparsified cover is .
Proof of Theorem 2
We organize the proof according to the structure and notation of our outline in Section 1.2. In particular, we proceed to provide the details of Stages 1 and 2, described in the outline. The reader should refer to Section 1.2 for notation.
Define and We define the expectations of the intermediate indicators as follows.
First, we set , for all . It follows that
Next, we define the probabilities , , using the following procedure:
Set ; and let be an arbitrary subset of cardinality .
Set , for all , and , for all .
We bound the total variation distance using the Poisson approximation to the Poisson Binomial distribution. In particular, Theorem 4 implies
Similarly, Finally, we use Lemma 3 (given below and proved in Section 6) to bound the distance
where we used that . Using the triangle inequality the above imply
We follow a similar rounding scheme to define from . That is, we round some of the ’s to and some of them to so that . As a result, we get (to see this, repeat the argument employed above to the variables and , )
Using (5), (6), (7) and Lemma 2 we get (2).
2 Details of Stage 2
Recall that and . Depending on on whether or we follow different strategies to define the expectations of indicators .
First we set , for all It follows that
For the definition of , we make use of Ehm’s Binomial approximation to the Poisson Binomial distribution, stated as Theorem 5 in Section 2.1. We start by partitioning as , where , and describe below a procedure for defining so that the following hold:
;
for all , is an integer multiple of .
To define , we apply the same procedure to to obtain . Assuming the correctness of our procedure for probabilities the following should also hold:
;
for all , is an integer multiple of .
Now that we have (9), using (8) and Lemma 2 we get (3).
So it suffices to define the properly. To do this, we define the partition where for all :
(Notice that the length of interval used in the definition of is .) Now, for each such that , we define via the following procedure:
Set , , , .
Set ; let be an arbitrary subset of cardinality .
Set , for all ;
for an arbitrary index , set ;
finally, set , for all .
;
for all , is an integer multiple of .
A similar derivation gives . So by the triangle inequality:
As Eq (10) holds for all , an application of Lemma 2 gives:
Moreover, the ’s defined above are integer multiples of , except maybe for . But we can round these to their closest multiple of , increasing by at most .
Let . We show that the random variable is within total variation distance from the Binomial distribution where
, by the Cauchy-Schwarz inequality; and
.
For fixed and , we set , for all , and , for all , and compare the distributions of and . For convenience we define
The following lemma compares the values , , , .
The proof of Lemma 4 is given in Section 6. To compare and we approximate both by Translated Poisson distributions. Theorem 6 implies that
where for the last inequality we assumed , but the bound of clearly also holds for . Similarly,
where for the last inequality we assumed , but the bound of clearly also holds for . By the triangle inequality we then have that
It remains to bound the total variation distance between the two Translated Poisson distributions. We make use of the following lemma.
where for the last inequality we assumed , but the bound clearly also holds for . Using (15) and (16) we get
Proof of Theorem 3
For all , by combining Theorem 7 and Lemma 6 and we get that
Plugging into Proposition 1, we get
But implies that . So we get in a similar fashion
Deferred Proofs
Proof of Lemma 1: Let and . It is obvious that, if , then the distributions of and are the same. In the other direction, we show that, if and have the same distribution, then . Consider the polynomials:
Since and have the same distribution, and are equal, so they have the same degree and roots. Notice that has degree and roots . Similarly, has degree and roots . Hence, .
Proof of Lemma 2: It follows from the coupling lemma that for any coupling of the variables :
We proceed to fix a specific coupling. For all , it follows from the optimal coupling theorem that there exists a coupling of and such that . Using these individual couplings for each we define a grand coupling of the variables such that , for all . This coupling is faithful because are mutually independent and are mutually independent. Under this coupling Eq (19) implies:
Proof of Claim 1: We use dynamic programming. Let us consider the following tensor of dimension :
Every cell of is assigned value or , as follows:
Inductively, to complete layer , we consider all the non-zero entries of layer and for every such non-zero entry and for every , we find which entry of layer we would transition to if we chose . We set that entry equal to and we also save a pointer to this entry from the corresponding entry of layer , labeling that pointer with the value . The bit operations required to complete layer are bounded by
Therefore, the overall time needed to complete is
Having completed , it is easy to check if there is a solution to . A solution exists if and only if
and can be found by tracing the pointers from this cell of back to level . The overall running time is dominated by the time needed to complete .
Proof of lemma 3: Without loss of generality assume that and denote . For all , denote
Finally, define .
Combining the above we get the result.
As and , we get Moreover, since ,
For all , Condition in the statement of Theorem 3 is equivalent to the following condition:
Next, for all , define to be the power sum symmetric polynomial of degree
The implication is established in a similar fashion. says that
Using the Binomial theorem and induction, we see that (25) implies:
Acknowledgement
We thank the anonymous reviewer for comments that helped improve the presentation.