Faster Algorithms for Privately Releasing Marginals
Justin Thaler, Jonathan Ullman, Salil Vadhan
Introduction
Consider a database in which each of the rows corresponds to an individual’s record, and each record consists of binary attributes. The goal of privacy-preserving data analysis is to enable rich statistical analyses on the database while protecting the privacy of the individuals. In this work, we seek to achieve differential privacy , which guarantees that no individual’s data has a significant influence on the information released about the database.
One of the most important classes of statistics on a dataset is its marginals. A marginal query is specified by a set and a pattern . The query asks, “What fraction of the individual records in has each of the attributes set to ?” A major open problem in privacy-preserving data analysis is to efficiently create a differentially private summary of the database that enables analysts to answer each of the marginal queries. A natural subclass of marginals are -way marginals, the subset of marginals specified by sets such that .
Privately answering marginal queries is a special case of the more general problem of privately answering counting queries on the database, which are queries of the form, “What fraction of individual records in satisfy some property ?” Early work in differential privacy showed how to approximately answer any set of of counting queries by perturbing the answers with appropriately calibrated noise, providing good accuracy (say, within of the true answer) as long as .
Given this state of affairs, it is natural to seek efficient algorithms capable of privately releasing approximate answers to marginal queries even when . A recent series of works have shown how to privately release answers to -way marginal queries with small average error (over various distributions on the queries) with both running time and minimum database size much smaller than (e.g. for product distributions and for arbitrary distributions ). Hardt et. al. also gave an algorithm for privately releasing -way marginal queries with small worst-case error and minimum database size much smaller than . However the running time of their algorithm is still , which is polynomial in the number of queries.
In this paper, we give faster algorithms for releasing marginals and other classes of counting queries.
For notational convenience, we focus on monotone -way disjunction queries. However, our results extend straightforwardly to general non-monotone -way disjunction queries (see Section 4.1), which are equivalent to -way marginals. A monotone -way disjunction is specified by a set of size and asks what fraction of records in have at least one of the attributes in set to .
Our algorithm is inspired by a series of works reducing the problem of private query release to various problems in learning theory. One ingredient in this line of work is a shift in perspective introduced by Gupta, Hardt, Roth, and Ullman . Instead of viewing disjunction queries as a set of functions on the database, they view the database as a function , in which each vector is interpreted as the indicator vector of a set , and equals the evaluation of the disjunction specified by on the database . They use the structure of the functions to privately learn an approximation that has small average error over any product distribution on disjunctions.In their learning algorithm, privacy is defined with respect to the rows of the database that defines , not with respect to the examples given to the learning algorithm (unlike earlier works on “private learning” ).
Cheraghchi, Klivans, Kothari, and Lee observed that the functions can be approximated by a low-degree polynomial with small average error over the uniform distribution on disjunctions. They then use a private learning algorithm for low-degree polynomials to release an approximation to ; and thereby obtain an improved dependence on the accuracy parameter, as compared to .
Hardt, Rothblum, and Servedio observe that is itself an average of disjunctions (each row of specifies a disjunction of bits in the indicator vector of the query), and thus develop private learning algorithms for threshold of sums of disjunctions. These learning algorithms are also based on low-degree approximations of sums of disjunctions. They show how to use their private learning algorithms to obtain a sanitizer with small average error over arbitrary distributions with running time and minimum database size . They then are able to apply the private boosting technique of Dwork, Rothblum, and Vadhan to obtain worst-case accuracy guarantees. Unfortunately, the boosting step incurs a blowup of in the running time.
We improve the above results by showing how to directly compute (a noisy version of) a polynomial that is privacy-preserving and still approximates on all -way disjunctions, as long as is sufficiently large. Specifically, the running time and the database size requirement of our algorithm are both polynomial in the number of monomials in , which is . By “directly”, we mean that we compute from the database itself and perturb its coefficients, rather than using a learning algorithm. Our construction of the polynomial uses the same low-degree approximations exploited by Hardt et. al. in the development of their private learning algorithms.
In summary, the main difference between prior work and ours is that prior work used learning algorithms that have restricted access to the database, and released the hypothesis output by the learning algorithm. In contrast, we do not make use of any learning algorithms, and give our release algorithm direct access to the database. This enables our algorithm to achieve a worst-case error guarantee while maintaining a minimal database size and running time much smaller than the size of the query set. Our algorithm is also substantially simpler than that of Hardt et. al.
We also consider other families of counting queries. We define the class of -of- queries. Like a monotone -way disjunction, an -of- query is defined by a set such that . The query asks what fraction of the rows of have at least of the attributes in set to . For , these queries are exactly monotone -way disjunctions, and -of- queries are a strict generalization.
As an example application, consider a database that allows high school students to express their preferences for colleges in the form of a decision list. For example, a student may say, “If the school is ranked in the top ten nationwide, I am willing to apply to it. Otherwise, if the school is rural, I am unwilling to apply. Otherwise, if the school has a good basketball team then I am willing to apply to it.” And so on. Each student is allowed to use up to attributes out of a set of binary attributes. Our sanitizer allows any college (represented by its binary attributes) to determine the fraction of students willing to apply.
For comparison, we note that all the results on releasing -way disjunctions (including ours) also apply to a dual setting where the database records specify a -way disjunction over bits and the queries are -bit strings (in this setting plays the role of ). Theorem 1.3 generalizes this dual version of Theorem 1.1, as length- decision lists are a strict generalization of -way disjunctions.
We prove the latter two results (Theorems 1.2 and 1.3) using the same approach outlined for marginals (Theorem 1.1), but with different low-degree polynomial approximations appropriate for the different types of queries.
An attractive type of summary is a synthetic database. A synthetic database is a new database whose rows are “fake”, but such that approximately preserves many of the statistical properties of the database (e.g. all the marginals). Some of the previous work on counting query release has provided synthetic data, starting with Barak et. al. and including .
Preliminaries
Let a database be a collection of rows from a data universe . We say that two databases are adjacent if they differ only on a single row, and we denote this by .
A sanitizer takes a database as input and outputs some data structure in . We are interested in sanitizers that satisfy differential privacy.
A sanitizer is -differentially private if for every two adjacent databases and every subset , In the case where we say that is -differentially private.
Since a sanitizer that always outputs satisfies Definition 2.1, we also need to define what it means for a sanitizer to be accurate. In particular, we are interested in sanitizers that give accurate answers to counting queries. A counting query is defined by a boolean predicate . We define the evaluation of the query on a database to be We use to denote a set of counting queries.
An output of a sanitizer is -accurate for the query set if for every . A sanitizer is -accurate for the query set if for every database ,
where the probability is taken over the coins of .
The choice of the norm in the accuracy guarantee of the lemma is for convenience, and doesn’t matter for the parameters of Theorems 1.1-1.3 (except for the hidden constants).
If the privacy requirement is relaxed to -differential privacy (for , then it is sufficient to perturb each coordinate of with noise from a Laplace distribution of smaller magnitude, leading to smaller error.
2 Query Function Families
We take the approach of Gupta et. al. and think of the database as specifying a function mapping queries to their answers , which we call the -representation of . We now describe this transformation more formally:
Let be a set of counting queries on a data universe , where each query is indexed by an -bit string. We define the index set of to be the set .
For some intuition about this transformation, when the queries are monotone -way disjunctions on a database , the queries are defined by sets , . In this case each query can be represented by the -bit indicator vector of the set , with at most non-zero entries. Thus we can take and .
3 Polynomial Approximations
Let be the family of all -variate real polynomials of degree and norm . In many cases, the functions can be approximated well on all the indices in by a family of polynomials with low degree and small norm. Formally:
Given a family of -variate functions and a set , we say that the family uniformly -approximates on if for every , there exists such that .
From Polynomial Approximations to Data Release Algorithms
In this section we present an algorithm for privately releasing any family of counting queries such that that can be efficiently and uniformly approximated by polynomials. The algorithm will take an -row database and, for each row , constructs a polynomial that uniformly approximates the function (recall that , for each ). From these, it constructs a polynomial that uniformly approximates . The final step is to perturb each of the coefficients of using noise from a Laplace distribution (Theorem 2.3) and bound the error introduced from the perturbation.
is -accurate for for
First we construct the sanitizer . See the relevant codebox below.
We establish that is -differentially private. This follows from the observation that for any two adjacent that differ only on row ,
The last inequality is from the fact that for every , is a vector of norm at most . Part 1 of the Theorem now follows directly from the properties of the Laplace Mechanism (Theorem 2.3). Now we construct the evaluator .
Efficiency.
Accuracy. Finally, we analyze the accuracy of the sanitizer . First, by the assumption that uniformly -approximates on , we have
where the probability is taken over the coins of . Part (3) of the Theorem will then follow by the triangle inequality.
The first inequality follows from the fact that every monomial evaluates to or at the point . This completes the proof of the theorem.
Using Theorem 2.4, we can improve the bound on the error at the expense of relaxing the privacy guarantee to -differential privacy. This improved error only affects the hidden constants in Theorems 1.1-1.3, so we only state those theorems for -differential privacy.
is -differentially private,
is -accurate for for
The proof of this theorem is identical to that of Theorem 3.1, but using the analysis of the Laplace mechanism from Theorem 2.4 in place of that of Theorem 2.3.
Applications
In this section we establish the existence of explicit families of low-degree polynomials approximating the families for some interesting query sets.
We define the class of monotone -way disjunctions as follows:
for every ,
for every , .
We can use Lemma 4.2 to approximate -way monotone disjunctions. Note that our result easily extends to monotone -way conjunctions via the identity . Moreover, it extends to non-monotone conjunctions and disjunctions: we may extend the data universe as in [15, Theorem 1.2] to , and include the negation of each item in the original domain. Non-monotone conjunctions over domain correspond to monotone conjunctions over the expanded domain .
Theorem 1.1 in the introduction follows by combining Theorems 3.1 and 4.3.
2 Releasing Monotone r𝑟r-of-k𝑘k Queries
We define the class of monotone -of- queries as follows:
Sherstov [20, Lemma 3.11] gives an explicit construction of polynomials that can be used to approximate the family over with low degree. It can be verified by inspecting the construction that the coefficients of the resulting polynomial are not too large.
,
for every , , and
for every , .
For completeness we include a proof of Lemma 4.5 in the appendix. We can use these polynomials to approximate monotone -of- queries.
The construction and proof is identical to that of Theorem 4.3 with the polynomials of Lemma 4.5 in place of the polynomials described in Lemma 4.2. ∎
Theorem 1.2 in the introduction now follows by combining Theorems 3.1 and 4.6. Note that our result also extends easily to non-monotone -of- queries in the same manner as Theorem 1.1.
Using the principle of inclusion-exclusion, the answer to a monotone -of- query can be written as a linear combination of the answers to monotone -way disjunctions. Thus, a sanitizer that is -accurate for monotone -way disjunctions implies a sanitizer that is -accurate for monotone -of- queries. However, combining this implication with Theorem 1.1 yields a sanitizer with running time , which has a worse dependence on than what we achieve in Theorem 1.2.
3 Releasing Decision Lists
We obtain Theorem 1.3 of the introduction by combining Theorems 3.1 and 4.9.
Generalizations and Limitations of Our Approach
Acknowledgements
We thank Vitaly Feldman, Moritz Hardt, Varun Kanade, Aaron Roth, Guy Rothblum, and Li-Yang Tan for helpful discussions.
References
Appendix A Polynomial Approximation of Decision Lists
At a high level, we treat each term of the above sum independently, using a transformation of the Chebyshev polynomials to approximate each term within additive error . This ensures that that the sum of the resulting polynomials approximates within additive error as desired. Details follow.
Let be the polynomial described in Lemma 4.2 with error parameter . Then the polynomial satisfies the following properties:
The degree of is ,
for every , .
Consider the polynomial defined as
Appendix B Polynomial Approximation of r𝑟r-of-k𝑘k Queries
,
for every , , and
for every , .
Let be the degree Chebyshev polynomial of the first kind (Fact B.2). We will use the following well-known properties of Chebyshev polynomials.
The Chebyshev polynomials of the first kind satisfy the following properties.
Each coefficient of has absolute value at most .
Let \Delta=\big{\lceil}\frac{\log(k/\gamma)}{\log n}\big{\rceil}, and z=3\Delta\big{\lceil}\log k\big{\rceil}. The construction proceeds in several steps, with the final polynomial defined in terms of multiple intermediate polynomials.
The coefficients of have absolute value at most .
The final polynomial is defined as