A Linear-Optical Proof that the Permanent is #P-Hard

Scott Aaronson

Introduction

Given an n×nn\times n matrix A=(ai,j)A=\left(a_{i,j}\right), the permanent of AA is defined as

A seminal result of Valiant says that computing Per⁡(A)\operatorname*{Per}\left(A\right) is #P\mathsf{\#P}-hard, if AA is a matrix over (say) the integers, the nonnegative integers, or the set {0,1}\left\{0,1\right\}.See Hrubes, Wigderson, and Yehudayoff for a recent, “modular” presentation of Valiant’s proof (which also generalizes the proof to the noncommutative and nonassociative case). Here #P\mathsf{\#P} means (informally) the class of counting problems—problems that involve summing exponentially-many nonnegative integers—and #P\mathsf{\#P}-hard means “at least as hard as any #P\mathsf{\#P} problem.”See the Complexity Zoo (www.complexityzoo.com) for the definitions of #P\mathsf{\#P} and other complexity classes used in this paper.,If AA is a nonnegative integer matrix, then Per⁡(A)\operatorname*{Per}\left(A\right) is itself a #P\mathsf{\#P} function, which implies that it is #P\mathsf{\#P}-complete (the term for functions that are both #P\mathsf{\#P}-hard and in #P\mathsf{\#P}). If AA can have negative or fractional entries, then strictly speaking Per⁡(A)\operatorname*{Per}\left(A\right) is no longer #P\mathsf{\#P}-complete, but it is still #P\mathsf{\#P}-hard and computable in the class FP#P\mathsf{FP}^{\mathsf{\#P}}.

More concretely, Valiant gave a polynomial-time algorithm that takes as input an instance φ(x1,…,xn)\varphi\left(x_{1},\ldots,x_{n}\right) of the Boolean satisfiability problem, and that outputs a matrix AφA_{\varphi} such that Per⁡(Aφ)\operatorname*{Per}\left(A_{\varphi}\right) encodes the number of satisfying assignments of φ\varphi. This means that computing the permanent is at least as hard as counting satisfying assignments.

Unfortunately, the standard proof that the permanent is #P\mathsf{\#P}-hard is notoriously opaque; it relies on a set of gadgets that seem to exist for “accidental” reasons. Could there be an alternative proof that gave more, or at least different, insight? In this paper, we try to answer that question by giving a new, quantum-computing-based proof that the permanent is #P\mathsf{\#P}-hard. In particular, we will derive the permanent’s #P\mathsf{\#P}-hardness as a consequence of the following three facts:

Postselected linear optics is capable of universal quantum computation, as shown in a celebrated 2001 paper of Knill, Laflamme, and Milburn (henceforth referred to as KLM).KLM actually prove the stronger (and more practically-relevant) result that linear optics with adaptive measurements is capable of universal quantum computation. For our purposes, however, we only need the weaker fact that postselected measurements suffice for universal QC, which KLM prove as a lemma along the way to their main result.

Quantum computations can encode #P\mathsf{\#P}-hard quantities in their amplitudes.

Amplitudes in nn-photon linear-optics circuits can be expressed as the permanents of n×nn\times n matrices.

Even though our proof is based on quantum computing, we stress that we have made it entirely self-contained: all of the results we need (including the KLM Theorem , and even the construction of the Toffoli gate from 11-qubit and CSIGN⁡\operatorname*{CSIGN} gates) are proved in this paper for completeness. We assume some familiarity with quantum computing notation (e.g., kets and quantum circuit diagrams), but not with linear optics.

If one counts the complexity of all of the individual pieces we use—especially the universality results for quantum gates—then our reduction from #P\mathsf{\#P} to the permanent ends up being at least as complicated as Valiant’s, and probably more so. In our view, however, this is similar to how writing a program in C++ tends to produce a longer, more complicated executable file than writing the same program in assembly language. Normally, one also cares about the length and readability of the source code! Our purpose in this paper is to illustrate how quantum computing provides a powerful “high-level programming language” in which one can, among other things, easily rederive the most celebrated result in the theory of #P\mathsf{\#P}-hardness.

But why does the world need a new proof that the permanent is #P\mathsf{\#P}-hard—especially a proof invoking what some might consider to be exotic concepts? Let us offer several answers:

Any theorem as basic as the #P\mathsf{\#P}-hardness of the permanent deserves several independent proofs. And our proof really is “independent” of the standard one: rather than composing variable and clause gadgets,Indeed, our proof does not even go through the Cook-Levin Theorem: it reduces a #P\mathsf{\#P} computation directly to the permanent, without first reducing #P\mathsf{\#P} to #3SAT\#3SAT. we multiply matrices corresponding to quantum gates, and use ideas from linear optics to keep track of how such multiplications affect the permanent. One way to see the difference is that our proof never uses the notion of a cycle cover.

While our proof, like the standard one, requires “gadgets” (one to simulate a Toffoli gate using CSIGN⁡\operatorname*{CSIGN} gates, another to simulate a CSIGN⁡\operatorname*{CSIGN} gate using postselected linear optics), the connection to quantum computing gives those gadgets a natural semantics. In other words, the gadgets were introduced for “practical” reasons having nothing to do with proving the permanent #P\mathsf{\#P}-hard, and can be motivated independently of that goal. If one already knows the quantum universality gadgets, then we offer what seems like a major advance in complexity-theoretic pedagogy: a proof that the permanent is #P\mathsf{\#P}-hard that can be reproduced on-the-spot from memory!

As Kuperberg pointed out, by their nature, any #P\mathsf{\#P}-hardness proofs (including ours) that are based on “quantum postselection” almost immediately yield hardness of approximation results as well.

We expect that the quantum postselection approach used here could lead to #P\mathsf{\#P}-hardness proofs for many other problems—including problems not already known to be #P\mathsf{\#P}-hard by other means. In this direction, one natural place to look would be special cases of the permanent.

2 Related Work

By now, there are many examples where quantum computing has been used to give new or simpler proofs of classical complexity theorems; see Drucker and de Wolf for an excellent survey. Within the area of counting complexity, Aaronson showed that the class PP\mathsf{PP} is equal to PostBQP\mathsf{PostBQP} (quantum polynomial-time with postselection), and then used that theorem to give a simpler proof of the landmark result of Beigel, Reingold, and Spielman that PP\mathsf{PP} is closed under intersection. Later, also using the PostBQP=PP\mathsf{PostBQP}=\mathsf{PP} theorem, Kuperberg gave a “quantum proof” of the result of Jaeger, Vertigan, and Welsh that computing the Jones polynomial is #P\mathsf{\#P}-hard, and even showed that a certain approximate version is #P\mathsf{\#P}-hard (which had not been shown previously). Kuperberg’s argument for the Jones polynomial is conceptually similar to our argument for the permanent.

There is also precedent for using linear optics as a tool to prove theorems about the permanent. Scheel observed that the unitarity of linear-optical quantum computing implies the interesting fact that ∣Per⁡(U)∣≤1\left|\operatorname*{Per}\left(U\right)\right|\leq 1 for all unitary matrices UU.

Rudolph showed how to encode quantum amplitudes directly as matrix permanents, and in the process, gave a “quantum-computing proof” that the permanent is #P\mathsf{\#P}-hard. However, a crucial difference is that Rudolph starts with Valiant’s proof based on cycle covers, then recasts it in quantum terms (with the goal of making Valiant’s proof more accessible to a physics audience). By contrast, our proof is independent of Valiant’s; the tools we use were invented for separate reasons in the quantum computing literature.

There has been a great deal of work on linear-optical quantum computing, beyond the seminal KLM Theorem on which this paper relies. Recently, Aaronson and Arkhipov studied the complexity of sampling from a linear-optical computer’s output distribution, assuming no adaptive measurements are available. By using the #P\mathsf{\#P}-hardness of the permanent as an “input axiom,” they showed that this sampling problem is classically intractable unless P#P=BPPNP\mathsf{P}^{\mathsf{\#P}}=\mathsf{BPP}^{\mathsf{NP}}. More relevant to this paper is an alternative proof that Aaronson and Arkhipov gave for their result. Inspired by work of Bremner, Jozsa, and Shepherd , the alternative proof combines Aaronson’s PostBQP=PP\mathsf{PostBQP}=\mathsf{PP} theorem with the fact that postselected linear optics is universal for PostBQP\mathsf{PostBQP}, and thereby avoids any direct appeal to the #P\mathsf{\#P}-hardness of the permanent. In retrospect, that proof was already much of the way toward a linear-optical proof that the permanent is #P\mathsf{\#P}-hard; this paper simply makes the connection explicit.

Background

Not by accident, this section constitutes the bulk of the paper. First, in Section 2.1, we fix some facts and notation about standard (qubit-based) quantum computing. Then, in Section 2.2, we give a short overview of those aspects of linear-optical quantum computing that are relevant for us, and (for completeness) prove the KLM Theorem in the specific form we will need.

Abusing notation, we will often identify a quantum circuit QQ with the unitary transformation that it induces: for example, ⟨0⋯0∣Q∣0⋯0⟩\left\langle 0\cdots 0\right|Q\left|0\cdots 0\right\rangle represents the amplitude with which QQ maps its initial state to itself. We use ∣Q∣\left|Q\right| to denote the number of gates in QQ.

The first ingredient we need for our proof is a convenient set of quantum gates (in the standard qubit model). Thus, let G\mathcal{G} be the set of gates consisting of (1) all 11-qubit gates, and (2) the 22-qubit controlled-sign gate

which flips the amplitude if and only if both qubits are ∣1⟩\left|1\right\rangle.A more common 22-qubit gate than CSIGN⁡\operatorname*{CSIGN} is the controlled-NOT (CNOT⁡\operatorname*{CNOT}) gate, which maps each basis state ∣x,y⟩\left|x,y\right\rangle to ∣x,y⊕x⟩\left|x,y\oplus x\right\rangle. However, CSIGN⁡\operatorname*{CSIGN} is more convenient for linear-optics purposes, and is equivalent to CNOT⁡\operatorname*{CNOT} by conjugating the second qubit with a Hadamard gate. Then Barenco et al. showed that G\mathcal{G} is a universal set of quantum gates, in the sense that G\mathcal{G} generates any unitary transformation on any number of qubits (without error). For our purposes, however, the following weaker result suffices.

G\mathcal{G} generates the Toffoli gate, the 33-qubit gate that maps each basis state ∣x,y,z⟩\left|x,y,z\right\rangle to ∣x,y,z⊕xy⟩\left|x,y,z\oplus xy\right\rangle.

Proof. The circuit can be found in Nielsen and Chuang for example, but we reproduce it in Figure 1 for completeness. In the diagram,

is another 11-qubit gate, and the six vertical bars represent CSIGN⁡\operatorname*{CSIGN} gates.

2 Linear-Optical Quantum Computing

We now give a brief overview of linear-optical quantum computing (LOQC), an alternative quantum computing model based on identical photons rather than qubits. For a detailed introduction to LOQC from a computer science perspective, see Aaronson and Arkhipov .

In LOQC, each basis state of our quantum computer has the form ∣S⟩=∣s1,…,sm⟩\left|S\right\rangle=\left|s_{1},\ldots,s_{m}\right\rangle, where s1,…,sms_{1},\ldots,s_{m} are nonnegative integers summing to nn. Here sis_{i} represents the number of photons in the ithi^{th} location or “mode,” and the fact that s1+⋯+sm=ns_{1}+\cdots+s_{m}=n means that photons are never created or destroyed. One should think of mm and nn as both polynomially-bounded. For this paper, it will be convenient to assume that mm is even, that n=m/2n=m/2, and that the initial state has the form ∣I⟩=∣0,1,0,1,…,0,1⟩\left|I\right\rangle=\left|0,1,0,1,\ldots,0,1\right\rangle: that is, one photon in each even-numbered mode, and no photons in the odd-numbered modes.

Let Φm,n\Phi_{m,n} be the set of nonnegative integer tuples S=(s1,…,sm)S=\left(s_{1},\ldots,s_{m}\right) such that s1+⋯+sm=ns_{1}+\cdots+s_{m}=n, and let Hm,n\mathcal{H}_{m,n} be the Hilbert space spanned by basis states ∣S⟩\left|S\right\rangle with S∈Φm,nS\in\Phi_{m,n}. Then a general state in LOQC is just a unit vector in Hm,n\mathcal{H}_{m,n}:

with ∑S∈Φm,n∣αS∣2=1\sum_{S\in\Phi_{m,n}}\left|\alpha_{S}\right|^{2}=1.

To transform ∣ψ⟩\left|\psi\right\rangle, one can select any m×mm\times m unitary transformation U=(uij)U=\left(u_{ij}\right). This UU then induces a larger unitary transformation φ(U)\varphi\left(U\right) on the Hilbert space Hm,n\mathcal{H}_{m,n} of nn-photon states. There are several ways to define φ(U)\varphi\left(U\right), but perhaps the simplest is the following formula:

for all tuples S=(s1,…,sm)S=\left(s_{1},\ldots,s_{m}\right) and T=(t1,…,tm)T=\left(t_{1},\ldots,t_{m}\right) in Φm,n\Phi_{m,n}. Here US,TU_{S,T} is the n×nn\times n matrix obtained from UU by taking sis_{i} copies of the ithi^{th} row of UU and tjt_{j} copies of the jthj^{th} column, for all i,j∈[m]i,j\in\left[m\right]. To illustrate, if

and ∣S⟩=∣T⟩=∣2,1⟩\left|S\right\rangle=\left|T\right\rangle=\left|2,1\right\rangle, then

Intuitively, the reason the permanent arises in formula (*) is that there are n!n! ways of mapping the nn photons in basis state ∣S⟩\left|S\right\rangle onto the nn photons in basis state ∣T⟩\left|T\right\rangle. Since the photons are identical bosons, quantum mechanics says that each of those n!n! ways contributes a term to the total ⟨S∣φ(U)∣T⟩\left\langle S|\varphi\left(U\right)|T\right\rangle, with the contribution given by the product of the transition amplitudes uiju_{ij} for each of the nn photons individually.

It turns out that φ(U)\varphi\left(U\right) is always unitary and that φ\varphi is a homomorphism. Both facts seem surprising viewed purely as algebraic consequences of formula (*), but of course they have natural physical interpretations: φ(U)\varphi\left(U\right) is unitary because it represents an actual physical transformation that can be applied, and φ\varphi is a homomorphism because generalizing from one photon to nn photons must commute with composing beamsplitters. In this paper, we will not need that φ(U)\varphi\left(U\right) is unitary; see Aaronson and Arkhipov for a proof of that fact. Below we prove that φ\varphi is a homomorphism.

Proof. We want to show that for all tuples S,T∈Φm,nS,T\in\Phi_{m,n} and all m×mm\times m unitaries U,VU,V,

By equation (*), the above is equivalent (after multiplying both sides by s1!⋯sm!t1!⋯tm!\sqrt{s_{1}!\cdots s_{m}!t_{1}!\cdots t_{m}!}) to the identity

We will prove identity (**) in the special case n=mn=m and S=T=I=(1,1,…,1)S=T=I=\left(1,1,\ldots,1\right), since the general case is analogous. We have

In the second line above, we decomposed the sum by thinking about each permutation σ∈Sn\sigma\in S_{n} as a product of two permutations: one, τ\tau, that maps nn particles in the initial configuration ∣I⟩\left|I\right\rangle to nn particles in the intermediate configuration ∣R⟩\left|R\right\rangle when UU is applied, and another, ξ\xi, that maps nn particles in the intermediate configuration ∣R⟩\left|R\right\rangle to nn particles in the final configuration ∣I⟩\left|I\right\rangle when VV is applied. This yields the same result, as long as we remember to sum over all possible intermediate configurations R∈Φn,nR\in\Phi_{n,n}, and also to divide each summand by r1!⋯rn!r_{1}!\cdots r_{n}!, which is the size of RR’s automorphism group (i.e., the number of ways to permute the nn particles within ∣R⟩\left|R\right\rangle that leave ∣R⟩\left|R\right\rangle unchanged).

In the standard qubit model, every unitary transformation can be decomposed as a product of gates, each of which acts nontrivially on only 11 or 22 qubits. Similarly, in LOQC, every unitary transformation can be decomposed as a product of linear-optics gates, each of which acts nontrivially on only 11 or 22 modes. Then a linear-optics circuit is simply a list of linear-optics gates applied to specified modes (or pairs of modes) starting from the initial state ∣I⟩=∣0,1,…,0,1⟩\left|I\right\rangle=\left|0,1,\ldots,0,1\right\rangle.A crucial difference between standard quantum circuits and linear-optics circuits is that, whereas a standard quantum gate is the tensor product of a small (say 4×44\times 4) unitary matrix with an exponentially-large (say 2n−2×2n−22^{n-2}\times 2^{n-2}) identity matrix, a linear-optics gate is the direct sum of a small (say 2×22\times 2) unitary matrix with a polynomially-large (say (m−2)×(m−2)\left(m-2\right)\times\left(m-2\right)) identity matrix. It is only the homomorphism U→φ(U)U\rightarrow\varphi\left(U\right) that produces exponentially-large matrices. One consequence, pointed out by Reck et al. , is that, whereas most nn-qubit unitary transformations require Ω(22n)\Omega\left(2^{2n}\right) gates to implement (as follows from an easy dimension argument), every mm-mode unitary transformation UU can be implemented using only O(m2)O\left(m^{2}\right) linear-optics gates.

The last notion we need is that of postselected LOQC. In our context, postselection simply means measuring the number of photons in a given mode ii, and conditioning on a particular result (for example, photons, or 11 photon). After we postselect on the number of photons in some mode, we will never use that mode for further computation.In physics language, all photon-number measurements are assumed to be “demolition” measurements. For this reason, without loss of generality, we can defer all postselected measurements until the end of the computation.

Our #P\mathsf{\#P}-hardness proof will fall out as a corollary of the following universality theorem, which is implicit in the work of KLM . Indeed, we could just appeal to the KLM construction as a “black box,” but we choose not to do so, since the properties of the construction that we want are slightly different from the properties KLM want, and we wish to verify in detail that the desired properties hold.

Postselected linear optics can simulate universal quantum computation. More concretely: there exists a polynomial-time classical algorithm that converts a quantum circuit QQ over the gate set G\mathcal{G} into a linear-optics circuit LL, so that

where Γ\Gamma is the number of CSIGN⁡\operatorname*{CSIGN} gates in QQ and ∣I⟩=∣0,1,…,0,1⟩\left|I\right\rangle=\left|0,1,\ldots,0,1\right\rangle is the standard initial state.

Proof. To encode a (qubit-based) quantum circuit by a postselected linear-optics circuit, KLM use the so-called dual-rail representation of a qubit using two optical modes. In this representation, the qubit ∣0⟩\left|0\right\rangle is represented as ∣0,1⟩\left|0,1\right\rangle, while the qubit ∣1⟩\left|1\right\rangle is represented as ∣1,0⟩\left|1,0\right\rangle. Thus, to simulate a quantum circuit that acts on kk qubits, we need 2k2k optical modes. (We will also need additional modes to handle postselection, but we can ignore those for now.) Let the modes corresponding to qubit ii be labeled (i,0)\left(i,0\right) and (i,1)\left(i,1\right) respectively. Notice that the initial state ∣0⋯0⟩\left|0\cdots 0\right\rangle in the qubit model maps onto the initial state ∣I⟩\left|I\right\rangle in the optical model.

Since φ\varphi is a homomorphism by Lemma 2, to prove the theorem it suffices to show how to simulate the gates in G\mathcal{G}. Simulating a 11-qubit gate is easy: simply apply the appropriate 2×22\times 2 unitary transformation to the Hilbert space spanned by ∣0,1⟩\left|0,1\right\rangle and ∣1,0⟩\left|1,0\right\rangle. The interesting part is how to simulate a CSIGN⁡\operatorname*{CSIGN} gate. To do so, KLM use another gate that they call NS⁡1\operatorname*{NS}\nolimits_{1}, which applies the following unitary transformation to a single mode:

(We do not care how NS⁡1\operatorname*{NS}\nolimits_{1} acts on ∣3⟩\left|3\right\rangle, ∣4⟩\left|4\right\rangle, and so on, since those basis states will never arise in our simulation.) Using NS⁡1\operatorname*{NS}\nolimits_{1}, it is not hard to simulate CSIGN⁡\operatorname*{CSIGN} on two qubits ii and jj. The procedure, shown in Figure 2, is this: first apply a Hadamard transformation to modes (i,0)\left(i,0\right) and (j,0)\left(j,0\right).

One can check that this induces the following transformation on the state of (i,0)\left(i,0\right) and (j,0)\left(j,0\right):

The key point is that we get a state involving 22 photons in the same mode, if and only if the modes (i,0)\left(i,0\right) and (j,0)\left(j,0\right) both contained a photon. Next, apply NS⁡1\operatorname*{NS}\nolimits_{1} gates to both (i,0)\left(i,0\right) and (j,0)\left(j,0\right). This flips the amplitude if and only if we started with ∣1,1⟩\left|1,1\right\rangle. Finally, apply a second Hadamard transformation to (i,0)\left(i,0\right) and (j,0)\left(j,0\right), to complete the implementation of CSIGN⁡\operatorname*{CSIGN}.

We now explain how to implement NS⁡1\operatorname*{NS}\nolimits_{1} on a given mode ii, using postselection. To do so, we need two additional modes jj and kk, which are initialized to the states ∣0⟩\left|0\right\rangle and ∣1⟩\left|1\right\rangle respectively. First we apply the following 3×33\times 3 unitary transformation to i,j,ki,j,k:

Then we postselect on jj and kk being returned to the state ∣0,1⟩\left|0,1\right\rangle. As shown in , this postselection always succeeds with amplitude 1/21/2 (corresponding to probability 1/41/4); and that conditioned on it succeeding, the effect is to apply NS⁡1\operatorname*{NS}\nolimits_{1} in mode ii. To prove this, observe that since the number of photons is conserved, the effect of WW on mode ii must have the form

for some λ0,λ1,λ2\lambda_{0},\lambda_{1},\lambda_{2}. Using formula (*), we then calculate

This implies that the CSIGN⁡\operatorname*{CSIGN} circuit shown in Figure 2 succeeds with amplitude 1/41/4 (corresponding to probability 1/161/16), and furthermore, we know when it succeeds.

In the proof of Theorem 3, the main reason the matrix WW looks complicated is simply that it needs to be unitary. However, notice that unitarity is irrelevant for our #P\mathsf{\#P}-hardness application—and if we drop the unitarity requirement, then we can replace WW by a simpler 2×22\times 2 matrix, such as

To implement NS⁡1\operatorname*{NS}\nolimits_{1} on a given mode ii, we would apply YY to ii as well as another mode jj that initially contains one photon, then postselect on jj still containing one photon after YY is applied. One can verify by calculation that the effect on mode ii is

where λ0=λ1=1\lambda_{0}=\lambda_{1}=1 and λ2=−1\lambda_{2}=-1.

Main Result

In this section we deduce the following theorem, as a straightforward consequence of Theorem 3.

In classical complexity theory, one is often more interested in various corollaries of Theorem 4: for example, that computing Per⁡(A)\operatorname*{Per}\left(A\right) remains #P\mathsf{\#P}-hard even if AA is a nonnegative integer matrix, or a {−1,0,1}\left\{-1,0,1\right\}-valued matrix, or a {0,1}\left\{0,1\right\}-valued matrix. Valiant gave simple reductions by which one can deduce all of these corollaries from Theorem 4. We do not know how to use the linear-optics perspective to get any additional insight into the corollaries.

Let CC be a classical circuit that computes a Boolean function C:{0,1}n→{−1,1}C:\left\{0,1\right\}^{n}\rightarrow\left\{-1,1\right\}, and let ΔC:=∑x∈{0,1}nC(x)\Delta_{C}:=\sum_{x\in\left\{0,1\right\}^{n}}C\left(x\right). Then computing ΔC\Delta_{C}, given CC as input, is a #P\mathsf{\#P}-hard problem essentially by definition. On the other hand, it is easy to encode ΔC\Delta_{C} as an amplitude in a quantum circuit:

There exists a classical algorithm that takes a circuit CC as input, runs in poly⁡(n,∣C∣)\operatorname*{poly}\left(n,\left|C\right|\right) time, and outputs a (qubit-based) quantum circuit QQ, consisting of gates from G\mathcal{G}, such that

Proof. Let DCD_{C} be a 2n×2n2^{n}\times 2^{n} diagonal unitary matrix whose (x,x)\left(x,x\right) entry is C(x)C\left(x\right). Then since the Toffoli gate is universal for classical computation, a quantum circuit consisting of 11-qubit gates and Toffoli gates can easily apply DCD_{C}. To do so, one uses the standard “uncomputing” trick:

Finally, by Lemma 1, we can simulate each of the Toffoli gates in QQ using gates from the set G\mathcal{G}.

Let QQ be the quantum circuit from Lemma 5, and assume QQ uses k=poly⁡(n,∣C∣)k=\operatorname*{poly}\left(n,\left|C\right|\right) qubits. By Theorem 3, we can simulate QQ by a linear-optics circuit LL such that

where Γ=poly⁡(n,∣C∣)\Gamma=\operatorname*{poly}\left(n,\left|C\right|\right) is the number of CSIGN⁡\operatorname*{CSIGN} gates in QQ. Furthermore, the circuit LL uses m:=2k+4Γm:=2k+4\Gamma optical modes. Let UU be the m×mm\times m unitary matrix induced by LL, and let VV be the (m/2)×(m/2)\left(m/2\right)\times\left(m/2\right) submatrix of UU obtained by taking the even-numbered rows and columns only. Then we have

where the first line follows from formula (*) and the third from Lemma 5. Since VV can be produced in polynomial time given CC, this already shows that computing Per⁡(V)\operatorname*{Per}\left(V\right) to sufficient precision is #P\mathsf{\#P}-hard.

However, we still need to deal with the issue that the entries of VV are real numbers.Indeed, the matrices that we multiply to obtain UU can be complex matrices, but UU itself (and hence the submatrix VV) will always be real. Let b:=⌈log⁡2(n!)+2n+2Γ⌉b:=\left\lceil\log_{2}\left(n!\right)+2n+2\Gamma\right\rceil. Then notice that truncating the entries of VV to bb bits of precision produces a matrix V~\widetilde{V} such that

For this reason, we can assume that each entry of VV has the form k/2bk/2^{b} for some integer k∈[−2b,2b]k\in\left[-2^{b},2^{b}\right]. Now set A:=2bVA:=2^{b}V. Then AA is an integer matrix satisfying Per⁡(A)=2bnPer⁡(V)\operatorname*{Per}\left(A\right)=2^{bn}\operatorname*{Per}\left(V\right), whose entries can be specified using b+O(1)=poly⁡(n,∣C∣)b+O\left(1\right)=\operatorname*{poly}\left(n,\left|C\right|\right) bits each. This completes the proof of Theorem 4.

We conclude by noticing that our proof yields not only Theorem 4, but also the following corollary:

Proof. By the above equivalences, it suffices to show that computing sgn⁡(ΔC)\operatorname*{sgn}\left(\Delta_{C}\right) is #P\mathsf{\#P}-hard. This is true because, given the ability to compute sgn⁡(ΔC)\operatorname*{sgn}\left(\Delta_{C}\right), we can determine ΔC\Delta_{C} exactly using binary search. In more detail, given a positive integer kk, let C[k]C\left[k\right] denote the circuit CC modified to contain kk additional inputs xx such that C(x)=1C\left(x\right)=1, and let C[−k]C\left[-k\right] denote CC modified to contain kk additional xx’s such that C(x)=−1C\left(x\right)=-1. Then clearly

Thus we can use the following strategy: compute the signs of ΔC[1],ΔC[−1],ΔC[2],ΔC[−2],ΔC[4],ΔC[−4],\Delta_{C\left[1\right]},\Delta_{C\left[-1\right]},\Delta_{C\left[2\right]},\Delta_{C\left[-2\right]},\Delta_{C\left[4\right]},\Delta_{C\left[-4\right]},and so on, increasing kk by successive factors of 22, until a kk is found such that sgn⁡(ΔC[k])≠sgn⁡(ΔC[2k])\operatorname*{sgn}\left(\Delta_{C\left[k\right]}\right)\neq\operatorname*{sgn}\left(\Delta_{C\left[2k\right]}\right). At that point, we know that ΔC\Delta_{C} must be between kk and 2k2k. Then by computing sgn⁡(ΔC[3k/2])\operatorname*{sgn}\left(\Delta_{C\left[3k/2\right]}\right), we can decide whether ΔC\Delta_{C} is between kk and 3k/23k/2 or between 3k/23k/2 and 2k2k, and so on recursively until ΔC\Delta_{C} has been determined exactly.

Corollary 6 implies, in particular, that approximating Per⁡(A)\operatorname*{Per}\left(A\right) to within any multiplicative factor is #P\mathsf{\#P}-hard—since to output a multiplicative approximation, at the least we would need to know whether Per⁡(A)\operatorname*{Per}\left(A\right) is positive or negative.

Using a more involved binary search strategy (which we omit), one can show that, for any β(N)∈[1,poly⁡(N)]\beta\left(N\right)\in\left[1,\operatorname*{poly}\left(N\right)\right], even approximating ∣ΔC∣\left|\Delta_{C}\right| or ΔC2\Delta_{C}^{2} to within a multiplicative factor of β(N)\beta\left(N\right) would let one compute ΔC\Delta_{C} exactly, and is therefore #P\mathsf{\#P}-hard under Turing reductions. It follows from this that approximating ∣Per⁡(A)∣\left|\operatorname*{Per}\left(A\right)\right| or Per⁡(A)2\operatorname*{Per}\left(A\right)^{2} to within a multiplicative factor of β(N)\beta\left(N\right) is #P\mathsf{\#P}-hard as well. (Aaronson and Arkhipov gave a related but more complicated proof of the #P\mathsf{\#P}-hardness of approximating ∣Per⁡(A)∣\left|\operatorname*{Per}\left(A\right)\right| and Per⁡(A)2\operatorname*{Per}\left(A\right)^{2}, which did not first replace Per⁡(A)\operatorname*{Per}\left(A\right) with ΔC\Delta_{C}.)

Acknowledgments

I am grateful to Alex Arkhipov and Michael Forbes for helpful discussions, and to Andy Drucker, Greg Kuperberg, Avi Wigderson, Ronald de Wolf, and the anonymous reviewers for their comments.

References