Informational derivation of Quantum Theory
G. Chiribella, G. M. D'Ariano, P. Perinotti
I Introduction
More than eighty years after its formulation, quantum theory is still mysterious. The theory has a solid mathematical foundation, addressed by Hilbert, von Neumann and Nordheim in 1928 HNN28 and brought to completion in the monumental work by von Neumann von32. However, this formulation is based on the abstract framework of Hilbert spaces and self-adjoint operators, which, to say the least, are far from having an intuitive physical meaning. For example, the postulate stating that the pure states of a physical system are represented by unit vectors in a suitable Hilbert space appears as rather artificial: which are the physical laws that lead to this very specific choice of mathematical representation? The problem with the standard textbook formulations of quantum theory is that the postulates therein impose particular mathematical structures without providing any fundamental reason for this choice: the mathematics of Hilbert spaces is adopted without further questioning as a prescription that “works well” when used as a black box to produce experimental predictions. In a satisfactory axiomatization of Quantum Theory, instead, the mathematical structures of Hilbert spaces (or C*-algebras) should emerge as consequences of physically meaningful postulates, that is, postulates formulated exclusively in the language of physics: this language refers to notions like physical system, experiment, or physical process and not to notions like Hilbert space, self-adjoint operator, or unitary operator. Note that any serious axiomatization has to be based on postulates that can be precisely translated in mathematical terms. However, the point with the present status of quantum theory is that there are postulates that have a precise mathematical statement, but cannot be translated back into language of physics. Those are the postulates that one would like to avoid.
The need for a deeper understanding of quantum theory in terms of fundamental principles was clear since the very beginning. Von Neumann himself expressed his dissatisfaction with his mathematical formulation of Quantum Theory with the surprising words “I don’t believe in Hilbert space anymore”, reported by Birkhoff in Birk61. Realizing the physical relevance of the axiomatization problem, Birkhoff and von Neumann made an attempt to understand quantum theory as a new form of logic BirkVN36: the key idea was that propositions about the physical world must be treated in a suitable logical framework, different from classical logics, where the operations AND and OR are no longer distributive. This work inaugurated the tradition of quantum logics, which led to several attempts to axiomatize quantum theory, notably by Mackey Mack63 and Jauch and Piron JauPir63 (see Ref. coeckerev for a review on the more recent progresses of quantum logics). In general, a certain degree of technicality, mainly related to the emphasis on infinite-dimensional systems, makes these results far from providing a clear-cut description of quantum theory in terms of fundamental principles. Later Ludwig initiated an axiomatization program Lud83 adopting an operational approach, where the basic notions are those of preparation devices and measuring devices and the postulates specify how preparations and measurements combine to give the probabilities of experimental outcomes. However, despite the original intent, Ludwig’s axiomatization did not succeed in deriving Hilbert spaces from purely operational notions, as some of the postulates still contained mathematical notions with no operational interpretation.
More recently, the rise of quantum information science moved the emphasis from logics to information processing. The new field clearly showed that the mathematical principles of quantum theory imply an enormous amount of information-theoretic consequences, such as the no-cloning theorem wootterszurek; dieks, the possibility of teleportation tele, secure key distribution wiesner; bb84; e91, or of factoring numbers in polynomial time shor. The natural question is whether the implication can be reversed: is it possible to retrieve quantum theory from a set of purely informational principles? Another contribution of quantum information has been to shift the emphasis to finite dimensional systems, which allow for a simpler treatment but still possess all the remarkable quantum features. In a sense, the study of finite dimensional systems allows one to decouple the conceptual difficulties in our understanding of quantum theory from the technical difficulties of infinite dimensional systems.
In this scenario, Hardy’s 2001 work Har01 re-opened the debate about the axiomatizations of quantum theory with fresh ideas. Hardy’s proposal was based on five main assumptions about the relation between dimension of the state space and the number of perfectly distinguishable states of a given system, about the structure of composite systems, and about the possibility of connecting any two pure states of a physical system through a continuous path of reversible transformations. However, some of these assumptions directly refer to the mathematical properties of the state space (in particular, the “Simplicity Axiom” 2, which is an abstract statement about the functional dependence of the state space dimension on the number of perfectly distinguishable states). Very recently, building on Hardy’s work there have been two new attempts of axiomatization by Dakic and Brukner DakBru09 and Masanes and Müller Mas10. Although these works succeeded in removing the “Simplicity Axiom”, they still contain mathematical assumptions that cannot be understood in elementary physical terms (see e.g. requirement 5 of Ref. Mas10, which assumes that “all mathematically well-defined measurements are allowed by the theory”).
Another approach to the axiomatization of quantum theory was pursued by one of the authors in a series of works maurofirst culminated in Ref. maurolast. These works tackled the problem using operational principles related to tomography and calibration of physical devices, experimental complexity, and to the composition of elementary transformations. In particular this research introduced the concept of dynamically faithful states, namely states that can be used for the complete tomography of physical processes. Although this approach went very close to deriving quantum theory also in this case one mathematical assumption without operational interpretation was needed (see the CJ postulate of Ref. maurolast).
In this paper we provide a complete derivation of finite dimensional quantum theory based of purely operational principles. Our principles do not refer to abstract properties of the mathematical structures that we use to represent states, transformations or measurements, but only to the way in which states, transformations and measurements combine with each other. More specifically, our principles are of informational nature: they assert basic properties of information-processing, such as the possibility or impossibility to carry out certain tasks by manipulating physical systems. In this approach the rules by which information can be processed determine the physical theory, in accordance with Wheeler’s program “it from bit”, for which he argued that “all things physical are information-theoretic in origin” wheeler. Note that, however, our axiomatization of quantum theory is relevant, as a rigorous result, also for those who do not share Wheeler’s ideas on the informational origin of physics. In particular, in the process of deriving quantum theory we provide alternative proofs for many key features of the Hilbert space formalism, such as the spectral decomposition of self-adjoint operators or the existence of projections. The interesting feature of these proofs is that they are obtained by manipulation of the principles, without assuming Hilbert spaces form the start.
The main message of our work is simple: within a standard class of theories of information processing, quantum theory is uniquely identified by a single postulate: purification. The purification postulate, introduced in Ref. purification, expresses a distinctive feature of quantum theory, namely that the ignorance about a part is always compatible with the maximal knowledge of the whole. The key role of this feature was noticed already in 1935 by Schrödinger in his discussion about entanglement Schr35, of which he famously wrote “I would not call that one but rather the characteristic trait of quantum mechanics, the one that enforces its entire departure from classical lines of thought”. In a sense, our work can be viewed as the concrete realization of Schrödinger’s claim: the fact that every physical state can be viewed as the marginal of some pure state of a compound system is indeed the key to single out quantum theory within a standard set of possible theories. It is worth stressing, however, that the purification principle assumed in this paper includes a requirement that was not explicitly mentioned in Schrödinger’s discussion: if two pure states of a composite system have the same marginal on system , then they are connected by some reversible transformation on system . In other words, we assume that all purifications of a given mixed state are equivalent under local reversible operations anchese….
The purification principle expresses a law of conservation of information, stating that at least in principle, irreversibility can always be reduced to the lack of control over an environment. More precisely, the purification principle is equivalent to the statement that every irreversible process can be simulated in an essentially unique way by a reversible interaction of the system with an environment, which is initially in a pure state purification. This statement can also be extended to include the case of measurement processes, and in that case it implies the possibility of arbitrarily shifting the cut between the observer and the observed system purification. The possibility of such a shift was considered by von Neumann as a “fundamental requirement of the scientific viewpoint” (see p. 418 of von32) and his discussion of the measurement process was exactly aimed to show that quantum theory fulfils this requirement.
Besides Schrödinger’s discussion on entanglement and von Neumann’s discussion of the measurement process, the purification principle is deeply rooted in the structure of quantum theory. At the purely mathematical level, it plays a crucial role in the theory of C*-algebras of operators on separable Hilbert spaces, where the purification principle is equivalent to the Gelfand-Naimark-Segal (GNS) construction arveson and implies the celebrated Stinespring’s theorem stine. On the other hand, purification is a cornerstone of quantum information, lying at the origin of most quantum protocols. As it was shown in Ref. purification, the purification principle directly implies crucial features like no-cloning, teleportation, no-information without disturbance, error correction, the impossibility of bit commitment, and the “no-programming” theorem of Ref. no-prog.
In addition to the purification postulate, our derivation of quantum theory is based on five informational axioms. The reason why we call them “axioms”, as opposed to the the purification “postulate”, is that they are not at all specific of quantum theory. These axioms represent standard features of information-processing that everyone would, more or less implicitly, assume. They define a class of theories of information-processing that includes, for example, classical information theory, quantum information theory, and quantum theory with superselection rules. The question whether there are other theories satisfying our five axioms and, in case of a positive answer, the full classification of these theories is currently an open problem.
Here we informally illustrate the five axioms, leaving the more detailed description to the remaining part of the paper:
Causality: the probability of a measurement outcome at a certain time does not depend on the choice of measurements that will be performed later.
Perfect distinguishability: if a state is not completely mixed (i.e. if it cannot be obtained as a mixture from any other state), then there exists at least one state that can be perfectly distinguished from it.
Ideal compression: every source of information can be encoded in a suitable physical system in a lossless and maximally efficient fashion. Here lossless means that the information can be decoded without errors and maximally efficient means that every state of the encoding system represents a state in the information source.
Local distinguishability: if two states of a composite system are different, then we can distinguish between them from the statistics of local measurements on the component systems.
Pure conditioning: if a pure state of system undergoes an atomic measurement on system , then each outcome of the measurement induces a pure state on system . (Here atomic measurement means a measurement that cannot be obtained as a coarse-graining of another measurement).
All these axioms are satisfied by classical information theory. Axiom 5 is even trivial for classical theory, because the only pure states of a composite system are the product of pure states of the component systems and , and hence the state of system will be pure irrespectively of what we do on system .
A stronger version of axiom 5, introduced in Ref. maurolast, is the following:
Atomicity of composition: the sequential composition of two atomic operations is atomic. (Here atomic transformation means a transformation that cannot be obtained from coarse-graining).
However, it turns out that Axiom 5 is enough for our derivation: thanks to the purification postulate we will be able to show the non-trivial implication: Axiom 5 Axiom 5’ (see lemma 16).
The paper is organized as follows. In Sec. II we review the framework of operational-probabilistic theories introduced in Ref. purification. This framework will provide the basic notions needed for the formulation of our principles. In Sec. III we introduce the principles from which we will derive Quantum Theory. In Sec. IV we prove some direct consequences of the principles that will be used later in the paper. In Sec. V we discuss the properties of perfectly distinguishable states, while in Sec. VI we prove the existence of a duality between pure states and atomic effects.
The results about distinguishability and duality of pure states and atomic effects allow us to show in Sec. VII that every system has a well defined informational dimension—the operational counterpart of the Hilbert space dimension. Sec. VIII contains the proof that every state can be decomposed as a convex combination of perfectly distinguishable pure states. Similarly, any element of the vector space spanned by the states can be written as a linear combination of perfectly distinguishable states. This result corresponds to the spectral theorem for self-adjoint operators on complex Hilbert spaces. In Sec. IX we prove some results about the maximum teleportation probability, which allow us to derive a functional relation between the dimension of the state space and the number of perfectly distinguishable states of the system. The mathematical representation of systems with two perfectly distinguishable states is derived in Sec. X, where we prove that such systems are indeed two-dimensional quantum systems—a.k.a. qubits. In Sec. XI we construct projections on the faces of the state space of any system and prove their main properties. These results lead to the derivation of the operational analogue of the superposition principle in Sec. XII which allows to prove that systems with the same number of perfectly distinguishable states are operationally equivalent (Subsec. XII.2). The properties of the projections and the superposition principle are then exploited in Sec. XIII—where we extend the density matrix representation from qubits to higher-dimensional systems, thus proving that a system with perfectly distinguishable states is indeed a quantum system with -dimensional Hilbert space. We conclude the paper with Sec. XIV, where we review our results, discussing future directions for this research.
II The framework
This Section provides a brief summary of the framework of operational-probabilistic theories, which was formulated in Ref. purification. We refer to Ref. purification for an exhaustive presentation of the details of the framework and of the ideas behind it. The operational-probabilistic framework combines the operational language of circuits with the toolbox of probability theory: on the one hand, experiments are described by circuits resulting from the connection of physical devices, on the other hand each device in the circuit can have classical outcomes and the theory provides the probability distribution of outcomes when the devices are connected to form closed circuits (that is, circuits that start with a preparation and end with a measurement).
The notions discussed in this section will allow us to draw a precise distinction between principles with an operational content and exclusively mathematical principles: with the expression ”operational principle” we will mean a principle that can be expressed using only the basic notions of the the operational-probabilistic framework.
A test represents one use of a physical device, like a Stern-Gerlach magnet, a beamsplitter, or a photon counter. The device will have an input system and an output system, labelled by capital letters. The corresponding test can have different classical outcomes, represented by different values of an index :
Each outcome corresponds to a possible event, represented as
We denote by the set of all events from to . The reason for this notation is that in the next subsection the elements of will be interpreted as transformations with input system and output system . If we simply write in place of .
A test with a single outcome will be called deterministic. This name is justified by the fact that, if there is a single possible outcome, then this outcome will occur with certainty (cf. the probabilistic structure introduced in the next subsection).
Two devices can be composed in a sequence, as long as the input system of the second device is equal to the output system of the first. The events in the composite test are represented as
and are written in formulas as .
For every system one can perform the identity-test (or simply, the identity), that is, a test with a single outcome, with the property
The subindex will be dropped from where there is no ambiguity.
The letter will be reserved for the trivial system, which simply means “nothing” nothingquantum. A device with input (resp. output) system is a device with no input (resp. no output). The corresponding tests will be called preparation-tests (resp. observation-tests). In this case we replace the input (resp. output) wire with a round portion:
In formulas we will write (resp. ). The sets and will be denoted as and , respectively. The reason for this special notation is that in the next subsection the elements of (resp. ) will be interpreted as the states (resp. effects) of system .
From every pair of systems and one can form a composite system, denoted by . Clearly, composing system with nothing still gives system , in formula . Two devices can be composed in parallel, thus obtaining a new device with composite input and composite output systems. The events in composite test are represented as
and are written in formulas as . In the special case of states we will often write in place of . Similarly, for effects we will write in place of .
Sequential and parallel composition commute: one has for every such that the output of (resp. ) coincides with the input of (resp. ).
When one of the two tests is the identity, we will omit the box and draw only a straight line, as in
The rules summarized in this section define the operational language of circuits, which has been discussed in detail in a series of inspiring works by Coecke (see in particular Refs. kinder; picturalism). The language of circuits allows one to represent the schematic of an experiment, like e.g.
and also to represent a particular outcome of the experiment
In formula, the above circuit is given by
II.2 Probabilistic structure: states, effects and transformations
On top of the language of circuits, we put a probabilistic structure purification: we declare that the composition of a preparation-test with an observation-test gives rise to a joint probability distribution:
with and . In formula we write . Moreover, if two experiments are run in parallel, we assume that the joint probability distribution is given by the product:
where .
II.3 Basic definitions in the operational-probabilistic framework
Here we summarize few elementary definitions that will be used later in the paper. The meaning of the definitions in the case of quantum theory is also discussed.
First, we start from the notions of coarse-graining and refinement. Coarse-graining arises when we join together some outcomes of a test: we say that the test is a coarse-graining of the test if there is a disjoint partition of such that
Conversely, if is a coarse-graining of , we say that is a refinement of . Intuitively, a test that refines another is a test that extracts information in a more precise way: it is a test with better “resolving power”.
The notion of refinement also applies to a single transformation: a refinement of the transformation is given by a test and a subset such that
Accordingly, we say that each transformation is a refinement of . A transformation is atomic if it has only trivial refinements: if refines , then for some probability . A test that consists of atomic transformations is a test whose “resolving power” cannot be further improved.
When discussing states (i.e. transformations with trivial input) we will use the word pure as a synonym of atomic. A pure state describes a situation of maximal knowledge about the system’s preparation, a knowledge that cannot be further refined.
As usual, a state that is not pure will be called mixed. An important notion is that of completely mixed state:
A state is completely mixed if any other state can refine it: precisely, is completely mixed if for every there is a non-zero probability such that is a refinement of .
Intuitively, a completely mixed state describes a situation of complete ignorance about the system’s preparation: if a system is described by a completely mixed state, then it means that we know so little about its preparation that, in fact, every preparation is possible.
We conclude this paragraph with a couple of definitions that will be used throughout the paper:
A transformation is reversible if there exists another transformation such that and . When the reversible transformations form a group, indicated as .
Two systems and are operationally equivalent if there exists a reversible transformation from to .
When two systems are operationally equivalent one can convert one into the other in a reversible fashion.
II.3.2 Examples in in quantum theory
Preparation-tests are often called quantum information sources in quantum information theory. A generic state is an unnormalized density matrix. A deterministic state, corresponding to a single-outcome preparation-test is a normalized density matrix , with .
Diagonalizing we then obtain that each matrix is a refinement of . More generally, every matrix such that is a refinement of . Up to a positive rescaling, all matrices with support contained in the support of are refinements of . A quantum state is atomic (pure) if and only if it is proportional to a rank-one projection. A quantum state is completely mixed if and only if its density matrix has full rank. Note that the quantum state , where is the identity matrix, is a particular example of completely mixed state, but not the only example. Precisely, is the unique unitarily invariant state in dimension .
Let us now consider the case of observation-tests: in quantum theory an observation-test is given by a POVM (positive operator-valued measure), namely by a collection of non-negative matrices such that
An effect is then a non-negative matrix upper bounded by the identity. In quantum theory there is only one deterministic effect, corresponding to a single-outcome observation test: the unique deterministic effect given by the identity matrix. As we will see in the following section, the fact that the deterministic effect is unique is equivalent to the fact that quantum theory is a causal theory.
An effect is atomic if and only if is proportional to a rank-one projector. An observation-test is atomic if it is a POVM with rank-one elements.
is trace-preserving. A general transformation is then given by a trace non-increasing map, called quantum operation, whereas a deterministic transformation, corresponding to a single-outcome test, is given by a trace-preserving map, called quantum channel.
Any quantum operation can be written in the Kraus form , where are the Kraus operators. Up to a positive scaling, every quantum operation such that the Kraus operators of belong to the linear span of the Kraus operators of is a refinement of . A map is atomic if and only if there is only one Kraus operator in its Kraus form. A reversible transformation in quantum theory is a unitary map , where is a unitary operator, that is and where is the identity operator on (). Two quantum systems are operationally equivalent if and only if the corresponding Hilbert spaces have the same dimension.
II.4 Operational principles
We are now in position to make precise the usage of the expression “operational principle” in the context of this paper. By “operational principle” we mean here a principle that can be stated using only the operational-probablistic language, i.e. using only
the notions of system, test, outcome, probability, state, effect, transformation
their specifications: atomic, pure, mixed, completely mixed
more complex notions constructed from the above terms (e.g. the notion of “reversible transformation”).
The distinction between operational principles and principles referring to abstract mathematical properties, mentioned in the introduction, should now be clear: for example, a statement like “the pure states of a system cannot be cloned” is a valid operational principle, because it can be analyzed in basic operational-probabilistic terms as “for every system there exists no transformation with input system and output system such that for every pure state of ”. On the contrary, a statement like “the state space of a system with two perfectly distinguishable states is a three-dimensional sphere” is not a valid operational principle, because there is no way to express what it means for a state space to be a three-dimensional sphere in terms of basic operational notions. The fact that a state spate is a sphere may be eventually derived from operational principles, but cannot be assumed as a starting point.
III The principles
We now state the principles used in our derivation. The first five principles express generic features that are shared by both classical and quantum theory. They could be even included in the definition of the background framework: they define the simple model of information processing in which we try to single out quantum theory. For this reason we will call them axioms. The sixth principle in our derivation has a different status: it expresses the genuinely quantum features. A major message of our work is that, within a broad class of theories of information processing, quantum theory is completely described by the purification principle. To emphasize the special role of the sixth principle we will call it postulate, in analogy with the parallel postulate of Euclidean geometry.
The first axiom of our list, causality purification, is so basic that could be considered as part of the background framework. We decided to explicitly present it as an axiom for two reasons: The first reason is that the framework of operational-probabilistic theories can be developed even without this requirement (see Ref.purification for the general framework and Refs. supermaps; comblong for two explicit examples of non-causal theories). The second reason is that we want to stress that causality is an essential ingredient in our derivation. This observation is important in view of possible extensions of quantum theory to quantum gravity scenarios where the causal structure is not defined from the start (see e.g. Hardy in Ref. causaloid).
The probability of preparations is independent of the choice of observations.
In technical terms: if is a preparation-test, then the conditional probability of the preparation given the choice of the observation-test is the marginal
The axiom states that the marginal probability is independent of the choice of the observation-test : if and are two different observation-tests, then one has . Loosely speaking, one may refer to causality as a requirement of no-signalling from the future: indeed, causality is equivalent to the fact that the probability of an outcome at a certain time does not depend on the choice of operations that will be done at later times maurolast.
An operational-probabilistic theory that satisfies the causality axiom 1 will be called causal. As we already mentioned, causality is a very basic requirement and could be considered as part of the framework: it provides the notions used to state the other axioms and it implies several facts that will be used frequently in the paper. In fact, in our derivation we do not use the causality axiom directly, but only through its consequences. In the following we briefly summarize the facts and the notations that characterize the framework of causal operational-probabilistic theories, introduced and discussed in detail in Ref. purification. Similar structures have been subsequently considered in Refs. duotenz; caucat within a formal description of circuits in foliable spacetime regions.
First, causality is equivalent to the existence of an effect such that for every observation-test . We call the effect the deterministic effect for system . By definition, the effect is unique. The subindex in will be dropped when no confusion can arise.
In a causal theory every test satisfies the condition
As a consequence, a transformation satisfies the condition
with the equality if and only if is a channel (i.e. a deterministic transformation, corresponding to a single-outcome test). In Eq. (4) we used the notation to mean for every .
In a causal theory the norm of a state is given by . Accordingly, one can define the normalized state
In a causal theory one can always allow for rescaled preparations: conditionally to the outcome in the preparation-test we can say that we prepared the normalized state . For this reason, every state in a causal theory is proportional to a normalized state.
The set of normalized states will be denoted by . Since the set of all states is closed in the operational norm, also the set of normalized states is closed. Moreover, the set is convex purification: this means that for every pair of normalized states and for every probability the convex combination is a normalized state. Operationally, the state is obtained by
performing a binary test with outcomes and outcome probabilities and
for outcome preparing , thus realizing the preparation-test
coarse-graining over the outcomes, thus obtaining .
The step 2 (preparation of a state conditionally on the outcome of a previous test) is possible because the theory is causal purification.
The pure normalized states are the extreme points of the convex set . For a normalized state we define the face identified by as follows:
The face identified by is the set of all normalized states such that , for some non-zero probability and some normalized state .
In other words, is the set of all normalized states that show up in the convex decompositions of . Clearly, if is a pure state, then one has . The opposite situation is that of completely mixed states: by definition 1, a state is completely mixed if every state can stay in its convex decomposition, that is, if . An equivalent condition for a state to be completely mixed is the following:
Proof. The condition is clearly necessary. It is also sufficient because for a state the relation implies (see Lemma 16 of Ref.purification).
A completely mixed state can never be distinguished from another state with zero error probability:
Let be a completely mixed state and be an arbitrary state. Then, the probability of error in distinguishing from is strictly greater than zero.
Proof. By contradiction, suppose that one can distinguish between and with zero error probability. This means that there exists a binary test such that . Since is completely mixed there exists a probability and a state such that . Hence, the condition implies . Therefore, we have . This is in contradiction with the normalization of the probabilities in the test , which would require .
III.1.2 Perfect distinguishability
Our second axiom regards the task of state discrimination. As we saw in proposition 1, if a state is completely mixed, then it is impossible to distinguish it perfectly from any other state. Axiom 2 states the converse:
Every state that is not completely mixed can be perfectly distinguished from some other state.
Note that the statement of axiom 2 holds for quantum and for classical information theory. In quantum theory a completely mixed state is a density matrix with full rank. If a density matrix has not full rank, then it must have a kernel: hence, every density matrix with support in the kernel of will be perfectly distinguishable from , as stated in Axiom 2. Applying the same reasoning for density matrices that are diagonal in a given basis, one can easily see that Axiom 2 is satisfied also by classical information theory.
III.1.3 Ideal compression
The third axiom is about information compression. An information source for system is a preparation-test , where each is an unnormalized state and . A compression scheme is given by an encoding operation from to a smaller system , that is, to a system such that . The compression scheme is lossless for the source if there exists a decoding operation from to such that for every value of the index . This means that the decoding allows one to perfectly retrieve the states . We say that a compression scheme is lossless for the state , if it is lossless for every source such that . Equivalently, this means that the restriction of to the face identified by is equal to the identity channel: for every .
A lossless compression scheme is maximally efficient if the encoding system has the smallest possible size, that is, if the system has no more states than exactly those needed to compress . This happens when every normalized state comes from the encoding of some normalized state , namely .
We say that a compression scheme that is lossless and maximally efficient is ideal. Our second axiom states that ideal compression is always possible:
For every state there exists an ideal compression scheme.
It is easy to see that this statement holds in quantum theory and in classical probability theory. For example, if is a density matrix on a -dimensional Hilbert space and , then the ideal compression is obtained by just encoding in an -dimensional Hilbert space. As long as we do not tolerate losses, this is the most efficient one-shot compression we can devise in quantum theory. Similar observations hold for classical information theory.
III.1.4 Local distinguishability
The fourth axiom consists in the assumption of local distinguishability, here presented in the formulation of Ref. purification.
(Local distinguishability) If two bipartite states are different, then they give different probabilities for at least one product experiment.
In more technical terms: if are states and , then there are two effects and such that
Local distinguishability is equivalent to the fact that two distant parties, holding systems and , respectively, can distinguish between the two states using only local operations and classical communication and achieving an error probability strictly larger than , the probability of error in random guess purification. Again, this statement holds in ordinary quantum theory (on complex Hilbert spaces) and in classical information theory.
Another equivalent condition to local distinguishability is the local tomography axiom, introduced in Refs. maurofirst; barrett. The local tomography axiom state that every bipartite state can be reconstructed from the statistics of local measurements on the component systems. Technically, local tomography is in turn equivalent to the relation Har01 and to the fact that every state can be written as
An important consequence of local distinguishability, observed in Ref. purification, is that a transformation is completely specified by its action on : thanks to local distinguishability we have the implication
(see Lemma 14 of Ref.purification for the proof). Note that Eq. (5) does not hold for quantum theory on real Hilbert spaces purification.
III.1.5 Pure conditioning
The fourth axiom states how the outcomes of a measurement on one side of a pure bipartite state can induce pure states on the other side. In this case we consider atomic measurements, that is, measurements described by observation-tests where each effect is atomic. Intuitively, atomic measurement are those with maximum “resolving power”.
If a bipartite system is in a pure state, then each outcome of an atomic measurement on one side induces a pure state on the other.
The pure conditioning property holds in quantum theory and in classical information theory as well. In fact, the statement is trivial in classical information theory, because the only pure bipartite states are the product of pure states: no matter which measurement is performed on one side, the remaining state on the other side will necessarily be pure.
The pure conditioning property, as formulated above, has been recently introduced in Ref. pcond. A stronger version of axiom 5 is the atomicity of composition introduced in Ref. maurolast:
Atomicity of composition: the sequential composition of two atomic operations is atomic.
Since pure states and atomic effects are a particular case of atomic transformations, Axiom 5’ implies Axiom 5. In our derivation, however, also the converse implication holds: indeed, thanks to the purification postulate we will be able to show that Axiom 5 implies Axiom 5’ (see lemma 16).
III.2 The purification postulate
The last postulate in our list is the purification postulate, which was introduced and explored in detail in Ref. purification. While the previous axioms were also satisfied by classical probability theory, the purification axiom introduces in our derivation the genuinely quantum features. A purification of the state is a pure state of some composite system , with the property that is the marginal of , that is,
Here we refer to the system as the purifying system. The purification axiom states that every state can be obtained as the marginal of a pure bipartite state in an essentially unique way:
Every state has a purification. For fixed purifying system, every two purifications of the same state are connected by a reversible transformation on the purifying system.
Informally speaking, our postulate states that the ignorance about a part is always compatible with a maximal knowledge of the whole. The existence of pure bipartite states with mixed marginal was already recognized by Schrödinger as the characteristic trait of quantum theory Schr35. Here, however, we also emphasize the importance of the uniqueness of purification up to reversible transformations: this property sets up a relation between pure states and reversible transformations that generates most of the structure of quantum theory. As shown in Ref. purification, an impressive number of quantum features are actually direct consequences of purification. In particular, purification implies the possibility of simulating any irreversible process through a reversible interaction of the system with an environment that is finally discarded.
IV First consequences of the principles
Let be a state and let (resp. ) be its encoding (resp. decoding) in the ideal compression scheme of Axiom 3.
Essentially, the encoding operation identifies the face with the state space . In the following we provide a list of elementary lemmas showing that all statements about can be translated into statements about and vice-versa.
The composition of decoding and encoding is the identity on , namely .
Proof. Since the compression is maximally efficient, for every state there is a state such that . Using the fact that (the compression is lossless) we then obtain . By local distinguishability [see Eq. (5)], this implies .
The image of under the decoding operation is .
Proof. Since the compression is maximally efficient, for all there exists such that . Then, . This implies that . On the other hand, since the compression is lossless, for every state one has . This implies the inclusion .
If the state is pure, then the state is pure. If the state is pure, then the state is pure.
Proof. Suppose that is pure and that can be written as for some and some . Applying on both sides we obtain . Since is pure we must have . Now, applying on all terms of the equality and using lemma 2 we obtain . This proves that is pure. Conversely, suppose that is pure and for some and some . Since is in the face (lemma 3), also and are in the same face. Applying on both sides of the equality and using lemma 2 we obtain . Since is pure we must have . Applying on all terms of the equality we then have , thus proving that is pure.
We say that a state is completely mixed relative to the face if every state can stay in the convex decomposition of . In other words, is completely mixed relative to if one has . Note that in general implies .
If the state is completely mixed relative to , then the state is completely mixed. If the state is completely mixed, then the state is completely mixed relative to .
Proof. Suppose that is completely mixed relative to . Then every state can stay in its convex decomposition, say with and . Applying we have
Since the compression is maximally efficient, for every state there exists a state such that . Choosing the suitable and substituting to in Eq. (6) we then obtain that for every state there exists probability and a state such that
This implies that is completely mixed. Suppose now that is completely mixed. Then every state can stay in its convex decomposition, say . with and . Applying on both sides we have
Now, using lemma 3 we have that every state can be written as for some . Choosing the suitable and substituting to in Eq. (7) we then obtain that for evert state there exists a probability and a state such that . Therefore, is completely mixed relative to .
We now show that the system used for ideal compression of the state is unique up to operational equivalence:
If two systems and allow for ideal compression of a state , then and are operationally equivalent.
Proof. Let and denote the encoding/decoding schemes for systems and , respectively. Define the transformations and . It is easy to see that is reversible and . Indeed, since the restriction of and to the face is the identity, using Lemma 3 one has and similarly . Hence, we have and .
It is useful to introduce the notion of equality upon input of . We say that two transformations are equal upon input of if their restrictions to the face identified by are equal, that is, if for every . If and are equal upon input of we write .
Using the notion of equality upon input of we can rephrase the fact that the compression is lossless for as . Similarly, we can state the following:
The encoding is deterministic upon input of , that is .
Proof. For every we have , having used Eq. (4) and the fact that the compression is lossless. Since probabilities are bounded by 1, this implies for every , that is, .
The decoding is deterministic, that is .
Proof. For every we have , having used Eq. (4) and lemma 2. Hence, .
IV.2 Results about purification
The purification postulate 1 implies a large number of quantum features, as it was shown in Ref. purification. Here we review only the facts that are useful for our derivation, referring to Ref. purification for the proofs.
An elementary consequence of the uniqueness of purification is that the group of reversible transformations on acts transitively on the set of pure states:
For every couple of pure states there is a reversible transformation such that .
Proof. See Lemma 20 of Ref. purification.
Transitivity implies that for every system there is a unique state that is invariant under reversible transformations, that is, a unique state such that for every :
For every system , there is a unique state invariant under all reversible transformations in . The invariant state has the following properties:
.
Proof. See Corollary 34 and Theorem 4 of Ref. purification. The proof of item 2 uses the local distinguishability axiom.
When there is no ambiguity we will drop the subindex and simply write .
The uniqueness of purification in postulate 1 requires that if are two purifications of , then there exists a reversible transformation such that . The following lemma extends the uniqueness property to purifications with different purifying systems:
(Uniqueness of the purification up to channels on the purifying systems) Let and be two purifications of . Then there exists a channel such that
Proof. See Lemma 21 of Ref. purification.
Another consequence of the uniqueness of purification is the fact that any ensemble decomposition of a given mixed state can be obtained by performing a measurement on the purifying system:
(Purification of preparation-tests) Let be a state and be a purification of . If be a preparation-test such that , then there exists an observation-test on the purifying system such that
Proof. See lemma 8 of Ref. purification.
If is a purification of and belongs to the face , then there exists an effect and a non-zero probability such that
An important consequence of purification and local distinguishability is the relation between equality upon input of and equality on the purifications of :
(Equality upon input of vs equality on purifications of ) Let be a purification of , and let be two transformations. Then one has
Proof. See theorem 1 of Ref. purification. The proof of the direction uses the local distinguishability axiom.
As a consequence, the purification of a completely mixed state allows for the tomography of transformations:
Let be completely mixed and is a purification of . Then, for all transformations one has
Proof. By theorem 1 the first condition is equivalent to . Since is completely mixed, this means for every . By local distinguishability [see Eq. (5)] this implies .
Corollary 2 shows that the state characterizes the transformation completely. We will express this fact by saying that the state is dynamically faithful maurolast, or just faithful, for short. Using this notion we can rephrase corollary 2 as:
If is pure and its marginal on system is completely mixed, then is dynamically faithful for system .
Let us choose a fixed faithful state for system , say . Then for every transformation we can define the Choi state as
(Choi isomorphism) For a given faithful state the map has the following properties:
it defines a bijective correspondence between tests from to and collections of states for satisfying
the transformation is atomic if and only if the corresponding state is pure.
Proof. See Theorem 17 of Ref. purification.
A simple consequence of the Choi isomorphism is the following:
Let be a collection of transformations. Then, is a test if and only if
In particular, let be a collection of effects. Then, is an observation-test if and only if
Proof. Apply item 1 of theorem 2 to the collection of states defined by .
A much deeper consequence of the Choi isomorphism is the following theorem:
(States specify the theory) Let be two theories satisfying the purification postulate. If and have the same sets of normalized states, then .
Proof. See Theorem 19 of Ref.purification.
Thanks to theorem 3 to derive quantum theory we will only need to prove that our principles imply that for every system the normalized states can be described as positive Hermitian matrices with unit trace. Once this is proved, theorem 3 automatically ensures that all the dynamics and all the measurements allowed by the theory are exactly the dynamics and the measurements allowed in quantum theory.
Note that in the definition of the Choi state we left the freedom to choose the faithful state . Among many possibilities, one convenient choice is to take a faithful state obtained as a purification of the invariant state . Moreover, as we will see in the next paragraph, we can always choose the purifying system in such a way that the marginal on is completely mixed.
IV.3 Results about the combination of compression and purification
An important consequence of the combination of the purification postulate with the compression axiom is the fact that one can always choose a purification of such that the marginal state on the purifying system is completely mixed. To prove this result we need the following lemma:
Let be a state and let be a purification of . If is the encoding operation in the compression scheme of axiom 3, then the state is pure.
Proof. Let be the decoding operation. Since the compression is lossless for we know that . By theorem 1 this is equivalent to the condition . Now, suppose that . Applying on both sides we then obtain , and, since is pure, for every we must have , where is some probability. Finally, since (lemma 2), one has . Hence, admits only decompositions with , that is, is pure.
We are now in position to prove the desired result:
For every state there exists a system and a purification of such that the marginal state on system is completely mixed. Moreover, the system is unique up to operational equivalence.
Let be a purification of and let be the encoding for . Then, the state is dynamically faithful for .
Proof. The marginal of on system is , which is completely mixed by lemma 5. Hence, is dynamically faithful by corollary 3.
The decoding transformation in the ideal compression for is atomic.
Proof. Let be a purification of , for some purifying system . Since (the compression is lossless), we have (theorem 1). Now, by corollary 5 is faithful for and by lemma 13 is pure. Using the Choi isomorphism with the faithful state we then obtain that is atomic.
IV.4 Teleportation and the link product
Proof. See Corollary 19 of Ref. purification.
Let us choose to be the faithful state in the definition of the Choi isomorphism. Then the sequential composition of transformation induces a composition of Choi states in following way:
For two transformations and the Choi state of is given by the link product
Proof. See Corollary 22 of Ref. purification.
We conclude this paragraph with an important result that follows from the combination of the link product structure with the pure conditioning axiom:
The composition of two atomic transformations is atomic.
Proof. Let and be two atomic transformations. By the Choi isomorphism, the (unnormalized) states and are pure. Since the teleportation effect in Eq. (9) is atomic (lemma 15), the pure conditioning axiom 5 implies the state is pure. By the Choi isomorphism this means that is atomic.
IV.5 No information without disturbance
We say that a test is non-disturbing upon input of if . If is completely mixed, we simply say that the test is non-disturbing.
A consequence of the purification postulate is the following “no-information without disturbance” result:
A test is non-disturbing upon input of if and only if there is a set of probabilities such that for every .
Proof. See Theorem 10 of Ref. purification.
The no-information without disturbance result implies the following geometrical limitation
For every system the convex set of states is not a segment.
Proof. The proof is by contradiction. Suppose that for some system the set is a segment. The segment has only two pure states, say and , and every other state is completely mixed. Then the distinguishability axiom 2 imposes that and are perfectly distinguishable. Take the binary test such that and define the “measure-and-prepare” test as , (the possibility of preparing a state depending on the outcome of a previous measurement is guaranteed by causality purification). Since every state in the segment can be written as convex combination of the two extreme points, we have that the test is non-disturbing: for every . This is in contradiction with lemma 17 because and are not proportional to the identity.
We know that no information can be extracted without disturbance. In the following we will prove a result in the converse direction: if a measurement extracts no information, than it can be realized in a non-disturbing fashion. To show this result we first need the following
For every observation test with finite outcome set there is a system and a test consisting of atomic transformations such that .
Proof. Let be a pure faithful state for system and let the Choi state of . Take a purification of , say for some purifying system singlepurifying. Then, by the Choi isomorphism there is a test , with input and output , such that
[see item 1 of theorem 2] Moreover, each transformation is atomic (item 2 of theorem 2). Applying the deterministic effect on both sides we then obtain . By definition of , this implies , and, since is dynamically faithful, .
Let be a state, be an effect, and be an atomic transformation such that . If for some , then there exists a channel such that .
Proof. Consider a purification of , say , and define the state by . By the atomicity of composition 16 the state is pure. Moreover, we have
having used theorem 1 in the last equality. This implies that and are different purifications of the same mixed state on system . Then, by lemma 11 there exists a channel such that . By theorem 1, the last equality implies .
We now make a simple observation that combined with Theorem 5 will lead to some interesting consequences:
If , then . Similarly, if , then .
Proof. By definition, iff there exists and such that . If , then we have . Since and cannot be larger than , the only way to have the equality is to have . By definition, this amounts to say . Similarly, if , one has , which is satisfied only if , that is, if .
Let be a state, be an effect, and be an atomic transformation such that . If , then is correctable upon input of , that is, there exists a correction operation such that .
Proof. If , then clearly . Lemma 19 then implies . Applying theorem 5 we finally obtain the thesis.
Let be a state, be an effect such that . Then there exists a transformation such that and .
Proof. Straightforward consequence of lemma 18 and of corollary 8.
Finally, we say that an observation-test is non-informative upon input of if we have for every . This means that the test is unable to distinguish the states in the face . As a consequence of theorem 5 we have the following “no disturbance without information” result:
If the test is non-informative upon input of then there is a test that is non-disturbing upon input of and satisfies for every .
Proof. By lemma 18 there exists a test such that each transformation is atomic and . By theorem 5, for each there is a correction channel such that . Defining we then obtain the thesis.
V Perfectly distinguishable states
In this section we prove some basic facts about perfectly distinguishable states. Let us start from the definition:
The normalized states are perfectly distinguishable if there exists an observation-test such that . The observation-test is called perfectly distinguishing.
From the distinguishability axiom 2 it is clear that every nontrivial system has at least two perfectly distinguishable states:
For every nontrivial system there are at least two perfectly distinguishable states.
Proof. Let be a pure state of . Obviously, is not completely mixed (unless the system has only one state, that is, unless is trivial). Hence, by axiom 2 there exists at least a state that is perfectly distinguishable from .
An equivalent condition for perfect distinguishability is the following:
The states are perfectly distinguishable if and only if there exists an observation-test such that for every .
Proof. The condition is clearly necessary. On the other hand, the condition implies
Since all probabilities are non-negative, we must have for , and therefore, .
A very general fact about state discrimination is expressed by the following:
If is perfectly distinguishable from and (resp. ) belongs to the face identified by (resp. ), then is perfectly distinguishable from .
Proof. Let be the binary observation-test that distinguishes perfectly between and . By definition, is such that and . Now, by lemma 19, and for all and .
Thanks to purification and to the local distinguishability axiom 4, we are also in position to show a much stronger result:
Let and be two sets of perfectly distinguishable states. If is perfectly distinguishable from , then the states are perfectly distinguishable.
Proof. Let be the observation-test such that and . Now, by corollary 9 there is a transformation such that and . Similarly, there exists a transformation such that and . We can then define the following observation-test
where (resp. ) is the observation-test that perfectly distinguishes among the states (resp. ). By corollary 4 [see in particular Eq. (8)], is indeed an observation-test: each is an effect and one has the normalization
Moreover, since and , one has for every . By lemma 21, this implies that the states are perfectly distinguishable.
A set of perfectly distinguishable states is maximal if there is no state such that the states are perfectly distinguishable
A set of perfectly distinguishable states is maximal if and only if the state is completely mixed.
Proof. We first prove that if is completely mixed, then the set must be maximal. Indeed, if there existed a state such that are perfectly distinguishable, then clearly would be distinguishable from . This is absurd because by proposition 1 no state can be perfectly distinguished from a completely mixed state. Conversely, if is maximal, then is completely mixed. If it were not, by the distinguishability axiom 2, would be perfectly distinguishable from some state . By lemma 23, this would imply that the states are perfectly distinguishable, in contradiction with the hypothesis that the set is maximal.
Every set of perfectly distinguishable pure states can be extended to a maximal set of perfectly distinguishable pure states.
Any pure state belongs to a maximal set of perfectly distinguishable pure states.
We conclude this section with a few elementary facts about how the ideal compression of axiom 3 preserves the distinguishability properties. In the following we will choose a state and (resp. ) will be the encoding (resp. decoding) in the ideal compression scheme for .
If the states are perfectly distinguishable, then the states are perfectly distinguishable. Conversely, if the states are perfectly distinguishable, then the states are perfectly distinguishable.
Proof. Let be the observation-test such that for every . Since the compression is lossless, we have and . Now, consider the test defined by . Clearly we have for every . By lemma 21 this means that the states are perfectly distinguishable. Similarly, let the observation-test that distinguishes the set . Since (lemma 2), we can conclude by the same argument that the states are perfectly distinguishable.
We say that a set of perfectly distinguishable states is maximal in the face if there is no state such that the states are perfectly distinguishable. We then have the following:
If is a maximal set of perfectly distinguishable states in the face , then is a maximal set of perfectly distinguishable states. Conversely, if is a maximal set of perfectly distinguishable states, then is a maximal set of perfectly distinguishable states in the face .
Proof. Distinguishability of the states and is proved by lemma 25. Let us now prove maximality. By contradiction, suppose that the set is maximal in the face while the set , is not maximal. This means that there exists a state such that the states are perfectly distinguishable. By lemma 25 the states are perfectly distinguishable. Since for every , this means that the states are perfectly distinguishable, in contradiction with the fact that is maximal. This proves that the set must be maximal. Conversely, if the set is maximal, using the same argument we can prove that the set must be maximal in .
VI Duality between pure states and atomic effects
We now show the existence of a one-to-one correspondence between states and effects of any system in the theory. Let us start from a simple observation:
If is atomic and for , then must be pure.
Proof. By lemma 19, the condition implies . By theorem 1, the condition implies
We are now in position to show that every atomic effect is associated to a unique pure state.
For every atomic effect , there exists a unique pure state such that .
Proof. Let be a state such that . By lemma 26, must be pure. Moreover, this pure state must be unique: suppose that and are pure states such that . Then for one has . Since must be pure, one has .
We now show the converse result: for every pure state there exists a unique atomic effect such that . Let us start from the existence:
Let be a maximal set of perfectly distinguishable pure states and let be the observation-test such that . Then each effect is atomic with .
Proof. It is obvious that , because of the condition . It remains to prove atomicity. Consider the state , which is completely mixed by theorem 6. Let be a purification of , chosen in such a way that the marginal on system is completely mixed (theorem 4). As a consequence of purification (lemma 12), there exists an observation-test on system such that . Since is dynamically faithful on system , each effect must be atomic. Now, define the normalized states and the probabilities by
Applying the deterministic effect on both sides one has . On the other hand, applying the effect one has instead . This implies for every . Since is atomic, lemma 26 forces each to be pure. Finally, each must be atomic since its Choi state is pure (theorem 2).
As a consequence, we can prove the following existence result:
For every pure state there exists an atomic effect such that .
Proof. By corollary 11, every pure state belongs to a maximal set of perfectly distinguishable pure states , say . The thesis then follows from lemma 27.
We now prove that the atomic effect such that is unique. For this purpose we need two auxiliary lemmas:
Let be an arbitrary pure state and let be the probability defined by
where is the invariant state of system . Then the value of the probability is independent of .
Proof. Since for every couple of pure states and one has for some reversible channel (lemma 9), and since is invariant, one has if and only if . The maximum probabilities for and are then equal.
Since for every couple of pure states, from now on we will write in place of .
Let be a pure state and be an atomic effect such that . Let be a purification of the invariant state , chosen in such a way that the marginal on system is completely mixed, and let be the unique atomic effect on such that
[note that exists by lemma 12is uniquely defined by Eq. (12) because is faithful for system ]. Then one has
where is the unique pure state such that .
Proof. Define the normalized pure state and the probability by
In order to prove the thesis we have to show that and . Applying on both sides of Eq. (14) and using Eq. (12) we obtain . This implies
with the equality if and only if . Let be an atomic effect such that (such an effect exists because of lemma 28 ). Define the normalized pure state and the probability by
Applying on both sides and using Eq. (14) we obtain , which implies , with the equality if and only if . Combining this with the inequality (15) we have . On the other hand, by Lemma 29 one has , and consequently . This also implies that and .
For every pure state there is a unique atomic effect such that .
Proof. Existence has been already proved in lemma 28. Let us prove uniqueness: suppose that and are two atomic effects such that . Then, applying lemma 30 to and we obtain
Since is dynamically faithful, this implies .
Finally, an important consequence of theorem 8 is
If are two atomic effects with , then there is a reversible channel such that .
Proof. Let and be the (unique) normalized states such that and , respectively. Now, there is a reversible channel such that . Hence, . By theorem 8, one has .
We conclude this section with an elementary result that will be used later in the paper:
Let and be the encoding and the decoding in the ideal compression scheme for . If is a pure state and is the atomic effect such that , then is a pure state and is the atomic effect such that .
Proof. The state is pure by lemma 4. The effect is atomic by lemmas 14 and 16. Since , one has .
VII Dimension
In this section we show that each system in our theory has given informational dimension, defined as the maximum number of perfectly distinguishable pure states available in the system. In the Hilbert space framework, the informational dimension will be the dimension of the Hilbert space.
All maximal sets of perfectly distinguishable pure states have the same number of elements.
Proof. Let be a maximal set of perfectly distinguishable pure states for system , and let the observation-test such that . By lemma 27, each is atomic and . Then, by corollary 13, one has , where each is a reversible channel and is a fixed atomic effect with . By the invariance of we then obtain . On the other hand, one has , which implies . Since is arbitrary, is independent of the choice of the set .
An immediate consequence of the proof of lemma 32 is
For every atomic effect with one has .
This simple fact has two very important consequences. The first is that the dimension of a composite system is the product of the dimensions of the components:
The dimension of the composite system is the product of the dimensions of and , namely .
Proof. From lemma 10 we know that is the unique invariant state of system . Now, if and are such that , then is such that . Hence we have .
The second consequence is the relation between the dimension and the maximum probability of a pure state in the convex decomposition of the invariant state :
For every system , the maximum probability of a pure state in the convex decomposition of the invariant state is .
Proof. Let be a purification of the invariant state , chosen in such a way that the marginal on system is completely mixed. Let be an atomic effect with . Then, equation (13) becomes
where is some normalized pure state of system . Applying the deterministic effect on system on both sides we obtain . Finally, corollary 14 states . By comparison, we obtain .
Thanks to the compression axiom 3, the notion of dimension can be applied not only to the whole state space but also to its faces. With face of the convex set we always mean the face identified by some state .
Let be a face of the convex set . Every maximal set of perfectly distinguishable pure states in has the same cardinality . Precisely, if is the face identified by and is the encoding in the ideal compression for , then we have
Proof. The set is perfectly distinguishable by lemma 25, and it is maximal by corollary 12. Moreover, the states are pure by lemma 4. Hence, the cardinality of the set must be .
From now on the maximum number of perfectly distinguishable states in the face will be called the dimension of the face and will be denoted by .
VIII Decomposition into perfectly distinguishable pure states
In this section we show that in a theory satisfying our principles any state can be written as a convex combination of perfectly distinguishable pure states. In quantum theory, this corresponds to the diagonalization of the density matrix.
To prove this result we need first a sufficient condition for the distinguishability of states, given in the following
Let be a set of states. If there exists a set of effects (not necessarily an observation-test) such that , then the states are perfectly distinguishable.
Proof. For each consider the binary test . Since by hypothesis , the test can perfectly distinguish from any mixture of the states . In particular, this means that, for every , can be perfectly distinguished from the mixture . Note that, by definition, the states belong to the face . We now prove by induction on that the states are perfectly distinguishable. This is true for . Now, suppose that the states are perfectly distinguishable. Since the state is perfectly distinguishable from , by lemma 23 we have that the states are perfectly distinguishable. Taking the thesis follows.
We now show that the invariant state is a mixture of perfectly distinguishable pure states.
For every maximal set of perfectly distinguishable pure states one has
Proof. Let be the observation test such that , and be a purification of , chosen in such a way that the marginal on system is completely mixed (theorem 4). Let be the pure states defined by
and, for each , let be the atomic effect such that
(here we used lemma 30 and the fact that ). Then we have
By lemma 35, this implies that the states are perfectly distinguishable. Now, since the marginal of on system is completely mixed, theorem 6 states that the set is maximal. Let the observation test such that . By lemma 27, each must be atomic. On the other hand, there is a unique atomic effect such that (theorem 8). Therefore, . This means that the effects form an observation test. Once this fact has been proved, using Eq. (16) we obtain
The distance between the invariant state and an arbitrary pure state is
Proof. Take a maximal set of perfectly distinguishable pure states such that (corollary 11). Since one has , where . Hence, one has , having used that and are perfectly distinguishable and therefore (see subsection II-I in Ref. purification).
We can now prove the following strong result:
For every system , every mixed state can be written as a convex combination of perfectly distinguishable pure states.
for some state on the border of , that is, for some state that is not completely mixed. But we know from the discussion of point (1) that the state is a mixture of perfectly distinguishable pure states, say . By lemma 24, this set can be extended to a maximal set of perfectly distinguishable pure states . On the other hand, theorem 9 states that . This implies the desired decomposition
where for , and otherwise.
It is easy to show that the marginals of a pure bipartite state have the same spectral decomposition:
Proof. Let be the observation-test such that , be the observation test such that for every . For , define the pure state and the probability via the relation
[Note that is pure due to the pure conditioning axiom] By definition, we have
The above relation implies and . Hence, the states are perfectly distinguishable. On the other hand, we have which implies , . Therefore, we obtained
which is the desired spectral decomposition.
The spectral decomposition of states has many consequences. Here we just discuss the simplest ones, which are needed for the purpose of the derivation of quantum theory.
A first consequence is the following lemma:
Let be a pure state and let be the unique atomic effect such that . If is perfectly distinguishable from , then .
Proof. Let us write , with perfectly distinguishable pure states and for each . Now, by lemma 23 the states are perfectly distinguishable, and by lemma 24 this set can be extended to a maximal set of perfectly distinguishable pure states , with for and . Denote by the observation test that perfectly distinguishes between the states . Note that, by definition, and for every . Also, recall that is atomic (lemma 27). By the duality of theorem 8 we have , and, therefore, .
Another consequence of theorem 10 is the following characterization of the completely mixed states as full rank states:
(Characterization of completely mixed states) A state , written as a mixture of a maximal set of perfectly distinguishable pure states , is completely mixed if and only if for every .
Proof. Necessity: If for some , then is perfectly distinguishable from . Hence, it cannot be completely mixed. Sufficiency: let . Then we have , where is the state defined by . Since contains in its convex decomposition, and since is completely mixed, we conclude that is completely mixed.
In particular, for two-dimensional systems we have the result:
For any state on the border of is pure.
Proof. Write as , where and and are normalized states. If there is nothing to prove, because is proportional to a state. Then, suppose that . Write as where are perfectly distinguishable and define . Then one has . Now, by definition is proportional to a state: indeed we have , and, by definition . Therefore is proportional to a state, say , with . Writing as , where is a maximal set of perfectly distinguishable pure states, we then obtain , which is the desired decomposition.
In quantum theory, corollary 21 is equivalent to the fact that every Hermitian matrix is diagonal in a suitable orthonormal basis. A simple consequence of corollary 21 is the following
For every system with there is a continuous set of pure states.
We conclude this section with the dual result to the “spectral decomposition” of corollary 21:
Since is dynamically faithful, this implies , where .
IX Teleportation revisited
We start by showing a probabilistic teleportation scheme that achieves success probability for every system :
For every system , probabilistic teleportation can be achieved with probability .
and, since is dynamically faithful,
as can be verified applying both members of Eq. (19) to , thus obtaining Eq. (18).
IX.2 Isotropic states and effects
Let us start from the definition of the transpose:
[note that the transposition is defined with respect to the given state ]
namely .
The conjugate is just defined as the inverse of the transpose:
We can now give the definition of isotropic pure state (isotropic atomic effect):
An example of isotropic state is : indeed, by definition of conjugate we have, for every ,
As a consequence, the teleportation effect is isotropic: indeed one has
which implies , since the state is dynamically faithful.
We now show that all isotropic pure states (isotropic atomic effects) are connected to the state (to the effect ) through a local reversible transformation.
Since is dynamically faithful, the above equation implies for every .
By the duality between states and effects, it is easy to obtain the following:
As a consequence, every isotropic effect is connected to the teleportation effect by a local reversible transformation:
Proof. Since and are both isotropic, lemma 39 implies that they are both connected to through a local reversible transformation, say and , respectively. Therefore, they are connected to each other through the transformation .
IX.3 Dimension of the state space
Due to local distinguishability, any bipartite state can be written as
with and . Finally, a transformation from to can be written as
In this matrix representation, the teleportation diagram of Eq. (19) becomes
where is the identity matrix in dimension . On the other hand, we also have
We now show that one has the equality, using the following standard lemma:
where is an orthogonal matrix.
where is the Haar measure on the compact group (see corollary 30 of Ref. purification for the proof of compactness) and denotes the transpose of . By definition, one has and for every . Let us now define the new representation
obtained from by a change of basis in the subspace spanned by . With this choice, each matrix is orthogonal:
having used Eq. (23) for the last equality. Using the inequality , that holds for every orthogonal matrix, we then obtain
ans, therefore
The dimension of the vector space generated by the states in is .
Proof. Using lemma 41 and Eq. (23) we obtain . Hence,
An interesting consequence of the relation is the following
Let us write an arbitrary state as , with . Then, the linear map defined by is not a physical transformation.
Since this quantity is negative for every , the map cannot be a physical transformation.
cannot represent a physical transformation of system .
X Derivation of the qubit
The first step is to prove that the set of normalized states is a sphere. The idea of the proof is a simple geometric observation: in the ordinary three-dimensional space the sphere is the only compact convex set that has an infinite number of pure states connected by orthogonal transformations. The complete proof is given in the following
Since the convex set of density matrices on a two-dimensional Hilbert space is a sphere, we can represent the states in as density matrices. Precisely, we can choose three orthogonal axes passing through the center of the sphere and call them axes, take , to be the two perfectly distinguishable pure states in the direction of the -axis and define . From the geometry of the sphere we know that any state can be written as
where the pure states are those for which . The Bloch representation of quantum state is then obtained by associating the basis vectors to the matrices
Proof. Clearly the matrix must be positive for every effect , since we have for every density matrix . Moreover, since we have for every density matrix , we must have . Finally, we know that for every couple of perfectly distinguishable pure states there exists an atomic effect such that and . Since the two pure states are represented by orthogonal rank-one projectors and , we must have . This proves that the atomic effects are the whole set of positive rank-one projectors. As a consequence, also every positive matrix with must represent some effect .
Note that we proved that all two-dimensional systems and in our theory have the same states (), the same effects (), and the same reversible transformations (), but we did not show that and are operationally equivalent. For example, and could be different when we compose them with a third system : the set of states and could be non-isomorphic. The fact that every couple of two-dimensional systems and are operationally equivalent will be proved later (cf. corollary 40).
We conclude this section with a simple fact that will be very useful later:
(Superposition principle for qubits) Let be two perfectly distinguishable pure states of a system with . Let be the observation-test such that . Then, for every probability there exists a pure state such that
Precisely, the set of pure states satisfying Eq. (28) is a circle in the Bloch sphere.
Proof. Elementary property of density matrices.
XI Projections
In this section we define the projection on a face of the convex set and we prove several properties of projections. The projection on the face will be defined as an atomic operation that acts as the identity on states in the face and that annihilates the states on the orthogonal face . In the following we first introduce the concept of orthogonal face, then prove the existence and uniqueness of projections, and finally give some useful results on the projection of a pure state on two orthogonal faces.
In order to introduce the notion of orthogonal face we need first a few elementary results. We start by showing that there is a canonical way to associate a state to a face :
Let be a face of the convex set and let be a maximal set of perfectly distinguishable pure states in . Then the state depends only on the face and not on the particular set . Morever, is the face identified by
Proof. Suppose that is the face identified by and let (resp. ) be the encoding (resp. decoding) in the ideal compression for . By lemma 4 and corollary 12, is a maximal set of perfectly distinguishable pure states of and by theorem 9 one has . Hence, . Since the right-hand side of the equality is independent of the particular set , the state in the left-hand side is independent too. To prove that is the face identified by it is enough to observe that is completely mixed relative to : this fact follows from the relation and from lemma 5
We now define the orthogonal complement of the state :
The orthogonal complement of the state is the state defined as follows:
if , then
if , then is defined by the relation
An easy way to write the orthogonal complement is
Take a maximal set of perfectly distinguishable pure states in and extend it to a maximal set of perfectly distinguishable pure states in , then for we have
Proof. By definition, for we have . Substituting the expressions and we then obtain the thesis.
Note, however, that by definition the orthogonal complement depends only on the face and not on the choice of the maximal set in lemma 43.
The states and are perfectly distinguishable.
Proof. Take a maximal set of perfectly distinguishable pure states in , extend it to a maximal set , and take the observation-test such that . Then the binary test , defined by distinguishes perfectly between and .
We say that a state is perfectly distinguishable from the face if is perfectly distinguishable from every state in the face . With this definition we have the following
is perfectly distinguishable from the face
is perfectly distinguishable from
belongs to the face identified by , i.e. .
Proof. () is perfectly distinguishable from if and only if then there exists a binary test such that and . By lemma 19 this is equivalent to the condition and , that is, is distinguishable from any state in the face identified by , which by definition is . () Let be a maximal set of perfectly distinguishable states in , , and let be the maximal set of perfectly distinguishable pure states in the spectral decomposition , with for every . Since is perfectly distinguishable from , by lemma 23 we have that the states are all perfectly distinguishable. Let us extend this set to a maximal set . By lemma 43 have . Hence, all the states are in the face . Since is a mixture of these states, it also belongs to the face . () Since and are perfectly distinguishable, if belongs to the face identified by , then by lemma 22 is perfectly distinguishable from .
If is perfectly distinguishable from and from , then is perfectly distinguishable from any convex mixture of and .
Proof. Let be the face identified by . Then by lemma 44 we have . Since is a convex set, any mixture of and belongs to it. By lemma 44, this means that any mixture of and is perfectly distinguishable from .
We are now ready to give the definition of orthogonal face:
The orthogonal face is the set of all states that are perfectly distinguishable from the face .
By lemma 44 it is clear that is the face identified by , that is .
In the following we list few elementary facts about orthogonal faces:
Proof. Item 1. If the thesis is obvious. If , take a maximal set (resp. ) of perfectly distinguishable pure states in (resp. ). Hence we have
By corollary 32 the states and are perfectly distinguishable. Hence, the states are perfectly distinguishable jointly (lemma 23). Now, we must have , otherwise there would be a pure state that is perfectly distinguishable from the states . This implies that belongs to and that states are perfectly distinguishable in , in contradiction with the hypotheses that the set is maximal in . Item 2 Immediate from item 1 and definition 9. Item 3 and 4 Both items follow by comparison of item 2 with Eq. 29. Item 5 By condition 3 of lemma 44, is the face identified by the state , which, by item 4, is . Since the face identified by is , we have .
We now show that there is a canonical way to associate an effect to a face :
We say that is the effect associated to the face if and only if and .
In other words, the definition imposes that for every and for every .
A state belongs to the face if and only if .
Proof. By definition, if belongs to , then . Conversely, if , then is perfectly distinguishable from , because . Now, we know that is equal to (item 4 of lemma 45). By item 2 of lemma 44 the fact that is perfectly distinguishable from implies that belongs to , which is just (item 5 of lemma 45).
We now show that the effect associated to the face exists and is unique. A preliminary result needed to this purpose is the following:
The effect must have the form , where is the atomic effect such that and is a maximal set of perfectly distinguishable pure states in .
Proof. By corollary 23 we have that can be written as where is a perfectly distinguishing test. Moreover, since is an effect, we must have forall . Now, by definition we have , which implies for every , that is, whenever . Let us focus on the values of for which . Let be the pure state such that . The condition implies that is perfectly distinguishable from . Therefore, belongs to , which is . Since by definition we must have , this also implies that . In summary, we proved that where the prime means that the sum is restricted to those values of such that . The condition also implies that the number of terms in the sum must be exactly . The thesis is then proved by suitably relabelling the effects , in such a way that belongs to for every .
The effect associated to the face is unique.
Proof. Suppose that and are two effects associated to the face , both written as in lemma 47. Let (resp. ) be the maximal set of perfectly distinguishable pure states in such that for every (resp. for every ), and let be a maximal set of perfectly distinguishable pure states in . Since and are perfectly distinguishable, the states (resp. are perfectly distinguishable (lemma 23). Moreover, the set is maximal since . Let be the atomic effect such that . Then, the test that distinguishes the states (resp. is given by (resp. and its normalization reads
By comparison we obtain .
XI.2 Projections
We are now in position to define the projection on a face:
Let be a face of . A projection on the face is an atomic transformation such that
When is the face identified by a pure state , we have and call a projection on the pure state .
The first condition in definition 12 means that the projection does not disturb the states in the face . The second condition means that annihilates all states in the orthogonal face . As a notation, we will indicate with the projection on the face , that is, we will use the definition .
An equivalent condition for to be a projection on the face is the following:
Let be a maximal set of perfectly distinguishable pure states for system . The transformation in is a projection on the face generated by the subset if and only if
Proof. The condition is clearly necessary, since by Definition 12 for . On the other hand, if for then by definition of we have , and, therefore .
Proof. is atomic, being the product of two atomic transformations. We now show that : Indeed, by the local tomography axiom it is easy to see that every state can be written as , where is a basis for and is a basis for . Since , we have
In the following we will show that for every face there exists a unique projection and we will prove several properties of projections. Let us start from an elementary observation:
Let be a pure state in the face and let be the atomic effect such that . If is an atomic transformation such that , then . Moreover, if is the effect associated to the face , then we have .
Proof. By lemma 16, the effect is atomic. Now, since , we have . However, by theorem 8 is the unique atomic effect such that . Hence, . Moreover, writing as with , (lemma 47), we obtain .
When applied to the case of projections, the above lemma gives the following
Let be a pure state in the face and let be the atomic effect such that . Then, we have . Moreover, if is the effect associated to the face , then we have .
The counterpart of corollary 34 is given as follows:
Let be a pure state in the face and let be the atomic effect such that . Then, we have . Moreover, if is the effect associated to the face , then we have .
Proof. By lemma 16, the effect is atomic. Hence, must be proportional to an atomic effect with , for some proportionality constant , that is . We want to prove that is zero. By contradiction, suppose that . Let be the pure state such that . Now, since , we have , which implies . Hence, is perfectly distinguishable from , which in turn implies that belongs to . We then have (the last equality follows from the fact that and belong to and , respectively, and hence are perfectly distinguishable). This is in contradiction with the assumption , thus concluding the proof that . Moreover, writing as with , , we obtain .
Combining corollary 34 and lemma 52 we obtain an important property of projections, expressed by the following:
If is a projection on the face , then one has .
Proof. The thesis follows from corollary 34 and lemma 52 and from the fact that .
In the following we will see that for every face there exists a unique projection. To prove that, let us start from the existence:
For every face of there exists a projection .
Proof. By lemma 18, there exists a system and an atomic transformation with . Then, if is a purification of , we can define the state . By lemma 16, is a pure state. Moreover, the pure states and have the same marginal on system : indeed, we have and, by definition, , which by theorem 1 implies . If and are two arbitrary pure states of and , respectively, the uniqueness of purification stated by Postulate 1 implies that there exists a reversible channel such that
Now, take the atomic effect such that , and define the transformation as
Applying on both sides of Eq. (30) we then obtain
and, therefore, . Moreover, the transformation is atomic, being the composition of atomic transformations (lemma 16). Finally, we have : indeed, by construction of we have
This implies , and, therefore, . In conclusion, is the desired projection.
To prove the uniqueness of the projection we need two auxiliary lemmas, given in the following.
Proof. The state is pure by lemma 16. Let us choose a maximal set of perfectly distinguishable pure states such that is maximal in . Now, we have
having used that (theorem 9), and the definition of .
Let be a projection. A transformation satisfies if and only if
Proof. Let be the purification of defined in lemma 54. Since , we have . In other words, we have . Since is dynamically faithful, this implies that . Conversely, Eq. (32) implies that for , , namely .
The projection satisfying Definition 12 is unique.
We now show a few simple properties of projections. In the following, given a maximal set of perfectly distinguishable pure states and any subset we define (with a slight abuse of notation) , and as the projection on the face . We will refer to as the face generated by .
For two arbitrary subsets one has
In particular, if one has .
Proof. First of all, is atomic, being the product of two atomic transformations. Moreover, since the face is contained in the faces and , we have for every . In other words, . Moreover, if we have . By lemma 49 and and by the uniqueness of projections (theorem 14) we then obtain that is the projection on the face generated by .
Every projection satisfies the identity .
Proof. Consider a maximal set of perfectly distinguishable pure states such that is maximal in . In this way is the face generated by , and, therefore . The thesis follows by taking in lemma 56.
For every state such that , the normalized state defined by
Proof. By corollary 35, we have . Since , we must have , and, therefore, the state in Eq. (33) is well defined. Moreover, using the definition of we obtain
having used corollaries 34 and 35 for the last equality. Finally, lemma 46 implies that belongs to the face .
Let be the projection on the pure state and be the atomic effect such that . Then for every state one has where .
Proof. Recall that, by corollary 35, we have . If then clearly . Otherwise, the proof is a straightforward application of corollary 37.
We conclude the present subsection with a result that will be useful in the next subsection.
An atomic transformation satisfies if and only if
Conversely, suppose that Eq. (34) is satisfied. Let be a pure state in and be the atomic effect such that . Then, we have
having used the relation (corollary 34). Then, by theorem 7, . Since is arbitrary, this implies .
XI.3 Projection of a pure state on two orthogonal faces
In Section X we proved a number of results concerning two-dimensional systems. Some properties of two-dimensional systems will be extended to the case of generic systems using the following lemma:
Consider a pure state and two complementary projections and . Then, belongs to the face identified by the state .
Proof. If (resp. ), then there is nothing to prove: this means that (resp. ) and the thesis is trivially true. Suppose now that and . Using the notation , , we can define the two pure states , , and the probabilities . In this way we have for and . Taking the atomic effect such that we have , where is the effect associated to the face . Recalling that for (corollary 34), we then conclude the following
Finally, lemma 46 yields .
A consequence of lemma 58 is the following
Let be a pure state, be the unique atomic effect such that , and be a face in . If is perfectly distinguishable from and from then is perfectly distinguishable from . In particular, one has .
Proof. Since is perfectly distinguishable from and , it is also perfectly distinguishable from any convex combination of them (corollary 33). Equivalently, is perfectly distinguishable from the face identified by . In particular, it must be perfectly distinguishable from , which belongs to by virtue of lemma 58. If is the atomic effect such that , then by lemma 36 we have .
A technical result that will be useful in the following is:
Let be a pure state such that and . Define the pure states and and the mixed state . Then, we have
Proof. Let be a maximal set of perfectly distinguishable pure states in , chosen in such a way that , and let be a maximal set of perfectly distinguishable pure states in , chosen in such a way that . Defining the sets , , and we then have , and . Using lemma 56 we obtain
We conclude this subsection with an important observation about the group of reversible transformations that act as the identity on two orthogonal faces and . If is a face of , let us define as the group of all reversible transformations such that
For every face such that and , the group is topologically equivalent to a circle.
We now prove that in fact they are the whole circle. Let be a state in . Since belongs to the face , we obtain
XII The superposition principle
The validity of the superposition principle, proved for two-dimensional systems using the geometry of the Bloch sphere (corollary 31), can be now extended to arbitrary systems thanks to lemma 58.
(Superposition principle for general systems) Let be a maximal set of perfectly distinguishable pure states and be the observation-test such that . Then, for every choice of probabilities , there exists at least one pure state such that
where is the projection on .
Proof. Let us first prove the equivalence between Eqs. (35) and (36). From Eq. (36) we obtain Eq. (35) using the relation , which follows from corollary 35. Conversely, from Eq. (35) we obtain Eq. (36) using corollary 38 . Now, we will prove Eq. (35) by induction. The statement for is proved by corollary 31. Assume that the statement holds for every system of dimension and suppose that . Let be the face identified by and be the orthogonal face, identified by the state . Now there are two cases: either or . If , then there is nothing to prove: the desired state is . Then, suppose that . Using the induction hypothesis and the compression axiom 3 we can find a state such that , with , . Let us then define a new maximal set of perfectly distinguishable pure states , with and . Note that one has , that is, is the face generated by the states . Now consider the two-dimensional face identified by . By corollary 31 (superposition principle for qubits) we know that there exists a pure state with and . Let us define and . Then, we have and , and by lemma 56,
having used corollary 38 for the last equality. Finally, for we have
On the other hand we have .
Using the superposition principle and the spectral decomposition of theorem 10 we can now show that every state of system has a purification in provided :
For every state and for every system with there exists a purification of in .
Proof. Take the spectral decomposition of , given by , where are probabilities and is a maximal set of perfectly distinguishable pure states. Let be a maximal set of perfectly distinguishable pure states and (resp. ) be the test such that (resp. ). Clearly, is a maximal set of perfectly distinguishable pure states for . Then, by the superposition principle (theorem 16) there exists a pure state such that . Equivalently, we have for every and for . Summing over we then obtain .
In the terminology of Ref. purification, lemma 61 states that a system with is complete for the purification of system .
As a consequence of lemma 61 we have the following:
XII.2 Equivalence of systems with equal dimension
We are now in position to prove that two systems and with the same dimension are operationally equivalent, namely that there is a reversible transformation from to . In other words, we prove that the informational dimension classifies the systems of our theory up to operational equivalence. The fact that this property is derived from the principles, rather than being assumed from the start, is one of the important differences of our work with respect to Refs. Har01; DakBru09; Mas10. Another difference is that here the equivalence of systems with the same dimension is proved after the derivation of the qubit, whereas in Refs. Har01; DakBru09; Mas10 the derivation of the qubit requires the equivalence of systems with the same dimension.
(Operational equivalence of systems with equal dimension) Every two systems an with are operationally equivalent.
XII.3 Reversible operations of perfectly distinguishable pure states
An important consequence of the superposition principle is the possibility of transforming an arbitrary maximal set of perfectly distinguishable pure states into another via a reversible transformation:
Let and be two systems with and let (resp. ) be a maximal set of perfectly distinguishable pure states in (resp. ). Then, there exists a reversible transformation such that .
Combining Eqs. (37), (38), (39) we finally obtain
that is, for every .
XIII Derivation of the density matrix formalism
The goal of this section is to show that our set of axioms implies that
the set of effects is the set of positive matrices bounded by the identity
the pairing between a state and an effect is given by the trace of the product of the corresponding matrices.
Using the result of theorem 3, we will then obtain that all the physical transformations in our theory are exactly the physical transformations allowed in quantum mechanics. This will conclude our derivation of quantum theory.
Let () be the two perfectly distinguishable states in the direction of the -axis (-axis) and define
An immediate observation is the following:
Proof. Linear independence is evident from the geometry of the Bloch sphere. Moreover, for the states are perfectly distinguishable from , and, therefore . If , since the states , lie on the equator of the Bloch sphere, we know that for . Hence, .
Let , and consider the projection . Then, for and , one has for .
Proof. Using lemma 56 and corollary 38 we obtain
Since the face is isomorphic to the Bloch sphere and the state since , lie on the equator of the Bloch sphere, we know that . This implies
Proof. Since the number of vectors is exactly , to prove that they form a basis it is enough to show that they are linearly independent. Suppose that there exists a vector of coefficients such that
Applying the projection on both sides and using lemma 63 we obtain
However, we know that the vectors are linearly independent. Consequently, for all .
XIII.2 The matrices
Since the state space for system spans a real vector space of dimension , we can decide to represent the vectors as Hermitian matrices. Precisely, we associate the vector to the matrix defined by
the vector to the matrix
and the vector to the matrix
where can take the values or . The freedom in the choice of will be useful in subsection XIII.3, where we will introduce the representation of composite systems of two qubits. However, this choice of sign plays no role in the present subsection, and for simplicity we will take the positive sign.
Recall that in principle any orthogonal direction in the plane orthogonal to the -axis can be chosen to be the -axis. In general, the other possible choices for the -axis will lead to matrices of the form
and the corresponding choice for the -axis will lead to a matrices of the form
and the expansion coefficients are all real. Hence, each state is in one-to-one correspondence with a Hermitian matrix, given by
Since effects are linear functionals on states, they are also represented by Hermitian matrices. We will indicate with the Hermitian matrix associated to the effect . The matrix is uniquely defined by the relation:
In the rest of the section we show that the set of matrices is the whole set of positive Hermitian matrices with unit trace and that the set of matrices is the set of positive Hermitian matrices bounded by the identity.
The invariant state has matrix representation , where is the identity matrix in dimension .
Proof. Obvious from the expression and from the matrix representation of the states in Eq. (41).
Let be the atomic effect such that . Then, the effect has matrix representation such that .
Proof. Let be an arbitrary state. Expanding as in Eq. (46) and using lemma 62 we obtain . On the other hand, by Eq. (47) we have that is the -th diagonal element of the matrix : by definition of [eq. (41)], this implies . Now, by construction we have for every . Hence, .
The deterministic effect has matrix representation .
Proof. Obvious from the expression , combined with lemma 66 and Eq. (41).
For every state one has
Proof. .
The matrix elements of for a pure state are , with , , and .
Proof. First of all, the diagonal elements of are given by [cf. Eqs. (46) and (47)]. Denoting the -th element by , we clearly have . Now, the projection is a state in the face , and, by our choice of representation, the corresponding matrix is proportional to a pure qubit state (non-negative rank-one matrix). On the other hand, it is easy to see from Eqs. (46) and (47) that is the matrix with the same elements as in the block corresponding to the qubit and 0 elsewhere. In order to be positive and rank-one the corresponding sub-matrix must have the off-diagonal elements , for some with . Repeating the same argument for all choices of indices , the thesis follows.
For a pure state , the corresponding atomic effect such that has a matrix representation with the property that .
Proof. We already know that the statement holds for , where we proved the Bloch sphere representation, equivalent to the fact that states and effects are represented as positive complex matrices, with the set of pure states identified with the set of all rank-one projectors. Let us now consider a generic system . For every , the face generated by can be encoded in a two-dimensional system. Therefore, the matrices and are positive (also, recall that all matrix elements outside the block are zero). Let be the pure state in the face that is perfectly distinguishable from . Note that, since belongs to the face , it is also perfectly distinguishable from . Hence, is perfectly distinguishable from and, in particular, (lemma 59). This implies the relation
Now, since the matrix is positive, the above relation implies , where . Finally, repeating the argument for all possible values of , we obtain that for every , that is, . Taking the trace on both sides we obtain . To prove that , we use the relation .
We conclude with a simple corollary that will be used in the next subsection:
Let be a pure state and let be a set of pure states. If the state can be written as
for some real coefficients , then the atomic effect such that is given by
where is the atomic effect such that .
Proof. For every by theorem 18 one has
thus implying the thesis.
XIII.3 Choice of axes for a two-qubit system
If and are two systems with , then we can use two different types of matrix representations for the states of the composite system:
The first type of representation is the representation introduced through lemma 64: here we will refer to it as the standard representation. Note that there are many different representations of this type because for every pair there is freedom in choice of the - and -axis [cf. Eqs. (44 ) and (45)]
The second type of representation is the tensor product representation , defined by the tensor product of matrices representing states of systems and : for a state , with , we have
We now show a few properties of the tensor representation. Let denote the matrix corresponding to the effect in the tensor representation, that is, the matrix defined by
It is easy to show that the matrix representation for effects must satisfy the analogue of Eq. (48):
Let be a bipartite effect, written as . Then one has
where (resp. ) is the matrix representing the single-qubit effect (resp. ) in the standard representation for qubit (resp. ).
Proof. For every bipartite state one has
which implies the thesis.
Let be a pure state and let be the atomic effect such that . Then one has .
For every bipartite state , one has .
Hence, , where is the identity matrix. By lemma 68, we then have and, therefore .
Finally, an immediate consequence of local distinguishability is the following:
Then, we have for every .
The rest of this subsection is aimed at showing that, with a suitable choice of matrix representation for system , the standard representation coincides with the tensor representation, that is, for every . This technical result is important because some properties used in our derivation are easily proved in the standard representation, while the property expressed by lemma 69 is easily proved in the tensor representation: it is then essential to show that we can construct a representation that enjoys both properties.
The four states are clearly a maximal set of perfectly distinguishable pure states in . In the following we will construct the standard representation starting from this set.
For a composite system with one can choose the standard representation in such a way that the following equalities hold
Proof. Let us choose single-qubit representations and that satisfy Eqs. (41), (42), and (43). On the other hand, choosing tho states in lexicographic order as the four distinguishable states for the standard representation, we have
With this choice, we get for every . This proves Eq. (51). Let us now prove Eqs. (52) and (53). Consider the two-dimensional face , generated by the states and . This face is the face identified by the state , and we have . Therefore we can choose the vectors , to satisfy the relation , . Now, in the standard representation we have
[cf. Eqs. (42) and (43)]. This implies , for . Repeating the same argument for the face , , and we obtain the proof of Eqs. (52) and (53).
In order to prove that, with a suitable choice of axes, the standard representation coincides with the tensor representation—i.e. for every —it remains to find a choice of axes such that , . This will be proved in the following.
Let be a pure state such that [such a state exists due to the superposition principle]. With a suitable choice of the matrix representation , the state is represented by the matrix
where is the identity matrix. The most general form for is then the following
having defined and .
Now, by construction the state satisfies the condition
By definition of the tensor representation, the conditional states are described by the diagonal blocks of the matrix :
Since the states and are pure, the above matrices must be be rank-one. Moreover, their trace must be equal to , . Then we have two possibilities. Either i) and or ii) . In the case i), Eq. 54 holds. In the case ii), to prove Eq. (54) we need to change our choice of matrix representation for the qubit . Precisely, we make the following change:
where . Note that the inversion of the axes, sending to for every is not an allowed physical transformation, but this is not a problem here, because Eq. (58) is just a new choice of matrix representation, in which the set of states of system is still represented by the Bloch sphere.
More concisely, the change of matrix representation can be expressed as
Note that in the new representation the physical transformation is still represented as : indeed we have
Let us now prove Eq. (55). Using the fact that by definition one can directly verify the relation
This is precisely the matrix version of Eq. (55).
Note that the choice of needed in Eq. (54) is compatible with the choice of needed in lemma 70: indeed, to prove compatibility we only have to show that the representation used in Eq. (54) has the property , . This property is automatically guaranteed by the relation , and by Eq. (57) with and .
In the standard representation the state is represented by the matrix
Proof. The thesis follows from theorem 17 and lemma 70.
We now define the reversible transformations and as follows
Also, we define the states , and as
The states , and have the following tensor representation
Proof. Eq. (72) is obtained from Eq. (54) by explicit calculation using lemma 69 and Eq. (60). Then, the validity of Eq. (72) is easily obtained from Eq. (55) using the relations
The states , and have a standard representation of the form
with as in corollary 59, and .
Proof. Let us start from . First, from Eq.(72) it is immediate to obtain and . This gives the diagonal elements of . Then, using theorem 17 we obtain that must be as in Eq. (63), for some value of . Let us now consider . Again, the diagonal elements of the matrix are obtained from Eq. (72), which in this case yields and . Hence, by theorem 17 we must have
for some value of . Now, denote by the effect such that . We then have
having used theorem 18, corollary 44, and Eq. (72). Hence, we have , which implies , as in Eq. (63). Finally, the same arguments can be used for : The diagonal elements of are obtained from the relations and , which follow from Eq. (72). This implies that the matrix has the form
for some . The relation then implies .
Let us now consider the four vectors defined as follows
By the previous results, it is immediate to obtain the matrix representations of these vectors. In the tensor representation, using Eqs. (54) and (72) we obtain
while in the standard representation, using Eqs. (59) and (63), we obtain
Comparing the two matrix representations we are now in position to prove the desired result:
With a suitable choice of axes, one has for every .
Proof. For the face , using the freedom coming from Eqs. (43) and (44), we redefine the and axes so that and . In this way we have
Likewise, for the face we redefine the and axes so that and , so that we have
Finally, using Eqs. (55), (72), and (64) we have the relations
Since and coincide on the right-hand side of each equality, they must also coincide on the left-hand side.
With a suitable choice of axes, the standard representation coincides with the tensor representation, that is, for every .
Proof. Combining lemma 70 with lemma 74 we obtain that and coincide on the tensor products basis , where . By linearity, and coincide on every state.
From now on, whenever we will consider a composite system where and are two-dimensional we will adopt the choice that guarantees that the standard representation coincides with the tensor representation.
XIII.4 Positivity of the matrices
In this paragraph we show that the states in our theory can be represented by positive matrices. This amounts to prove that for every system , the set of states can be represented as a subset of the set of density matrices in dimension . This result will be completed in subsection XIII.5, where we will see that, in fact, every density matrix in dimension corresponds to some state of .
The starting point to prove positivity is the following:
Let and be two-dimensional systems. Then, for every pure state one has .
where and are the reversible transformations defined by and , respectively ( and are physical transformations by virtue of corollary 30). Here we used the fact that the standard two-qubit representation coincides with the tensor representation and, therefore, . Denoting the pure state by we then have
Since by theorem 17 we have , we conclude
Let be a system of dimension . Then, with a suitable choice of matrix representation the pure states of are represented by positive matrices.
Proof. The system is operationally equivalent to the composite system , where . Let be the reversible transformation implementing the equivalence. Now, we know that the states of are represented by positive matrices. If we define the basis vectors for by applying to the basis for , then we obtain that the states of are represented by the same matrices representing the states of .
Let be a system with . With a suitable choice of matrix representation, the matrix is positive for every pure state .
Proof. Let be a system with . By corollary 47 the states of are represented by positive matrices. Define the state , where are four perfectly distinguishable pure states. By the compression axiom, the face can be encoded in a three-dimensional system (corollary 40). In fact, since is operationally equivalent to , the face can be encoded in . Let and be the encoding and decoding operation, respectively. If we define the basis vectors for by applying to the basis vectors for the face , then we obtain that the states of are represented by the same matrices representing the states in the face . Since these matrices are positive, the thesis follows.
From now on, for every three-dimensional system we will choose the and axes so that is positive for every .
Let be a pure state with . Then, the corresponding matrix , given by
Proof. The relation can be trivially satisfied when for some . Hence, let us assume . Computing the determinant of one obtains . Since is positive, we must have . If the only possibility is .
Corollary 49 can be easily extended to systems of arbitrary dimension. To this purpose, we choose the and axes in such a way that the projection of every state on a three-dimensional face is represented by a positive matrix.
Proof. Consider a triple . Then the state is proportional to a pure state of a three dimensional system, whose representation is the square sub-matrix of with elements , . Now, corollary 49 forces the relation . Since this relation must hold for every choice of the triple , if we define , then we have . It is then immediate to verify that , where .
For every system , the state space can be represented as a subset of the set of density matrices in dimension .
Proof. For every state the matrix is Hermitian by construction, with unit trace by corollary 42, and positive since it is a convex mixture of positive matrices.
XIII.5 Quantum theory in finite dimensions
Here we conclude our derivation of quantum theory by showing that every density matrix in dimension corresponds to some state .
We already know from the superposition principle (lemma 35) that for every choice probabilities there is a pure state such that are the diagonal elements of . Thus, the set of density matrices corresponding to pure states contains at least one matrix of the form , with . It only remains to prove that every possible choice of phases corresponds to some pure state.
Recall that for a face we defined the group to be the group of reversible transformations such that and . We then have the following
Consider a system with . Let be a maximal set of perfectly distinguishable pure states, be the face identified by and its orthogonal face, identified by the state . If is a reversible transformation in , then the action of is given by
where is the identity matrix and .
Proof. Consider an arbitrary state and its matrix representation
Let us start from the case . Since , we have (lemma 51). This implies that sends states in the face to states in the face : indeed, for every one has , which implies (lemma 46). In other words, the restriction of to the face is a reversible qubit transformation. Therefore, the action of on a state must be given by
for some . Similarly, we can see that sends states in the face to states in the face . Hence, for every we have
for some . We now show that . To see that, consider a generic state , with the property that for every (such state exists due to the superposition principle of theorem 16). Writing as in Eq. (65) we then have
Now, since and are pure states, by corollary 49 we must have
By comparison we obtain . This proves Eq. (66) for . The proof for is then immediate: for every three-dimensional face the action of is given Eq. (66) for some . However, since the two faces and overlap on we must have . Similarly . We conclude that for every . This proves Eq. (66) in the general case.
We now show that every possible phase shift in Eq. (66) corresponds to a physical transformation:
A transformation of the form of Eq. (66) is a reversible transformation for every .
Proof. By lemma 77, the group is a subgroup of . Now, there are two possibilities: either is a (finite) cyclic group or coincides with . However, we know from theorem 15 that has a continuum of elements. Hence, and can take every value in .
An obvious corollary of the previous lemmas is the following
The transformation defined by
where is the diagonal matrix with diagonal elements is a reversible transformation for every vector .
This leads directly to the conclusion of our derivation:
Proof. Let . For every choice of probabilities there exists at least one pure state such that for every (lemma 35 ). This state is represented by the matrix with (lemma 76). Finally, we can transform with every reversible transformation defined in Eq. (67), thus obtaining where . Since and are arbitrary, this means that every rank-one density matrix corresponds to some pure state. Taking the possible convex mixtures we obtain that every density matrix corresponds to some state of system .
Choosing a suitable representation , we proved that for every system the set of normalized states is the whole set of density matrices in dimension . Thanks to the purification postulate, this is enough to prove that all the effects and all the transformations allowed in our theory are exactly the effects and the transformations allowed in quantum theory. Precisely we have the following
Proof. We proved that our theory has the same normalized states of quantum theory. On the other hand, quantum theory is a theory with purification and in quantum theory the possible physical transformations are quantum operations, i.e. completely positive trace-preserving maps. The thesis then follows from the fact that two theories with purification that have the same set of normalized states are necessarily the same (theorem 3).
XIV Conclusion
Quantum theory can be derived from purely informational principles. In particular, it belongs to a broad class of theories of information-processing that includes classical and quantum information theory as special cases. Within this class, quantum theory is identified uniquely by the purification postulate, stating that the ignorance about a part is always compatible with the maximal knowledge of the whole in an essentially unique way. This postulate appears as the origin of the key features of quantum information processing, such as no-cloning, teleportation, and error correction (see also Ref. purification). The general vision underlying the present work is that the main primitives of quantum information processing should be derived directly from the principles, without the abstract mathematics of Hilbert spaces, in order to make the revolutionary aspects of quantum information immediately accessible and to place them in the broader context of the fundamental laws of physics.
Finally, we would like to comment on possible generalizations of our work. As in any axiomatic construction, one can ask how the results change when the principles are modified. For example, one may be interested in relaxing the local distinguishability axiom and in considering theories, like quantum theory on real Hilbert spaces, where global measurements are essential to characterize the state of a composite system. In this direction, the results of Ref. purification suggest that also quantum theory on real Hilbert spaces can be derived from the purification principle, after that the local distinguishability requirement has been suitably relaxed. A possible way to weaken the local distinguishability requirement is to assume only the property of local distinguishability from pure states proposed in Ref. purification: this property states that the probability of distinguishing two states by local measurements is larger than whenever one of the two states is pure. A different way to relax local distinguishability would be to assume the property of 2-local tomography proposed in Ref. hw, which requires that the state of a multipartite system can be completely characterized using only measurements on bipartite subsystems. This property is equivalent to 2-local distinguishability, defined as the requirement that two different states of a multipartite system can be distinguished with probability of success larger than using only local measurements or measurements on bipartite subsystems.
A more radical generalization of our work would be to relax the assumption of causality. This would be particularly important for the discussion of quantum gravity scenarios, where the causal structure is not given a priori but is part of the dynamical variables of the theory. In this respect, the contribution of our work is twofold. First, it makes evident how fundamental is the assumption of causality in the ordinary formulation of quantum theory: the whole formalism of quantum states as density matrices with unit trace, quantum measurements as resolutions of the identity, and quantum channels as trace-preserving maps is crucially based on it. Technically speaking, the fact that the normalization of a state is given by a single linear functional (the trace, in quantum theory) is the signature of causality. This partly explains the troubles and paradoxes encountered when trying to combine the formalism of density matrices with non-causal evolutions, as in Deutsch’s model for close timelike curves deutsch; bennett. Moreover, given that the usual notion of normalization has to be abandoned in the non-causal scenario, and that the ordinary quantum formalism becomes inadequate, one may ask in what sense a theory of quantum gravity would be “quantum”. The suggestion coming from our work is that a “quantum” theory is a theory satisfying the purification principle, which can be suitably formulated even in the absence of causality inprep. The discussion of theories with purification in the non-causal scenario is an exciting avenue of future research.