Informational derivation of Quantum Theory

G. Chiribella, G. M. D'Ariano, P. Perinotti

I Introduction

More than eighty years after its formulation, quantum theory is still mysterious. The theory has a solid mathematical foundation, addressed by Hilbert, von Neumann and Nordheim in 1928 HNN28 and brought to completion in the monumental work by von Neumann von32. However, this formulation is based on the abstract framework of Hilbert spaces and self-adjoint operators, which, to say the least, are far from having an intuitive physical meaning. For example, the postulate stating that the pure states of a physical system are represented by unit vectors in a suitable Hilbert space appears as rather artificial: which are the physical laws that lead to this very specific choice of mathematical representation? The problem with the standard textbook formulations of quantum theory is that the postulates therein impose particular mathematical structures without providing any fundamental reason for this choice: the mathematics of Hilbert spaces is adopted without further questioning as a prescription that “works well” when used as a black box to produce experimental predictions. In a satisfactory axiomatization of Quantum Theory, instead, the mathematical structures of Hilbert spaces (or C*-algebras) should emerge as consequences of physically meaningful postulates, that is, postulates formulated exclusively in the language of physics: this language refers to notions like physical system, experiment, or physical process and not to notions like Hilbert space, self-adjoint operator, or unitary operator. Note that any serious axiomatization has to be based on postulates that can be precisely translated in mathematical terms. However, the point with the present status of quantum theory is that there are postulates that have a precise mathematical statement, but cannot be translated back into language of physics. Those are the postulates that one would like to avoid.

The need for a deeper understanding of quantum theory in terms of fundamental principles was clear since the very beginning. Von Neumann himself expressed his dissatisfaction with his mathematical formulation of Quantum Theory with the surprising words “I don’t believe in Hilbert space anymore”, reported by Birkhoff in Birk61. Realizing the physical relevance of the axiomatization problem, Birkhoff and von Neumann made an attempt to understand quantum theory as a new form of logic BirkVN36: the key idea was that propositions about the physical world must be treated in a suitable logical framework, different from classical logics, where the operations AND and OR are no longer distributive. This work inaugurated the tradition of quantum logics, which led to several attempts to axiomatize quantum theory, notably by Mackey Mack63 and Jauch and Piron JauPir63 (see Ref. coeckerev for a review on the more recent progresses of quantum logics). In general, a certain degree of technicality, mainly related to the emphasis on infinite-dimensional systems, makes these results far from providing a clear-cut description of quantum theory in terms of fundamental principles. Later Ludwig initiated an axiomatization program Lud83 adopting an operational approach, where the basic notions are those of preparation devices and measuring devices and the postulates specify how preparations and measurements combine to give the probabilities of experimental outcomes. However, despite the original intent, Ludwig’s axiomatization did not succeed in deriving Hilbert spaces from purely operational notions, as some of the postulates still contained mathematical notions with no operational interpretation.

More recently, the rise of quantum information science moved the emphasis from logics to information processing. The new field clearly showed that the mathematical principles of quantum theory imply an enormous amount of information-theoretic consequences, such as the no-cloning theorem wootterszurek; dieks, the possibility of teleportation tele, secure key distribution wiesner; bb84; e91, or of factoring numbers in polynomial time shor. The natural question is whether the implication can be reversed: is it possible to retrieve quantum theory from a set of purely informational principles? Another contribution of quantum information has been to shift the emphasis to finite dimensional systems, which allow for a simpler treatment but still possess all the remarkable quantum features. In a sense, the study of finite dimensional systems allows one to decouple the conceptual difficulties in our understanding of quantum theory from the technical difficulties of infinite dimensional systems.

In this scenario, Hardy’s 2001 work Har01 re-opened the debate about the axiomatizations of quantum theory with fresh ideas. Hardy’s proposal was based on five main assumptions about the relation between dimension of the state space and the number of perfectly distinguishable states of a given system, about the structure of composite systems, and about the possibility of connecting any two pure states of a physical system through a continuous path of reversible transformations. However, some of these assumptions directly refer to the mathematical properties of the state space (in particular, the “Simplicity Axiom” 2, which is an abstract statement about the functional dependence of the state space dimension on the number of perfectly distinguishable states). Very recently, building on Hardy’s work there have been two new attempts of axiomatization by Dakic and Brukner DakBru09 and Masanes and Müller Mas10. Although these works succeeded in removing the “Simplicity Axiom”, they still contain mathematical assumptions that cannot be understood in elementary physical terms (see e.g. requirement 5 of Ref. Mas10, which assumes that “all mathematically well-defined measurements are allowed by the theory”).

Another approach to the axiomatization of quantum theory was pursued by one of the authors in a series of works maurofirst culminated in Ref. maurolast. These works tackled the problem using operational principles related to tomography and calibration of physical devices, experimental complexity, and to the composition of elementary transformations. In particular this research introduced the concept of dynamically faithful states, namely states that can be used for the complete tomography of physical processes. Although this approach went very close to deriving quantum theory also in this case one mathematical assumption without operational interpretation was needed (see the CJ postulate of Ref. maurolast).

In this paper we provide a complete derivation of finite dimensional quantum theory based of purely operational principles. Our principles do not refer to abstract properties of the mathematical structures that we use to represent states, transformations or measurements, but only to the way in which states, transformations and measurements combine with each other. More specifically, our principles are of informational nature: they assert basic properties of information-processing, such as the possibility or impossibility to carry out certain tasks by manipulating physical systems. In this approach the rules by which information can be processed determine the physical theory, in accordance with Wheeler’s program “it from bit”, for which he argued that “all things physical are information-theoretic in origin” wheeler. Note that, however, our axiomatization of quantum theory is relevant, as a rigorous result, also for those who do not share Wheeler’s ideas on the informational origin of physics. In particular, in the process of deriving quantum theory we provide alternative proofs for many key features of the Hilbert space formalism, such as the spectral decomposition of self-adjoint operators or the existence of projections. The interesting feature of these proofs is that they are obtained by manipulation of the principles, without assuming Hilbert spaces form the start.

The main message of our work is simple: within a standard class of theories of information processing, quantum theory is uniquely identified by a single postulate: purification. The purification postulate, introduced in Ref. purification, expresses a distinctive feature of quantum theory, namely that the ignorance about a part is always compatible with the maximal knowledge of the whole. The key role of this feature was noticed already in 1935 by Schrödinger in his discussion about entanglement Schr35, of which he famously wrote “I would not call that one but rather the characteristic trait of quantum mechanics, the one that enforces its entire departure from classical lines of thought”. In a sense, our work can be viewed as the concrete realization of Schrödinger’s claim: the fact that every physical state can be viewed as the marginal of some pure state of a compound system is indeed the key to single out quantum theory within a standard set of possible theories. It is worth stressing, however, that the purification principle assumed in this paper includes a requirement that was not explicitly mentioned in Schrödinger’s discussion: if two pure states of a composite system AB{\rm A}{\rm B} have the same marginal on system A{\rm A}, then they are connected by some reversible transformation on system B{\rm B}. In other words, we assume that all purifications of a given mixed state are equivalent under local reversible operations anchese….

The purification principle expresses a law of conservation of information, stating that at least in principle, irreversibility can always be reduced to the lack of control over an environment. More precisely, the purification principle is equivalent to the statement that every irreversible process can be simulated in an essentially unique way by a reversible interaction of the system with an environment, which is initially in a pure state purification. This statement can also be extended to include the case of measurement processes, and in that case it implies the possibility of arbitrarily shifting the cut between the observer and the observed system purification. The possibility of such a shift was considered by von Neumann as a “fundamental requirement of the scientific viewpoint” (see p. 418 of von32) and his discussion of the measurement process was exactly aimed to show that quantum theory fulfils this requirement.

Besides Schrödinger’s discussion on entanglement and von Neumann’s discussion of the measurement process, the purification principle is deeply rooted in the structure of quantum theory. At the purely mathematical level, it plays a crucial role in the theory of C*-algebras of operators on separable Hilbert spaces, where the purification principle is equivalent to the Gelfand-Naimark-Segal (GNS) construction arveson and implies the celebrated Stinespring’s theorem stine. On the other hand, purification is a cornerstone of quantum information, lying at the origin of most quantum protocols. As it was shown in Ref. purification, the purification principle directly implies crucial features like no-cloning, teleportation, no-information without disturbance, error correction, the impossibility of bit commitment, and the “no-programming” theorem of Ref. no-prog.

In addition to the purification postulate, our derivation of quantum theory is based on five informational axioms. The reason why we call them “axioms”, as opposed to the the purification “postulate”, is that they are not at all specific of quantum theory. These axioms represent standard features of information-processing that everyone would, more or less implicitly, assume. They define a class of theories of information-processing that includes, for example, classical information theory, quantum information theory, and quantum theory with superselection rules. The question whether there are other theories satisfying our five axioms and, in case of a positive answer, the full classification of these theories is currently an open problem.

Here we informally illustrate the five axioms, leaving the more detailed description to the remaining part of the paper:

Causality: the probability of a measurement outcome at a certain time does not depend on the choice of measurements that will be performed later.

Perfect distinguishability: if a state is not completely mixed (i.e. if it cannot be obtained as a mixture from any other state), then there exists at least one state that can be perfectly distinguished from it.

Ideal compression: every source of information can be encoded in a suitable physical system in a lossless and maximally efficient fashion. Here lossless means that the information can be decoded without errors and maximally efficient means that every state of the encoding system represents a state in the information source.

Local distinguishability: if two states of a composite system are different, then we can distinguish between them from the statistics of local measurements on the component systems.

Pure conditioning: if a pure state of system AB{\rm A}{\rm B} undergoes an atomic measurement on system A{\rm A}, then each outcome of the measurement induces a pure state on system B{\rm B}. (Here atomic measurement means a measurement that cannot be obtained as a coarse-graining of another measurement).

All these axioms are satisfied by classical information theory. Axiom 5 is even trivial for classical theory, because the only pure states of a composite system AB{\rm A}{\rm B} are the product of pure states of the component systems A{\rm A} and B{\rm B}, and hence the state of system B{\rm B} will be pure irrespectively of what we do on system A{\rm A}.

A stronger version of axiom 5, introduced in Ref. maurolast, is the following:

Atomicity of composition: the sequential composition of two atomic operations is atomic. (Here atomic transformation means a transformation that cannot be obtained from coarse-graining).

However, it turns out that Axiom 5 is enough for our derivation: thanks to the purification postulate we will be able to show the non-trivial implication: Axiom 5 ⇒\Rightarrow Axiom 5’ (see lemma 16).

The paper is organized as follows. In Sec. II we review the framework of operational-probabilistic theories introduced in Ref. purification. This framework will provide the basic notions needed for the formulation of our principles. In Sec. III we introduce the principles from which we will derive Quantum Theory. In Sec. IV we prove some direct consequences of the principles that will be used later in the paper. In Sec. V we discuss the properties of perfectly distinguishable states, while in Sec. VI we prove the existence of a duality between pure states and atomic effects.

The results about distinguishability and duality of pure states and atomic effects allow us to show in Sec. VII that every system has a well defined informational dimension—the operational counterpart of the Hilbert space dimension. Sec. VIII contains the proof that every state can be decomposed as a convex combination of perfectly distinguishable pure states. Similarly, any element of the vector space spanned by the states can be written as a linear combination of perfectly distinguishable states. This result corresponds to the spectral theorem for self-adjoint operators on complex Hilbert spaces. In Sec. IX we prove some results about the maximum teleportation probability, which allow us to derive a functional relation between the dimension of the state space and the number of perfectly distinguishable states of the system. The mathematical representation of systems with two perfectly distinguishable states is derived in Sec. X, where we prove that such systems are indeed two-dimensional quantum systems—a.k.a. qubits. In Sec. XI we construct projections on the faces of the state space of any system and prove their main properties. These results lead to the derivation of the operational analogue of the superposition principle in Sec. XII which allows to prove that systems with the same number of perfectly distinguishable states are operationally equivalent (Subsec. XII.2). The properties of the projections and the superposition principle are then exploited in Sec. XIII—where we extend the density matrix representation from qubits to higher-dimensional systems, thus proving that a system with dd perfectly distinguishable states is indeed a quantum system with dd-dimensional Hilbert space. We conclude the paper with Sec. XIV, where we review our results, discussing future directions for this research.

II The framework

This Section provides a brief summary of the framework of operational-probabilistic theories, which was formulated in Ref. purification. We refer to Ref. purification for an exhaustive presentation of the details of the framework and of the ideas behind it. The operational-probabilistic framework combines the operational language of circuits with the toolbox of probability theory: on the one hand, experiments are described by circuits resulting from the connection of physical devices, on the other hand each device in the circuit can have classical outcomes and the theory provides the probability distribution of outcomes when the devices are connected to form closed circuits (that is, circuits that start with a preparation and end with a measurement).

The notions discussed in this section will allow us to draw a precise distinction between principles with an operational content and exclusively mathematical principles: with the expression ”operational principle” we will mean a principle that can be expressed using only the basic notions of the the operational-probabilistic framework.

A test represents one use of a physical device, like a Stern-Gerlach magnet, a beamsplitter, or a photon counter. The device will have an input system and an output system, labelled by capital letters. The corresponding test can have different classical outcomes, represented by different values of an index i∈Xi\in{\rm X}:

Each outcome i∈Xi\in{\rm X} corresponds to a possible event, represented as

We denote by Transf(A,B){\mathsf{Transf}}({\rm A},{\rm B}) the set of all events from A{\rm A} to B{\rm B}. The reason for this notation is that in the next subsection the elements of Transf(A,B){\mathsf{Transf}}({\rm A},{\rm B}) will be interpreted as transformations with input system A{\rm A} and output system B{\rm B}. If A=B{\rm A}={\rm B} we simply write Transf(A){\mathsf{Transf}}({\rm A}) in place of Transf(A,A){\mathsf{Transf}}({\rm A},{\rm A}).

A test with a single outcome will be called deterministic. This name is justified by the fact that, if there is a single possible outcome, then this outcome will occur with certainty (cf. the probabilistic structure introduced in the next subsection).

Two devices can be composed in a sequence, as long as the input system of the second device is equal to the output system of the first. The events in the composite test are represented as

and are written in formulas as DjCi\mathscr{D}_{j}\mathscr{C}_{i}.

For every system A{\rm A} one can perform the identity-test (or simply, the identity), that is, a test {IA}\{\mathscr{I}_{\rm A}\} with a single outcome, with the property

The subindex A{\rm A} will be dropped from IA\mathscr{I}_{\rm A} where there is no ambiguity.

The letter I{\rm I} will be reserved for the trivial system, which simply means “nothing” nothingquantum. A device with input (resp. output) system I{\rm I} is a device with no input (resp. no output). The corresponding tests will be called preparation-tests (resp. observation-tests). In this case we replace the input (resp. output) wire with a round portion:

In formulas we will write ∣ρi)B\left|\rho_{i}\right)_{\rm B} (resp. (aj∣A\left(a_{j}\right|_{\rm A}). The sets Transf(I,A){\mathsf{Transf}}({\rm I},{\rm A}) and Transf(A,I){\mathsf{Transf}}({\rm A},{\rm I}) will be denoted as St(A){\mathsf{St}}({\rm A}) and Eff(A){\mathsf{Eff}}({\rm A}), respectively. The reason for this special notation is that in the next subsection the elements of St(A){\mathsf{St}}({\rm A}) (resp. Eff(A){\mathsf{Eff}}({\rm A})) will be interpreted as the states (resp. effects) of system A{\rm A}.

From every pair of systems A{\rm A} and B{\rm B} one can form a composite system, denoted by AB{\rm A}{\rm B}. Clearly, composing system A{\rm A} with nothing still gives system A{\rm A}, in formula AI=IA=A{\rm A}{\rm I}={\rm I}{\rm A}={\rm A}. Two devices can be composed in parallel, thus obtaining a new device with composite input and composite output systems. The events in composite test are represented as

and are written in formulas as Ci⊗Dj\mathscr{C}_{i}\otimes\mathscr{D}_{j}. In the special case of states we will often write ∣ρi)∣σj)\left|\rho_{i}\right)\left|\sigma_{j}\right) in place of ρi⊗σj\rho_{i}\otimes\sigma_{j}. Similarly, for effects we will write (ai∣(bj∣\left(a_{i}\right|\left(b_{j}\right| in place of ai⊗bja_{i}\otimes b_{j}.

Sequential and parallel composition commute: one has (Ai⊗Bj)(Ck⊗Dl)=AiCk⊗BjDl(\mathscr{A}_{i}\otimes\mathscr{B}_{j})(\mathscr{C}_{k}\otimes\mathscr{D}_{l})=\mathscr{A}_{i}\mathscr{C}_{k}\otimes\mathscr{B}_{j}\mathscr{D}_{l} for every Ai,Bj,Ck,Dl\mathscr{A}_{i},\mathscr{B}_{j},\mathscr{C}_{k},\mathscr{D}_{l} such that the output of Ai\mathscr{A}_{i} (resp. Bj\mathscr{B}_{j}) coincides with the input of Ck\mathscr{C}_{k} (resp. Dl\mathscr{D}_{l}).

When one of the two tests is the identity, we will omit the box and draw only a straight line, as in

The rules summarized in this section define the operational language of circuits, which has been discussed in detail in a series of inspiring works by Coecke (see in particular Refs. kinder; picturalism). The language of circuits allows one to represent the schematic of an experiment, like e.g.

and also to represent a particular outcome of the experiment

In formula, the above circuit is given by

II.2 Probabilistic structure: states, effects and transformations

On top of the language of circuits, we put a probabilistic structure purification: we declare that the composition of a preparation-test {ρi}i∈X\{\rho_{i}\}_{i\in{\rm X}} with an observation-test {aj}j∈Y\{a_{j}\}_{j\in{\rm Y}} gives rise to a joint probability distribution:

with p(i,j)≥0p(i,j)\geq 0 and ∑i∈X∑j∈Yp(i,j)=1\sum_{i\in{\rm X}}\sum_{j\in{\rm Y}}p(i,j)=1. In formula we write p(i,j)=(aj∣ ⁣ρi)p(i,j)=\left(a_{j}\right|\left.\!\rho_{i}\right). Moreover, if two experiments are run in parallel, we assume that the joint probability distribution is given by the product:

where p(i,k):=(ak∣ ⁣ρi),q(j,l):=(bl∣ ⁣σj)p(i,k):=\left(a_{k}\right|\left.\!\rho_{i}\right),q(j,l):=\left(b_{l}\right|\left.\!\sigma_{j}\right).

II.3 Basic definitions in the operational-probabilistic framework

Here we summarize few elementary definitions that will be used later in the paper. The meaning of the definitions in the case of quantum theory is also discussed.

First, we start from the notions of coarse-graining and refinement. Coarse-graining arises when we join together some outcomes of a test: we say that the test {Dj}j∈Y\{\mathscr{D}_{j}\}_{j\in{\rm Y}} is a coarse-graining of the test {Ci}i∈X\{\mathscr{C}_{i}\}_{i\in{\rm X}} if there is a disjoint partition {Xj}j∈Y\{{\rm X}_{j}\}_{j\in{\rm Y}} of X{\rm X} such that

Conversely, if {Dj}j∈Y\{\mathscr{D}_{j}\}_{j\in{\rm Y}} is a coarse-graining of {Ci}i∈X\{\mathscr{C}_{i}\}_{i\in{\rm X}}, we say that {Ci}i∈X\{\mathscr{C}_{i}\}_{i\in{\rm X}} is a refinement of {Dj}j∈Y\{\mathscr{D}_{j}\}_{j\in{\rm Y}}. Intuitively, a test that refines another is a test that extracts information in a more precise way: it is a test with better “resolving power”.

The notion of refinement also applies to a single transformation: a refinement of the transformation C\mathscr{C} is given by a test {Ci}i∈X\{\mathscr{C}_{i}\}_{i\in{\rm X}} and a subset X0{\rm X}_{0} such that

Accordingly, we say that each transformation Ci,i∈X0\mathscr{C}_{i},i\in{\rm X}_{0} is a refinement of C\mathscr{C}. A transformation C\mathscr{C} is atomic if it has only trivial refinements: if Ci\mathscr{C}_{i} refines C\mathscr{C}, then Ci=pC\mathscr{C}_{i}=p\mathscr{C} for some probability p≥0p\geq 0. A test that consists of atomic transformations is a test whose “resolving power” cannot be further improved.

When discussing states (i.e. transformations with trivial input) we will use the word pure as a synonym of atomic. A pure state describes a situation of maximal knowledge about the system’s preparation, a knowledge that cannot be further refined.

As usual, a state that is not pure will be called mixed. An important notion is that of completely mixed state:

A state is completely mixed if any other state can refine it: precisely, ω∈St(A)\omega\in{\mathsf{St}}({\rm A}) is completely mixed if for every ρ∈St(A)\rho\in{\mathsf{St}}({\rm A}) there is a non-zero probability p>0p>0 such that pρp\rho is a refinement of ω\omega.

Intuitively, a completely mixed state describes a situation of complete ignorance about the system’s preparation: if a system is described by a completely mixed state, then it means that we know so little about its preparation that, in fact, every preparation is possible.

We conclude this paragraph with a couple of definitions that will be used throughout the paper:

A transformation U∈Transf(A,B)\mathscr{U}\in{\mathsf{Transf}}({\rm A},{\rm B}) is reversible if there exists another transformation U−1∈Transf(B,A)\mathscr{U}^{-1}\in{\mathsf{Transf}}({\rm B},{\rm A}) such that U−1U=IA\mathscr{U}^{-1}\mathscr{U}=\mathscr{I}_{\rm A} and UU−1=IB\mathscr{U}\mathscr{U}^{-1}=\mathscr{I}_{\rm B}. When A=B{\rm A}={\rm B} the reversible transformations form a group, indicated as GA{\mathbf{G}}_{\rm A}.

Two systems A{\rm A} and B{\rm B} are operationally equivalent if there exists a reversible transformation U\mathscr{U} from A{\rm A} to B{\rm B}.

When two systems are operationally equivalent one can convert one into the other in a reversible fashion.

II.3.2 Examples in in quantum theory

Preparation-tests are often called quantum information sources in quantum information theory. A generic state ρ\rho is an unnormalized density matrix. A deterministic state, corresponding to a single-outcome preparation-test is a normalized density matrix ρ\rho, with Tr[ρ]=1{\rm Tr}[\rho]=1.

Diagonalizing ρ=∑iαi∣ψi⟩⟨ψi∣\rho=\sum_{i}\alpha_{i}|\psi_{i}\rangle\langle\psi_{i}| we then obtain that each matrix αi∣ψi⟩⟨ψi∣\alpha_{i}|\psi_{i}\rangle\langle\psi_{i}| is a refinement of ρ\rho. More generally, every matrix σ\sigma such that σ≤ρ\sigma\leq\rho is a refinement of ρ\rho. Up to a positive rescaling, all matrices with support contained in the support of ρ\rho are refinements of ρ\rho. A quantum state ρ\rho is atomic (pure) if and only if it is proportional to a rank-one projection. A quantum state is completely mixed if and only if its density matrix has full rank. Note that the quantum state χ=Idd\chi=\frac{I_{d}}{d}, where Id{I_{d}} is the identity d×dd\times d matrix, is a particular example of completely mixed state, but not the only example. Precisely, χ=Idd\chi=\frac{I_{d}}{d} is the unique unitarily invariant state in dimension dd.

Let us now consider the case of observation-tests: in quantum theory an observation-test is given by a POVM (positive operator-valued measure), namely by a collection {Pj}j∈Y\{P_{j}\}_{j\in{\rm Y}} of non-negative d×dd\times d matrices such that

An effect is then a non-negative matrix P≥0P\geq 0 upper bounded by the identity. In quantum theory there is only one deterministic effect, corresponding to a single-outcome observation test: the unique deterministic effect given by the identity matrix. As we will see in the following section, the fact that the deterministic effect is unique is equivalent to the fact that quantum theory is a causal theory.

An effect PP is atomic if and only if PP is proportional to a rank-one projector. An observation-test is atomic if it is a POVM with rank-one elements.

is trace-preserving. A general transformation is then given by a trace non-increasing map, called quantum operation, whereas a deterministic transformation, corresponding to a single-outcome test, is given by a trace-preserving map, called quantum channel.

Any quantum operation C\mathscr{C} can be written in the Kraus form C(ρ)=∑iCiρCi†\mathscr{C}(\rho)=\sum_{i}C_{i}\rho C_{i}^{\dagger}, where Ci:H1→H2C_{i}:{\mathscr{H}}_{1}\to{\mathscr{H}}_{2} are the Kraus operators. Up to a positive scaling, every quantum operation D\mathscr{D} such that the Kraus operators of D\mathscr{D} belong to the linear span of the Kraus operators of C\mathscr{C} is a refinement of C\mathscr{C}. A map C\mathscr{C} is atomic if and only if there is only one Kraus operator in its Kraus form. A reversible transformation in quantum theory is a unitary map U(ρ)=UρU†\mathscr{U}(\rho)=U\rho U^{\dagger}, where U:H1→H2U:{\mathscr{H}}_{1}\to{\mathscr{H}}_{2} is a unitary operator, that is U†U=I1U^{\dagger}U=I_{1} and UU†=I2UU^{\dagger}=I_{2} where I1I_{1} (I2)(I_{2}) is the identity operator on H1{\mathscr{H}}_{1} (H2{\mathscr{H}}_{2}). Two quantum systems are operationally equivalent if and only if the corresponding Hilbert spaces have the same dimension.

II.4 Operational principles

We are now in position to make precise the usage of the expression “operational principle” in the context of this paper. By “operational principle” we mean here a principle that can be stated using only the operational-probablistic language, i.e. using only

the notions of system, test, outcome, probability, state, effect, transformation

their specifications: atomic, pure, mixed, completely mixed

more complex notions constructed from the above terms (e.g. the notion of “reversible transformation”).

The distinction between operational principles and principles referring to abstract mathematical properties, mentioned in the introduction, should now be clear: for example, a statement like “the pure states of a system cannot be cloned” is a valid operational principle, because it can be analyzed in basic operational-probabilistic terms as “for every system A{\rm A} there exists no transformation C\mathscr{C} with input system A{\rm A} and output system AA{\rm A}{\rm A} such that C∣φ)=∣φ)∣φ)\mathscr{C}\left|\varphi\right)=\left|\varphi\right)\left|\varphi\right) for every pure state φ\varphi of A{\rm A} ”. On the contrary, a statement like “the state space of a system with two perfectly distinguishable states is a three-dimensional sphere” is not a valid operational principle, because there is no way to express what it means for a state space to be a three-dimensional sphere in terms of basic operational notions. The fact that a state spate is a sphere may be eventually derived from operational principles, but cannot be assumed as a starting point.

III The principles

We now state the principles used in our derivation. The first five principles express generic features that are shared by both classical and quantum theory. They could be even included in the definition of the background framework: they define the simple model of information processing in which we try to single out quantum theory. For this reason we will call them axioms. The sixth principle in our derivation has a different status: it expresses the genuinely quantum features. A major message of our work is that, within a broad class of theories of information processing, quantum theory is completely described by the purification principle. To emphasize the special role of the sixth principle we will call it postulate, in analogy with the parallel postulate of Euclidean geometry.

The first axiom of our list, causality purification, is so basic that could be considered as part of the background framework. We decided to explicitly present it as an axiom for two reasons: The first reason is that the framework of operational-probabilistic theories can be developed even without this requirement (see Ref.purification for the general framework and Refs. supermaps; comblong for two explicit examples of non-causal theories). The second reason is that we want to stress that causality is an essential ingredient in our derivation. This observation is important in view of possible extensions of quantum theory to quantum gravity scenarios where the causal structure is not defined from the start (see e.g. Hardy in Ref. causaloid).

The probability of preparations is independent of the choice of observations.

In technical terms: if {ρi}i∈X⊂St(A)\{\rho_{i}\}_{i\in{\rm X}}\subset{\mathsf{St}}({\rm A}) is a preparation-test, then the conditional probability of the preparation ρi\rho_{i} given the choice of the observation-test {aj}j∈Y\{a_{j}\}_{j\in{\rm Y}} is the marginal

The axiom states that the marginal probability p(i∣{aj})p\left(i|\{a_{j}\}\right) is independent of the choice of the observation-test {aj}\{a_{j}\}: if {aj}j∈Y\{a_{j}\}_{j\in{\rm Y}} and {bk}k∈Z\{b_{k}\}_{k\in{\rm Z}} are two different observation-tests, then one has p(i∣{aj})=p(i∣{bk})p(i|\{a_{j}\})=p(i|\{b_{k}\}). Loosely speaking, one may refer to causality as a requirement of no-signalling from the future: indeed, causality is equivalent to the fact that the probability of an outcome at a certain time does not depend on the choice of operations that will be done at later times maurolast.

An operational-probabilistic theory that satisfies the causality axiom 1 will be called causal. As we already mentioned, causality is a very basic requirement and could be considered as part of the framework: it provides the notions used to state the other axioms and it implies several facts that will be used frequently in the paper. In fact, in our derivation we do not use the causality axiom directly, but only through its consequences. In the following we briefly summarize the facts and the notations that characterize the framework of causal operational-probabilistic theories, introduced and discussed in detail in Ref. purification. Similar structures have been subsequently considered in Refs. duotenz; caucat within a formal description of circuits in foliable spacetime regions.

First, causality is equivalent to the existence of an effect eAe_{\rm A} such that eA=∑j∈Xaje_{\rm A}=\sum_{j\in{\rm X}}a_{j} for every observation-test {aj}j∈Y\{a_{j}\}_{j\in{\rm Y}}. We call the effect eAe_{\rm A} the deterministic effect for system A{\rm A}. By definition, the effect eAe_{\rm A} is unique. The subindex A{\rm A} in eAe_{\rm A} will be dropped when no confusion can arise.

In a causal theory every test {Ci}i∈X⊂Transf(A,B)\{\mathscr{C}_{i}\}_{i\in{\rm X}}\subset{\mathsf{Transf}}({\rm A},{\rm B}) satisfies the condition

As a consequence, a transformation C∈Transf(A,B)\mathscr{C}\in{\mathsf{Transf}}({\rm A},{\rm B}) satisfies the condition

with the equality if and only if C\mathscr{C} is a channel (i.e. a deterministic transformation, corresponding to a single-outcome test). In Eq. (4) we used the notation (a∣≤(a′∣\left(a\right|\leq\left(a^{\prime}\right| to mean (a∣ ⁣ρ)≤(a′∣ ⁣ρ)\left(a\right|\left.\!\rho\right)\leq\left(a^{\prime}\right|\left.\!\rho\right) for every ρ∈St(A)\rho\in{\mathsf{St}}({\rm A}).

In a causal theory the norm of a state ρi∈St(A)\rho_{i}\in{\mathsf{St}}({\rm A}) is given by ∣ ⁣∣ρi∣ ⁣∣=(e∣ ⁣ρi)|\!|\rho_{i}|\!|=\left(e\right|\left.\!\rho_{i}\right). Accordingly, one can define the normalized state

In a causal theory one can always allow for rescaled preparations: conditionally to the outcome i∈Xi\in{\rm X} in the preparation-test {ρi}i∈X\{\rho_{i}\}_{i\in{\rm X}} we can say that we prepared the normalized state ρˉi\bar{\rho}_{i}. For this reason, every state in a causal theory is proportional to a normalized state.

The set of normalized states will be denoted by St1(A){\mathsf{St}}_{1}({\rm A}). Since the set of all states St(A){\mathsf{St}}({\rm A}) is closed in the operational norm, also the set of normalized states St1(A){\mathsf{St}}_{1}({\rm A}) is closed. Moreover, the set St1(A){\mathsf{St}}_{1}({\rm A}) is convex purification: this means that for every pair of normalized states ρ1,ρ2∈St1(A)\rho_{1},\rho_{2}\in{\mathsf{St}}_{1}({\rm A}) and for every probability p∈p\in the convex combination ρp=pρ1+(1−p)ρ2\rho_{p}=p\rho_{1}+(1-p)\rho_{2} is a normalized state. Operationally, the state ρp\rho_{p} is obtained by

performing a binary test with outcomes {1,2}\{1,2\} and outcome probabilities p1=pp_{1}=p and p2=1−pp_{2}=1-p

for outcome ii preparing ρi\rho_{i}, thus realizing the preparation-test {piρi}i=1,2\{p_{i}\rho_{i}\}_{i=1,2}

coarse-graining over the outcomes, thus obtaining ρp=pρ1+(1−p)ρ2\rho_{p}=p\rho_{1}+(1-p)\rho_{2}.

The step 2 (preparation of a state conditionally on the outcome of a previous test) is possible because the theory is causal purification.

The pure normalized states are the extreme points of the convex set St1(A){\mathsf{St}}_{1}({\rm A}). For a normalized state ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) we define the face identified by ρ\rho as follows:

The face identified by ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) is the set FρF_{\rho} of all normalized states σ∈St1(A)\sigma\in{\mathsf{St}}_{1}({\rm A}) such that ρ=pσ+(1−p)τ\rho=p\sigma+(1-p)\tau, for some non-zero probability p>0p>0 and some normalized state τ∈St1(A)\tau\in{\mathsf{St}}_{1}({\rm A}).

In other words, FρF_{\rho} is the set of all normalized states that show up in the convex decompositions of ρ\rho. Clearly, if φ\varphi is a pure state, then one has Fφ={φ}F_{\varphi}=\{\varphi\}. The opposite situation is that of completely mixed states: by definition 1, a state ω∈St1(A)\omega\in{\mathsf{St}}_{1}({\rm A}) is completely mixed if every state σ∈St1(A)\sigma\in{\mathsf{St}}_{1}({\rm A}) can stay in its convex decomposition, that is, if Fω=St1(A)F_{\omega}={\mathsf{St}}_{1}({\rm A}). An equivalent condition for a state to be completely mixed is the following:

Proof. The condition is clearly necessary. It is also sufficient because for a state σ∈St1(A)\sigma\in{\mathsf{St}}_{1}({\rm A}) the relation σ∈Span(Fω)\sigma\in\mathsf{Span}(F_{\omega}) implies σ∈Fω\sigma\in F_{\omega} (see Lemma 16 of Ref.purification).  ■\,\blacksquare

A completely mixed state can never be distinguished from another state with zero error probability:

Let ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) be a completely mixed state and σ∈St1(A)\sigma\in{\mathsf{St}}_{1}({\rm A}) be an arbitrary state. Then, the probability of error in distinguishing ρ\rho from σ\sigma is strictly greater than zero.

Proof. By contradiction, suppose that one can distinguish between ρ\rho and σ\sigma with zero error probability. This means that there exists a binary test {aρ,aσ}\{a_{\rho},a_{\sigma}\} such that (aρ∣ ⁣σ)=(aσ∣ ⁣ρ)=0\left(a_{\rho}\right|\left.\!\sigma\right)=\left(a_{\sigma}\right|\left.\!\rho\right)=0. Since ρ\rho is completely mixed there exists a probability p>0p>0 and a state τ∈St1(A)\tau\in{\mathsf{St}}_{1}({\rm A}) such that ρ=pσ+(1−p)τ\rho=p\sigma+(1-p)\tau. Hence, the condition (aσ∣ ⁣ρ)=0\left(a_{\sigma}\right|\left.\!\rho\right)=0 implies (aσ∣ ⁣σ)=0\left(a_{\sigma}\right|\left.\!\sigma\right)=0. Therefore, we have (aρ∣ ⁣σ)+(aσ∣ ⁣σ)=0\left(a_{\rho}\right|\left.\!\sigma\right)+\left(a_{\sigma}\right|\left.\!\sigma\right)=0. This is in contradiction with the normalization of the probabilities in the test {aρ,aσ}\{a_{\rho},a_{\sigma}\}, which would require (aρ∣ ⁣σ)+(aσ∣ ⁣σ)=1\left(a_{\rho}\right|\left.\!\sigma\right)+\left(a_{\sigma}\right|\left.\!\sigma\right)=1.  ■\,\blacksquare

III.1.2 Perfect distinguishability

Our second axiom regards the task of state discrimination. As we saw in proposition 1, if a state is completely mixed, then it is impossible to distinguish it perfectly from any other state. Axiom 2 states the converse:

Every state that is not completely mixed can be perfectly distinguished from some other state.

Note that the statement of axiom 2 holds for quantum and for classical information theory. In quantum theory a completely mixed state is a density matrix with full rank. If a density matrix ρ\rho has not full rank, then it must have a kernel: hence, every density matrix σ\sigma with support in the kernel of ρ\rho will be perfectly distinguishable from ρ\rho, as stated in Axiom 2. Applying the same reasoning for density matrices that are diagonal in a given basis, one can easily see that Axiom 2 is satisfied also by classical information theory.

III.1.3 Ideal compression

The third axiom is about information compression. An information source for system A{\rm A} is a preparation-test {ρi}i∈X\{\rho_{i}\}_{i\in{\rm X}}, where each ρi∈St(A)\rho_{i}\in{\mathsf{St}}({\rm A}) is an unnormalized state and ∑i∈X(e∣ ⁣ρi)=1\sum_{i\in{\rm X}}\left(e\right|\left.\!\rho_{i}\right)=1. A compression scheme is given by an encoding operation E\mathscr{E} from A{\rm A} to a smaller system C{\rm C}, that is, to a system C{\rm C} such that DC≤DAD_{\rm C}\leq D_{\rm A}. The compression scheme is lossless for the source {ρi}i∈X\{\rho_{i}\}_{i\in{\rm X}} if there exists a decoding operation D\mathscr{D} from C{\rm C} to A{\rm A} such that DE∣ρi)=∣ρi)\mathscr{D}\mathscr{E}\left|\rho_{i}\right)=\left|\rho_{i}\right) for every value of the index i∈Xi\in{\rm X}. This means that the decoding allows one to perfectly retrieve the states {ρi}i∈X\{\rho_{i}\}_{i\in{\rm X}}. We say that a compression scheme is lossless for the state ρ\rho, if it is lossless for every source {ρi}i∈X\{\rho_{i}\}_{i\in{\rm X}} such that ρ=∑i∈Xρi\rho=\sum_{i\in{\rm X}}\rho_{i}. Equivalently, this means that the restriction of DE\mathscr{D}\mathscr{E} to the face identified by ρ\rho is equal to the identity channel: DE∣σ)=σ\mathscr{D}\mathscr{E}\left|\sigma\right)=\sigma for every σ∈Fρ\sigma\in F_{\rho}.

A lossless compression scheme is maximally efficient if the encoding system C{\rm C} has the smallest possible size, that is, if the system C{\rm C} has no more states than exactly those needed to compress ρ\rho. This happens when every normalized state τ∈St1(C)\tau\in{\mathsf{St}}_{1}({\rm C}) comes from the encoding of some normalized state σ∈Fρ\sigma\in F_{\rho}, namely ∣τ)=E∣σ)\left|\tau\right)=\mathscr{E}\left|\sigma\right).

We say that a compression scheme that is lossless and maximally efficient is ideal. Our second axiom states that ideal compression is always possible:

For every state there exists an ideal compression scheme.

It is easy to see that this statement holds in quantum theory and in classical probability theory. For example, if ρ\rho is a density matrix on a dd-dimensional Hilbert space and rank(ρ)=r\mathsf{rank}(\rho)=r, then the ideal compression is obtained by just encoding ρ\rho in an rr-dimensional Hilbert space. As long as we do not tolerate losses, this is the most efficient one-shot compression we can devise in quantum theory. Similar observations hold for classical information theory.

III.1.4 Local distinguishability

The fourth axiom consists in the assumption of local distinguishability, here presented in the formulation of Ref. purification.

(Local distinguishability) If two bipartite states are different, then they give different probabilities for at least one product experiment.

In more technical terms: if ρ,σ∈St1(AB)\rho,\sigma\in{\mathsf{St}}_{1}({\rm A}{\rm B}) are states and ρ≠σ\rho\not=\sigma, then there are two effects a∈Eff(A)a\in{\mathsf{Eff}}({\rm A}) and b∈Eff(B)b\in{\mathsf{Eff}}({\rm B}) such that

Local distinguishability is equivalent to the fact that two distant parties, holding systems A{\rm A} and B{\rm B}, respectively, can distinguish between the two states ρ,σ∈St1(AB)\rho,\sigma\in{\mathsf{St}}_{1}({\rm A}{\rm B}) using only local operations and classical communication and achieving an error probability strictly larger than pran=1/2p_{ran}=1/2, the probability of error in random guess purification. Again, this statement holds in ordinary quantum theory (on complex Hilbert spaces) and in classical information theory.

Another equivalent condition to local distinguishability is the local tomography axiom, introduced in Refs. maurofirst; barrett. The local tomography axiom state that every bipartite state can be reconstructed from the statistics of local measurements on the component systems. Technically, local tomography is in turn equivalent to the relation DAB=DADBD_{{\rm A}{\rm B}}=D_{\rm A}D_{\rm B} Har01 and to the fact that every state ρ∈St(AB)\rho\in{\mathsf{St}}({\rm A}{\rm B}) can be written as

An important consequence of local distinguishability, observed in Ref. purification, is that a transformation C∈Transf(AB)\mathscr{C}\in{\mathsf{Transf}}({\rm A}{\rm B}) is completely specified by its action on St(A){\mathsf{St}}({\rm A}): thanks to local distinguishability we have the implication

(see Lemma 14 of Ref.purification for the proof). Note that Eq. (5) does not hold for quantum theory on real Hilbert spaces purification.

III.1.5 Pure conditioning

The fourth axiom states how the outcomes of a measurement on one side of a pure bipartite state can induce pure states on the other side. In this case we consider atomic measurements, that is, measurements described by observation-tests {ai}i∈X\{a_{i}\}_{i\in{\rm X}} where each effect aia_{i} is atomic. Intuitively, atomic measurement are those with maximum “resolving power”.

If a bipartite system is in a pure state, then each outcome of an atomic measurement on one side induces a pure state on the other.

The pure conditioning property holds in quantum theory and in classical information theory as well. In fact, the statement is trivial in classical information theory, because the only pure bipartite states are the product of pure states: no matter which measurement is performed on one side, the remaining state on the other side will necessarily be pure.

The pure conditioning property, as formulated above, has been recently introduced in Ref. pcond. A stronger version of axiom 5 is the atomicity of composition introduced in Ref. maurolast:

Atomicity of composition: the sequential composition of two atomic operations is atomic.

Since pure states and atomic effects are a particular case of atomic transformations, Axiom 5’ implies Axiom 5. In our derivation, however, also the converse implication holds: indeed, thanks to the purification postulate we will be able to show that Axiom 5 implies Axiom 5’ (see lemma 16).

III.2 The purification postulate

The last postulate in our list is the purification postulate, which was introduced and explored in detail in Ref. purification. While the previous axioms were also satisfied by classical probability theory, the purification axiom introduces in our derivation the genuinely quantum features. A purification of the state ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) is a pure state Ψρ\Psi_{\rho} of some composite system AB{\rm A}{\rm B}, with the property that ρ\rho is the marginal of Ψρ\Psi_{\rho}, that is,

Here we refer to the system B{\rm B} as the purifying system. The purification axiom states that every state can be obtained as the marginal of a pure bipartite state in an essentially unique way:

Every state has a purification. For fixed purifying system, every two purifications of the same state are connected by a reversible transformation on the purifying system.

Informally speaking, our postulate states that the ignorance about a part is always compatible with a maximal knowledge of the whole. The existence of pure bipartite states with mixed marginal was already recognized by Schrödinger as the characteristic trait of quantum theory Schr35. Here, however, we also emphasize the importance of the uniqueness of purification up to reversible transformations: this property sets up a relation between pure states and reversible transformations that generates most of the structure of quantum theory. As shown in Ref. purification, an impressive number of quantum features are actually direct consequences of purification. In particular, purification implies the possibility of simulating any irreversible process through a reversible interaction of the system with an environment that is finally discarded.

IV First consequences of the principles

Let ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) be a state and let E∈Transf(A,C)\mathscr{E}\in{\mathsf{Transf}}({\rm A},{\rm C}) (resp. D∈Transf(C,A)\mathscr{D}\in{\mathsf{Transf}}({\rm C},{\rm A})) be its encoding (resp. decoding) in the ideal compression scheme of Axiom 3.

Essentially, the encoding operation E∈Transf(A,C)\mathscr{E}\in{\mathsf{Transf}}({\rm A},{\rm C}) identifies the face FρF_{\rho} with the state space St1(C){\mathsf{St}}_{1}({\rm C}). In the following we provide a list of elementary lemmas showing that all statements about FρF_{\rho} can be translated into statements about St1(C){\mathsf{St}}_{1}({\rm C}) and vice-versa.

The composition of decoding and encoding is the identity on C{\rm C}, namely ED=IC\mathscr{E}\mathscr{D}=\mathscr{I}_{\rm C}.

Proof. Since the compression is maximally efficient, for every state τ∈St1(C)\tau\in{\mathsf{St}}_{1}({\rm C}) there is a state σ∈Fρ\sigma\in F_{\rho} such that Eσ=τ\mathscr{E}\sigma=\tau. Using the fact that DEσ=σ\mathscr{D}\mathscr{E}\sigma=\sigma (the compression is lossless) we then obtain EDτ=EDEσ=Eσ=τ\mathscr{E}\mathscr{D}\tau=\mathscr{E}\mathscr{D}\mathscr{E}\sigma=\mathscr{E}\sigma=\tau. By local distinguishability [see Eq. (5)], this implies ED=IC\mathscr{E}\mathscr{D}=\mathscr{I}_{\rm C}.  ■\,\blacksquare

The image of St1(C){\mathsf{St}}_{1}({\rm C}) under the decoding operation D\mathscr{D} is FρF_{\rho}.

Proof. Since the compression is maximally efficient, for all τ∈St1(C)\tau\in{\mathsf{St}}_{1}({\rm C}) there exists σ∈Fρ\sigma\in F_{\rho} such that τ=Eσ\tau=\mathscr{E}\sigma. Then, Dτ=DEσ=σ\mathscr{D}\tau=\mathscr{D}\mathscr{E}\sigma=\sigma. This implies that D(St1(C))⊆Fρ\mathscr{D}({\mathsf{St}}_{1}({\rm C}))\subseteq F_{\rho}. On the other hand, since the compression is lossless, for every state σ∈Fρ\sigma\in F_{\rho} one has DEσ=σ\mathscr{D}\mathscr{E}\sigma=\sigma. This implies the inclusion Fρ⊆D(St1(C))F_{\rho}\subseteq\mathscr{D}({\mathsf{St}}_{1}({\rm C})).  ■\,\blacksquare

If the state φ∈Fρ\varphi\in F_{\rho} is pure, then the state Eφ∈St1(C)\mathscr{E}\varphi\in{\mathsf{St}}_{1}({\rm C}) is pure. If the state ψ∈St1(C)\psi\in{\mathsf{St}}_{1}({\rm C}) is pure, then the state Dψ∈Fρ\mathscr{D}\psi\in F_{\rho} is pure.

Proof. Suppose that φ∈Fρ\varphi\in F_{\rho} is pure and that Eφ\mathscr{E}\varphi can be written as Eφ=pσ+(1−p)τ\mathscr{E}\varphi=p\sigma+(1-p)\tau for some p>0p>0 and some σ,τ∈St1(C)\sigma,\tau\in{\mathsf{St}}_{1}({\rm C}). Applying D\mathscr{D} on both sides we obtain φ=pDσ+(1−p)Dτ\varphi=p\mathscr{D}\sigma+(1-p)\mathscr{D}\tau. Since φ\varphi is pure we must have Dσ=Dτ=φ\mathscr{D}\sigma=\mathscr{D}\tau=\varphi. Now, applying E\mathscr{E} on all terms of the equality and using lemma 2 we obtain σ=τ=Eφ\sigma=\tau=\mathscr{E}\varphi. This proves that Eφ\mathscr{E}\varphi is pure. Conversely, suppose that ψ∈St1(C)\psi\in{\mathsf{St}}_{1}({\rm C}) is pure and Dψ=pσ+(1−p)τ\mathscr{D}\psi=p\sigma+(1-p)\tau for some p>0p>0 and some σ,τ∈St1(A)\sigma,\tau\in{\mathsf{St}}_{1}({\rm A}). Since Dψ\mathscr{D}\psi is in the face FρF_{\rho} (lemma 3), also σ\sigma and τ\tau are in the same face. Applying E\mathscr{E} on both sides of the equality Dψ=pσ+(1−p)τ\mathscr{D}\psi=p\sigma+(1-p)\tau and using lemma 2 we obtain ψ=EDψ=pEσ+(1−p)Eτ\psi=\mathscr{E}\mathscr{D}\psi=p\mathscr{E}\sigma+(1-p)\mathscr{E}\tau. Since ψ\psi is pure we must have Eσ=Eτ=ψ\mathscr{E}\sigma=\mathscr{E}\tau=\psi. Applying D\mathscr{D} on all terms of the equality we then have σ=τ=Dψ\sigma=\tau=\mathscr{D}\psi, thus proving that Dψ\mathscr{D}\psi is pure.  ■\,\blacksquare

We say that a state σ∈Fρ\sigma\in F_{\rho} is completely mixed relative to the face FρF_{\rho} if every state τ∈Fρ\tau\in F_{\rho} can stay in the convex decomposition of σ\sigma. In other words, σ\sigma is completely mixed relative to FρF_{\rho} if one has Fσ=FρF_{\sigma}=F_{\rho}. Note that in general σ∈Fρ\sigma\in F_{\rho} implies Fσ⊆FρF_{\sigma}\subseteq F_{\rho}.

If the state ω∈Fρ\omega\in F_{\rho} is completely mixed relative to FρF_{\rho}, then the state Eω∈St1(C)\mathscr{E}\omega\in{\mathsf{St}}_{1}({\rm C}) is completely mixed. If the state υ∈St1(C)\upsilon\in{\mathsf{St}}_{1}({\rm C}) is completely mixed, then the state Dυ∈Fρ\mathscr{D}\upsilon\in F_{\rho} is completely mixed relative to FρF_{\rho}.

Proof. Suppose that ω\omega is completely mixed relative to FρF_{\rho}. Then every state σ∈Fρ\sigma\in F_{\rho} can stay in its convex decomposition, say ω=pσ+(1−p)σ′\omega=p\sigma+(1-p)\sigma^{\prime} with p>0p>0 and σ′∈Fρ\sigma^{\prime}\in F_{\rho}. Applying E\mathscr{E} we have

Since the compression is maximally efficient, for every state τ∈St1(C)\tau\in{\mathsf{St}}_{1}({\rm C}) there exists a state σ∈Fρ\sigma\in F_{\rho} such that τ=Eσ\tau=\mathscr{E}\sigma. Choosing the suitable σ∈Fρ\sigma\in F_{\rho} and substituting τ\tau to Eσ\mathscr{E}\sigma in Eq. (6) we then obtain that for every state τ∈St1(C)\tau\in{\mathsf{St}}_{1}({\rm C}) there exists probability p>0p>0 and a state σ′∈Fρ\sigma^{\prime}\in F_{\rho} such that

This implies that Eω\mathscr{E}\omega is completely mixed. Suppose now that υ∈St1(C)\upsilon\in{\mathsf{St}}_{1}({\rm C}) is completely mixed. Then every state τ∈St1(C)\tau\in{\mathsf{St}}_{1}({\rm C}) can stay in its convex decomposition, say υ=pτ+(1−p)τ′\upsilon=p\tau+(1-p)\tau^{\prime}. with p>0p>0 and τ′∈St1(C)\tau^{\prime}\in{\mathsf{St}}_{1}({\rm C}). Applying D\mathscr{D} on both sides we have

Now, using lemma 3 we have that every state σ∈Fρ\sigma\in F_{\rho} can be written as σ=Dτ\sigma=\mathscr{D}\tau for some τ∈St1(C)\tau\in{\mathsf{St}}_{1}({\rm C}). Choosing the suitable τ∈St1(C)\tau\in{\mathsf{St}}_{1}({\rm C}) and substituting σ\sigma to Dτ\mathscr{D}\tau in Eq. (7) we then obtain that for evert state σ∈Fρ\sigma\in F_{\rho} there exists a probability p>0p>0 and a state τ′∈St1(C)\tau^{\prime}\in{\mathsf{St}}_{1}({\rm C}) such that Dυ=pσ+(1−p)Dτ′\mathscr{D}\upsilon=p\sigma+(1-p)\mathscr{D}\tau^{\prime}. Therefore, Dυ\mathscr{D}\upsilon is completely mixed relative to FρF_{\rho}.  ■\,\blacksquare

We now show that the system C{\rm C} used for ideal compression of the state ρ\rho is unique up to operational equivalence:

If two systems C{\rm C} and C′{\rm C}^{\prime} allow for ideal compression of a state ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}), then C{\rm C} and C′{\rm C}^{\prime} are operationally equivalent.

Proof. Let E,D\mathscr{E},\mathscr{D} and E′,D′\mathscr{E}^{\prime},\mathscr{D}^{\prime} denote the encoding/decoding schemes for systems C{\rm C} and C′{\rm C}^{\prime}, respectively. Define the transformations U:=E′D∈Transf(C,C′)\mathscr{U}:=\mathscr{E}^{\prime}\mathscr{D}\in{\mathsf{Transf}}({\rm C},{\rm C}^{\prime}) and V=ED′∈Transf(C′,C)\mathscr{V}=\mathscr{E}\mathscr{D}^{\prime}\in{\mathsf{Transf}}({\rm C}^{\prime},{\rm C}). It is easy to see that U\mathscr{U} is reversible and U−1=V\mathscr{U}^{-1}=\mathscr{V}. Indeed, since the restriction of D′E′\mathscr{D}^{\prime}\mathscr{E}^{\prime} and DE\mathscr{D}\mathscr{E} to the face FρF_{\rho} is the identity, using Lemma 3 one has D′E′D=D\mathscr{D}^{\prime}\mathscr{E}^{\prime}\mathscr{D}=\mathscr{D} and similarly DED′=D′\mathscr{D}\mathscr{E}\mathscr{D}^{\prime}=\mathscr{D}^{\prime}. Hence, we have UV=E′DED′=E′D′=IC′\mathscr{U}\mathscr{V}=\mathscr{E}^{\prime}\mathscr{D}\mathscr{E}\mathscr{D}^{\prime}=\mathscr{E}^{\prime}\mathscr{D}^{\prime}=\mathscr{I}_{{\rm C}^{\prime}} and VU=ED′E′D=ED=IC\mathscr{V}\mathscr{U}=\mathscr{E}\mathscr{D}^{\prime}\mathscr{E}^{\prime}\mathscr{D}=\mathscr{E}\mathscr{D}=\mathscr{I}_{{\rm C}}. ■\,\blacksquare

It is useful to introduce the notion of equality upon input of ρ\rho. We say that two transformations A,A′∈Transf(A,B)\mathscr{A},\mathscr{A}^{\prime}\in{\mathsf{Transf}}({\rm A},{\rm B}) are equal upon input of ρ∈St(A)\rho\in{\mathsf{St}}({\rm A}) if their restrictions to the face identified by ρ\rho are equal, that is, if Aσ=A′σ\mathscr{A}\sigma=\mathscr{A}^{\prime}\sigma for every σ∈Fρ\sigma\in F_{\rho}. If A\mathscr{A} and A′\mathscr{A}^{\prime} are equal upon input of ρ\rho we write A=ρA′\mathscr{A}=_{\rho}\mathscr{A}^{\prime}.

Using the notion of equality upon input of ρ\rho we can rephrase the fact that the compression is lossless for ρ\rho as DE=ρIA\mathscr{D}\mathscr{E}=_{\rho}\mathscr{I}_{\rm A}. Similarly, we can state the following:

The encoding E\mathscr{E} is deterministic upon input of ρ\rho, that is (eC∣E=ρ(eA∣\left(e_{\rm C}\right|\mathscr{E}=_{\rho}\left(e_{\rm A}\right|.

Proof. For every σ∈Fρ\sigma\in F_{\rho} we have (eC∣E∣σ)≥(eA∣DE∣σ)=(eA∣ ⁣σ)=1\left(e_{\rm C}\right|\mathscr{E}\left|\sigma\right)\geq\left(e_{\rm A}\right|\mathscr{D}\mathscr{E}\left|\sigma\right)=\left(e_{\rm A}\right|\left.\!\sigma\right)=1, having used Eq. (4) and the fact that the compression is lossless. Since probabilities are bounded by 1, this implies (eC∣E∣σ)=(eA∣ ⁣σ)\left(e_{\rm C}\right|\mathscr{E}\left|\sigma\right)=\left(e_{\rm A}\right|\left.\!\sigma\right) for every σ∈Fρ\sigma\in F_{\rho}, that is, (eC∣E=ρ(eA∣\left(e_{\rm C}\right|\mathscr{E}=_{\rho}\left(e_{\rm A}\right|.  ■\,\blacksquare

The decoding D\mathscr{D} is deterministic, that is (eA∣D=(eC∣\left(e_{\rm A}\right|\mathscr{D}=\left(e_{\rm C}\right|.

Proof. For every τ∈St1(A)\tau\in{\mathsf{St}}_{1}({\rm A}) we have (eA∣D∣τ)≥(eC∣ED∣τ)=(eC∣ ⁣τ)\left(e_{\rm A}\right|\mathscr{D}\left|\tau\right)\geq\left(e_{\rm C}\right|\mathscr{E}\mathscr{D}\left|\tau\right)=\left(e_{\rm C}\right|\left.\!\tau\right), having used Eq. (4) and lemma 2. Hence, (eA∣D=(eC∣\left(e_{\rm A}\right|\mathscr{D}=\left(e_{\rm C}\right|.  ■\,\blacksquare

IV.2 Results about purification

The purification postulate 1 implies a large number of quantum features, as it was shown in Ref. purification. Here we review only the facts that are useful for our derivation, referring to Ref. purification for the proofs.

An elementary consequence of the uniqueness of purification is that the group GA{\mathbf{G}}_{\rm A} of reversible transformations on A{\rm A} acts transitively on the set of pure states:

For every couple of pure states φ,φ′∈St1(A)\varphi,\varphi^{\prime}\in{\mathsf{St}}_{1}({\rm A}) there is a reversible transformation U∈GA\mathscr{U}\in{\mathbf{G}}_{\rm A} such that φ′=Uφ\varphi^{\prime}=\mathscr{U}\varphi.

Proof. See Lemma 20 of Ref. purification.  ■\,\blacksquare

Transitivity implies that for every system A{\rm A} there is a unique state χA∈St1(A)\chi_{\rm A}\in{\mathsf{St}}_{1}({\rm A}) that is invariant under reversible transformations, that is, a unique state such that UχA=χA\mathscr{U}\chi_{\rm A}=\chi_{\rm A} for every U∈GA\mathscr{U}\in{\mathbf{G}}_{\rm A}:

For every system A{\rm A}, there is a unique state χA\chi_{\rm A} invariant under all reversible transformations in GA{\mathbf{G}}_{\rm A}. The invariant state has the following properties:

χAB=χA⊗χB\chi_{{\rm A}{\rm B}}=\chi_{\rm A}\otimes\chi_{\rm B}.

Proof. See Corollary 34 and Theorem 4 of Ref. purification. The proof of item 2 uses the local distinguishability axiom.  ■\,\blacksquare

When there is no ambiguity we will drop the subindex A{\rm A} and simply write χ\chi.

The uniqueness of purification in postulate 1 requires that if Ψρ,Ψρ′∈St1(AB)\Psi_{\rho},\Psi_{\rho}^{\prime}\in{\mathsf{St}}_{1}({\rm A}{\rm B}) are two purifications of ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}), then there exists a reversible transformation U∈GB\mathscr{U}\in{\mathbf{G}}_{\rm B} such that Ψρ′=(IA⊗U)Ψρ\Psi_{\rho}^{\prime}=(\mathscr{I}_{\rm A}\otimes\mathscr{U})\Psi_{\rho}. The following lemma extends the uniqueness property to purifications with different purifying systems:

(Uniqueness of the purification up to channels on the purifying systems) Let Ψ∈St1(AB)\Psi\in{\mathsf{St}}_{1}({\rm A}{\rm B}) and Ψ′∈St1(AC)\Psi^{\prime}\in{\mathsf{St}}_{1}({\rm A}{\rm C}) be two purifications of ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}). Then there exists a channel C∈Transf(B,C)\mathscr{C}\in{\mathsf{Transf}}({\rm B},{\rm C}) such that

Proof. See Lemma 21 of Ref. purification.  ■\,\blacksquare

Another consequence of the uniqueness of purification is the fact that any ensemble decomposition of a given mixed state can be obtained by performing a measurement on the purifying system:

(Purification of preparation-tests) Let ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) be a state and Ψρ∈St1(AB)\Psi_{\rho}\in{\mathsf{St}}_{1}({\rm A}{\rm B}) be a purification of ρ\rho. If {ρi}i∈X\{\rho_{i}\}_{i\in{\rm X}} be a preparation-test such that ∑i∈Xρi=ρ\sum_{i\in{\rm X}}\rho_{i}=\rho, then there exists an observation-test {ai}i∈X\{a_{i}\}_{i\in{\rm X}} on the purifying system such that

Proof. See lemma 8 of Ref. purification.  ■\,\blacksquare

If Ψρ∈St1(AB)\Psi_{\rho}\in{\mathsf{St}}_{1}({\rm A}{\rm B}) is a purification of ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) and σ\sigma belongs to the face FρF_{\rho}, then there exists an effect bb and a non-zero probability p>0p>0 such that

An important consequence of purification and local distinguishability is the relation between equality upon input of ρ\rho and equality on the purifications of ρ\rho:

(Equality upon input of ρ\rho vs equality on purifications of ρ\rho) Let Ψ∈St1(AC)\Psi\in{\mathsf{St}}_{1}({\rm A}{\rm C}) be a purification of ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}), and let A,A′∈Transf(A,B)\mathscr{A},\mathscr{A}^{\prime}\in{\mathsf{Transf}}({\rm A},{\rm B}) be two transformations. Then one has

Proof. See theorem 1 of Ref. purification. The proof of the direction ⟸\Longleftarrow uses the local distinguishability axiom.  ■\,\blacksquare

As a consequence, the purification of a completely mixed state allows for the tomography of transformations:

Let ω∈St1(A)\omega\in{\mathsf{St}}_{1}({\rm A}) be completely mixed and Ψω∈St1(AC)\Psi_{\omega}\in{\mathsf{St}}_{1}({\rm A}{\rm C}) is a purification of ω\omega. Then, for all transformations A,A′∈Transf(A,B)\mathscr{A},\mathscr{A}^{\prime}\in{\mathsf{Transf}}({\rm A},{\rm B}) one has

Proof. By theorem 1 the first condition is equivalent to A=ωA′\mathscr{A}=_{\omega}\mathscr{A}^{\prime}. Since ω\omega is completely mixed, this means Aσ=A′σ\mathscr{A}\sigma=\mathscr{A}^{\prime}\sigma for every σ∈St1(A)\sigma\in{\mathsf{St}}_{1}({\rm A}). By local distinguishability [see Eq. (5)] this implies A=A′\mathscr{A}=\mathscr{A}^{\prime}.  ■\,\blacksquare

Corollary 2 shows that the state (A⊗IC)Ψω(\mathscr{A}\otimes\mathscr{I}_{\rm C})\Psi_{\omega} characterizes the transformation A\mathscr{A} completely. We will express this fact by saying that the state Ψω\Psi_{\omega} is dynamically faithful maurolast, or just faithful, for short. Using this notion we can rephrase corollary 2 as:

If Ψ∈St1(AC)\Psi\in{\mathsf{St}}_{1}({\rm A}{\rm C}) is pure and its marginal on system A{\rm A} is completely mixed, then Ψ\Psi is dynamically faithful for system A{\rm A}.

Let us choose a fixed faithful state for system A{\rm A}, say Ψ∈St1(AC)\Psi\in{\mathsf{St}}_{1}({\rm A}{\rm C}). Then for every transformation C∈Transf(A,B)\mathscr{C}\in{\mathsf{Transf}}({\rm A},{\rm B}) we can define the Choi state RC∈St(BC)R_{\mathscr{C}}\in{\mathsf{St}}({\rm B}{\rm C}) as

(Choi isomorphism) For a given faithful state Ψ∈St1(AC)\Psi\in{\mathsf{St}}_{1}({\rm A}{\rm C}) the map C↦RC:=(C⊗IC)Ψ\mathscr{C}\mapsto R_{\mathscr{C}}:=(\mathscr{C}\otimes\mathscr{I}_{\rm C})\Psi has the following properties:

it defines a bijective correspondence between tests {Ci}i∈X\{\mathscr{C}_{i}\}_{i\in{\rm X}} from A{\rm A} to B{\rm B} and collections of states {Ri}i∈X\{R_{i}\}_{i\in{\rm X}} for BC{\rm B}{\rm C} satisfying

the transformation C\mathscr{C} is atomic if and only if the corresponding state RCR_{\mathscr{C}} is pure.

Proof. See Theorem 17 of Ref. purification.

A simple consequence of the Choi isomorphism is the following:

Let {Ci}i∈X⊂Transf(A,B)\{\mathscr{C}_{i}\}_{i\in{\rm X}}\subset{\mathsf{Transf}}({\rm A},{\rm B}) be a collection of transformations. Then, {Ci}i∈X\{\mathscr{C}_{i}\}_{i\in{\rm X}} is a test if and only if

In particular, let {ai}i∈X⊂Eff(A)\{a_{i}\}_{i\in{\rm X}}\subset{\mathsf{Eff}}({\rm A}) be a collection of effects. Then, {ai}i∈X\{a_{i}\}_{i\in{\rm X}} is an observation-test if and only if

Proof. Apply item 1 of theorem 2 to the collection of states {Ri}i∈X\{R_{i}\}_{i\in{\rm X}} defined by Ri:=(Ci⊗IC)ΨR_{i}:=(\mathscr{C}_{i}\otimes\mathscr{I}_{\rm C}){\Psi}.  ■\,\blacksquare

A much deeper consequence of the Choi isomorphism is the following theorem:

(States specify the theory) Let Θ,Θ′\Theta,\Theta^{\prime} be two theories satisfying the purification postulate. If Θ\Theta and Θ′\Theta^{\prime} have the same sets of normalized states, then Θ′=Θ\Theta^{\prime}=\Theta.

Proof. See Theorem 19 of Ref.purification. ■\,\blacksquare

Thanks to theorem 3 to derive quantum theory we will only need to prove that our principles imply that for every system A{\rm A} the normalized states St1(A){\mathsf{St}}_{1}({\rm A}) can be described as positive Hermitian matrices with unit trace. Once this is proved, theorem 3 automatically ensures that all the dynamics and all the measurements allowed by the theory are exactly the dynamics and the measurements allowed in quantum theory.

Note that in the definition of the Choi state we left the freedom to choose the faithful state Ψ∈St1(AC)\Psi\in{\mathsf{St}}_{1}({\rm A}{\rm C}). Among many possibilities, one convenient choice is to take a faithful state Φ∈St1(AC)\Phi\in{\mathsf{St}}_{1}({\rm A}{\rm C}) obtained as a purification of the invariant state χ∈St1(A)\chi\in{\mathsf{St}}_{1}({\rm A}). Moreover, as we will see in the next paragraph, we can always choose the purifying system C{\rm C} in such a way that the marginal on C{\rm C} is completely mixed.

IV.3 Results about the combination of compression and purification

An important consequence of the combination of the purification postulate with the compression axiom is the fact that one can always choose a purification of ρ\rho such that the marginal state on the purifying system is completely mixed. To prove this result we need the following lemma:

Let ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) be a state and let Ψρ∈St1(AB)\Psi_{\rho}\in{\mathsf{St}}_{1}({\rm A}{\rm B}) be a purification of ρ\rho. If E∈Transf(A,C)\mathscr{E}\in{\mathsf{Transf}}({\rm A},{\rm C}) is the encoding operation in the compression scheme of axiom 3, then the state Ψρ′:=(E⊗IB)Ψρ\Psi_{\rho}^{\prime}:=(\mathscr{E}\otimes\mathscr{I}_{\rm B})\Psi_{\rho} is pure.

Proof. Let D∈Transf(C,A)\mathscr{D}\in{\mathsf{Transf}}({\rm C},{\rm A}) be the decoding operation. Since the compression is lossless for ρ\rho we know that DE=ρIA\mathscr{D}\mathscr{E}=_{\rho}\mathscr{I}_{\rm A}. By theorem 1 this is equivalent to the condition (DE⊗IB)Ψρ=Ψρ(\mathscr{D}\mathscr{E}\otimes\mathscr{I}_{\rm B})\Psi_{\rho}=\Psi_{\rho}. Now, suppose that (E⊗IB)Ψρ=∑i∈XΓi(\mathscr{E}\otimes\mathscr{I}_{\rm B})\Psi_{\rho}=\sum_{i\in{\rm X}}\Gamma_{i}. Applying D\mathscr{D} on both sides we then obtain Ψρ=∑i∈X(D⊗IB)Γi\Psi_{\rho}=\sum_{i\in{\rm X}}(\mathscr{D}\otimes\mathscr{I}_{\rm B})\Gamma_{i}, and, since Ψρ\Psi_{\rho} is pure, for every i∈Xi\in{\rm X} we must have (D⊗IB)Γi=piΨρ(\mathscr{D}\otimes\mathscr{I}_{\rm B})\Gamma_{i}=p_{i}\Psi_{\rho}, where pi≥0p_{i}\geq 0 is some probability. Finally, since ED=IC\mathscr{E}\mathscr{D}=\mathscr{I}_{{\rm C}} (lemma 2), one has Γi=pi(E⊗IB)Ψρ\Gamma_{i}=p_{i}(\mathscr{E}\otimes\mathscr{I}_{\rm B})\Psi_{\rho}. Hence, (E⊗IB)Ψρ(\mathscr{E}\otimes\mathscr{I}_{\rm B})\Psi_{\rho} admits only decompositions with Γi=pi(E⊗IB)Ψρ\Gamma_{i}=p_{i}(\mathscr{E}\otimes\mathscr{I}_{\rm B})\Psi_{\rho}, that is, (E⊗IB)Ψρ(\mathscr{E}\otimes\mathscr{I}_{\rm B})\Psi_{\rho} is pure. ■\,\blacksquare

We are now in position to prove the desired result:

For every state ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) there exists a system C{\rm C} and a purification Ψρ∈St1(AC)\Psi_{\rho}\in{\mathsf{St}}_{1}({\rm A}{\rm C}) of ρ\rho such that the marginal state on system C{\rm C} is completely mixed. Moreover, the system C{\rm C} is unique up to operational equivalence.

Let Ψρ∈St1(AB)\Psi_{\rho}\in{\mathsf{St}}_{1}({\rm A}{\rm B}) be a purification of ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) and let E∈Transf(A,C)\mathscr{E}\in{\mathsf{Transf}}({\rm A},{\rm C}) be the encoding for ρ\rho. Then, the state (E⊗IB)Ψρ∈St1(CB)(\mathscr{E}\otimes\mathscr{I}_{\rm B})\Psi_{\rho}\in{\mathsf{St}}_{1}({\rm C}{\rm B}) is dynamically faithful for C{\rm C}.

Proof. The marginal of (E⊗IB)Ψρ(\mathscr{E}\otimes\mathscr{I}_{\rm B})\Psi_{\rho} on system C{\rm C} is Eρ\mathscr{E}\rho, which is completely mixed by lemma 5. Hence, (E⊗IB)Ψρ(\mathscr{E}\otimes\mathscr{I}_{\rm B})\Psi_{\rho} is dynamically faithful by corollary 3.  ■\,\blacksquare

The decoding transformation D∈Transf(C,A)\mathscr{D}\in{\mathsf{Transf}}({\rm C},{\rm A}) in the ideal compression for ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) is atomic.

Proof. Let Ψρ∈St1(AB)\Psi_{\rho}\in{\mathsf{St}}_{1}({\rm A}{\rm B}) be a purification of ρ\rho, for some purifying system B{\rm B}. Since DE=ρIA\mathscr{D}\mathscr{E}=_{\rho}\mathscr{I}_{\rm A} (the compression is lossless), we have (DE⊗IB)∣Ψρ)=∣Ψρ)(\mathscr{D}\mathscr{E}\otimes\mathscr{I}_{{\rm B}})|\Psi_{\rho})=|\Psi_{\rho}) (theorem 1). Now, by corollary 5 (E⊗IB)∣Ψρ)(\mathscr{E}\otimes\mathscr{I}_{\rm B})|\Psi_{\rho}) is faithful for C{\rm C} and by lemma 13 (E⊗IB)∣Ψρ)(\mathscr{E}\otimes\mathscr{I}_{\rm B})|\Psi_{\rho}) is pure. Using the Choi isomorphism with the faithful state Ψ:=(E⊗IB)Ψρ\Psi:=(\mathscr{E}\otimes\mathscr{I}_{\rm B})\Psi_{\rho} we then obtain that D\mathscr{D} is atomic.  ■\,\blacksquare

IV.4 Teleportation and the link product

Proof. See Corollary 19 of Ref. purification.  ■\,\blacksquare

Let us choose Ψ(A)\Psi^{({\rm A})} to be the faithful state in the definition of the Choi isomorphism. Then the sequential composition of transformation induces a composition of Choi states in following way:

For two transformations C∈Transf(A,B)\mathscr{C}\in{\mathsf{Transf}}({\rm A},{\rm B}) and D∈Transf(B,C)\mathscr{D}\in{\mathsf{Transf}}({\rm B},{\rm C}) the Choi state of DC∈Transf(A,C)\mathscr{D}\mathscr{C}\in{\mathsf{Transf}}({\rm A},{\rm C}) is given by the link product

Proof. See Corollary 22 of Ref. purification.  ■\,\blacksquare

We conclude this paragraph with an important result that follows from the combination of the link product structure with the pure conditioning axiom:

The composition of two atomic transformations is atomic.

Proof. Let C∈Transf(A,B)\mathscr{C}\in{\mathsf{Transf}}({\rm A},{\rm B}) and D∈Transf(B,C)\mathscr{D}\in{\mathsf{Transf}}({\rm B},{\rm C}) be two atomic transformations. By the Choi isomorphism, the (unnormalized) states RCR_{\mathscr{C}} and RDR_{\mathscr{D}} are pure. Since the teleportation effect E(B)E^{({\rm B})} in Eq. (9) is atomic (lemma 15), the pure conditioning axiom 5 implies the state RDCR_{\mathscr{D}\mathscr{C}} is pure. By the Choi isomorphism this means that DC\mathscr{D}\mathscr{C} is atomic.  ■\,\blacksquare

IV.5 No information without disturbance

We say that a test {Ci}i∈X⊂Transf(A)\{\mathscr{C}_{i}\}_{i\in{\rm X}}\subset{\mathsf{Transf}}({\rm A}) is non-disturbing upon input of ρ\rho if ∑i∈XCi=ρIA\sum_{i\in{\rm X}}\mathscr{C}_{i}=_{\rho}\mathscr{I}_{\rm A}. If ρ\rho is completely mixed, we simply say that the test is non-disturbing.

A consequence of the purification postulate is the following “no-information without disturbance” result:

A test {Ci}i∈X⊂Transf(A)\{\mathscr{C}_{i}\}_{i\in{\rm X}}\subset{\mathsf{Transf}}({\rm A}) is non-disturbing upon input of ρ\rho if and only if there is a set of probabilities {pi}i∈X\{p_{i}\}_{i\in{\rm X}} such that Ci=ρpiIA\mathscr{C}_{i}=_{\rho}p_{i}\mathscr{I}_{\rm A} for every i∈Xi\in{\rm X}.

Proof. See Theorem 10 of Ref. purification.  ■\,\blacksquare

The no-information without disturbance result implies the following geometrical limitation

For every system A{\rm A} the convex set of states St1(A){\mathsf{St}}_{1}({\rm A}) is not a segment.

Proof. The proof is by contradiction. Suppose that for some system A{\rm A} the set St1(A){\mathsf{St}}_{1}({\rm A}) is a segment. The segment has only two pure states, say φ1\varphi_{1} and φ2\varphi_{2}, and every other state ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) is completely mixed. Then the distinguishability axiom 2 imposes that φ1\varphi_{1} and φ2\varphi_{2} are perfectly distinguishable. Take the binary test {a1,a2}\{a_{1},a_{2}\} such that (ai∣ ⁣φj)=δij\left(a_{i}\right|\left.\!\varphi_{j}\right)=\delta_{ij} and define the “measure-and-prepare” test {C1,C2}\{\mathscr{C}_{1},\mathscr{C}_{2}\} as Ci=∣φi)(ai∣\mathscr{C}_{i}=\left|\varphi_{i}\right)\left(a_{i}\right|, i=1,2i=1,2 (the possibility of preparing a state depending on the outcome of a previous measurement is guaranteed by causality purification). Since every state ρ\rho in the segment can be written as convex combination of the two extreme points, we have that the test {C1,C2}\{\mathscr{C}_{1},\mathscr{C}_{2}\} is non-disturbing: (C1+C2)ρ=ρ(\mathscr{C}_{1}+\mathscr{C}_{2})\rho=\rho for every ρ\rho. This is in contradiction with lemma 17 because C1\mathscr{C}_{1} and C2\mathscr{C}_{2} are not proportional to the identity.  ■\,\blacksquare

We know that no information can be extracted without disturbance. In the following we will prove a result in the converse direction: if a measurement extracts no information, than it can be realized in a non-disturbing fashion. To show this result we first need the following

For every observation test {ai}i∈X⊂Eff(A)\{a_{i}\}_{i\in{\rm X}}\subset{\mathsf{Eff}}({\rm A}) with finite outcome set X{\rm X} there is a system C{\rm C} and a test {Ai}i∈X⊂Transf(A,C)\{\mathscr{A}_{i}\}_{i\in{\rm X}}\subset{\mathsf{Transf}}({\rm A},{\rm C}) consisting of atomic transformations such that (ai∣=(eC∣Ai\left(a_{i}\right|=\left(e_{\rm C}\right|\mathscr{A}_{i}.

Proof. Let ∣Ψ)AB\left|\Psi\right)_{{\rm A}{\rm B}} be a pure faithful state for system A{\rm A} and let ∣Ri)B=(ai∣A∣Ψ)AB\left|R_{i}\right)_{\rm B}=\left(a_{i}\right|_{\rm A}\left|\Psi\right)_{{\rm A}{\rm B}} the Choi state of aia_{i}. Take a purification of RiR_{i}, say ∣Ψi)BC\left|\Psi_{i}\right)_{{\rm B}{\rm C}} for some purifying system C{\rm C} singlepurifying. Then, by the Choi isomorphism there is a test {Ai}i∈X\{\mathscr{A}_{i}\}_{i\in{\rm X}}, with input A{\rm A} and output C{\rm C}, such that

[see item 1 of theorem 2] Moreover, each transformation Ai:A→C\mathscr{A}_{i}:{\rm A}\to{\rm C} is atomic (item 2 of theorem 2). Applying the deterministic effect (eC∣\left(e_{\rm C}\right| on both sides we then obtain ∣Ri)B=(eC∣∣Ψi)CA=(eC∣Ai∣Ψ)AB\left|R_{i}\right)_{\rm B}=\left(e_{\rm C}\right|\left|\Psi_{i}\right)_{{\rm C}{\rm A}}=\left(e_{\rm C}\right|\mathscr{A}_{i}\left|\Psi\right)_{{\rm A}{\rm B}}. By definition of RiR_{i}, this implies (ai∣A∣Ψ)AB=(eC∣Ai∣Ψ)AB\left(a_{i}\right|_{{\rm A}}\left|\Psi\right)_{{\rm A}{\rm B}}=\left(e_{\rm C}\right|\mathscr{A}_{i}\left|\Psi\right)_{{\rm A}{\rm B}}, and, since Ψ\Psi is dynamically faithful, (ai∣A=(eC∣Ai\left(a_{i}\right|_{{\rm A}}=\left(e_{\rm C}\right|\mathscr{A}_{i}. ■\,\blacksquare

Let ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) be a state, a∈Eff(A)a\in{\mathsf{Eff}}({\rm A}) be an effect, and A∈Transf(A,B)\mathscr{A}\in{\mathsf{Transf}}({\rm A},{\rm B}) be an atomic transformation such that (a∣A=(e∣BA\left(a\right|_{\rm A}=\left(e\right|_{\rm B}\mathscr{A}. If (a∣=ρp(e∣\left(a\right|=_{\rho}p\left(e\right| for some p≥0p\geq 0, then there exists a channel C∈Transf(B,A)\mathscr{C}\in{\mathsf{Transf}}({\rm B},{\rm A}) such that CA=ρpIA\mathscr{C}\mathscr{A}=_{\rho}p\mathscr{I}_{\rm A}.

Proof. Consider a purification of ρ\rho, say Ψρ∈St(AC)\Psi_{\rho}\in{\mathsf{St}}({\rm A}{\rm C}), and define the state Σ∈St1(BC)\Sigma\in{\mathsf{St}}_{1}({\rm B}{\rm C}) by ∣Σ):=1p(A⊗IC)∣Ψρ)|\Sigma):=\frac{1}{p}(\mathscr{A}\otimes\mathscr{I}_{\rm C})|\Psi_{\rho}). By the atomicity of composition 16 the state Σ\Sigma is pure. Moreover, we have

having used theorem 1 in the last equality. This implies that Ψρ\Psi_{\rho} and Σ\Sigma are different purifications of the same mixed state on system C{\rm C}. Then, by lemma 11 there exists a channel C∈Transf(B,A)\mathscr{C}\in{\mathsf{Transf}}({\rm B},{\rm A}) such that ∣Ψρ)=(C⊗IC)∣Σ)=1p(CA⊗IC)∣Ψρ)\left|\Psi_{\rho}\right)=(\mathscr{C}\otimes\mathscr{I}_{\rm C})|\Sigma)=\frac{1}{p}(\mathscr{C}\mathscr{A}\otimes\mathscr{I}_{\rm C})|\Psi_{\rho}). By theorem 1, the last equality implies CA=ρpIA\mathscr{C}\mathscr{A}=_{\rho}p\mathscr{I}_{\rm A}.  ■\,\blacksquare

We now make a simple observation that combined with Theorem 5 will lead to some interesting consequences:

If (a∣ρ)=∥a∥(a|\rho)=\|a\|, then a=ρ∥a∥ea=_{\rho}\|a\|e. Similarly, if (a∣ρ)=0(a|\rho)=0, then a=ρ0a=_{\rho}0.

Proof. By definition, σ∈Fρ\sigma\in F_{\rho} iff there exists p>0p>0 and τ∈St1(A)\tau\in{\mathsf{St}}_{1}({\rm A}) such that ρ=pσ+(1−p)τ\rho=p\sigma+(1-p)\tau. If (a∣ ⁣ρ)=∥a∥\left(a\right|\left.\!\rho\right)=\|a\|, then we have ∥a∥=p(a∣ ⁣σ)+(1−p)(a∣ ⁣τ)\|a\|=p\left(a\right|\left.\!\sigma\right)+(1-p)\left(a\right|\left.\!\tau\right). Since (a∣ ⁣σ)\left(a\right|\left.\!\sigma\right) and (a∣ ⁣τ)\left(a\right|\left.\!\tau\right) cannot be larger than ∥a∥\|a\|, the only way to have the equality is to have (a∣ ⁣σ)=(a∣ ⁣τ)=∥a∥\left(a\right|\left.\!\sigma\right)=\left(a\right|\left.\!\tau\right)=\|a\|. By definition, this amounts to say a=ρ∥a∥ea=_{\rho}\|a\|e. Similarly, if (a∣ ⁣ρ)=0\left(a\right|\left.\!\rho\right)=0, one has 0=p(a∣ ⁣σ)+(1−p)(a∣ ⁣τ)0=p\left(a\right|\left.\!\sigma\right)+(1-p)\left(a\right|\left.\!\tau\right), which is satisfied only if (a∣ ⁣σ)=(a∣ ⁣τ)=0\left(a\right|\left.\!\sigma\right)=\left(a\right|\left.\!\tau\right)=0, that is, if a=ρ0a=_{\rho}0.  ■\,\blacksquare

Let ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) be a state, a∈Eff(A)a\in{\mathsf{Eff}}({\rm A}) be an effect, and A∈Transf(A,B)\mathscr{A}\in{\mathsf{Transf}}({\rm A},{\rm B}) be an atomic transformation such that (a∣A=(e∣BA\left(a\right|_{{\rm A}}=(e|_{\rm B}\mathscr{A}. If (a∣ρ)=1(a|\rho)=1, then A\mathscr{A} is correctable upon input of ρ\rho, that is, there exists a correction operation C∈Transf(B,A)\mathscr{C}\in{\mathsf{Transf}}({\rm B},{\rm A}) such that CA=ρIA\mathscr{C}\mathscr{A}=_{\rho}\mathscr{I}_{\rm A}.

Proof. If (a∣ ⁣ρ)=1\left(a\right|\left.\!\rho\right)=1, then clearly ∥a∥=1\|a\|=1. Lemma 19 then implies (a∣=ρ(e∣\left(a\right|=_{\rho}\left(e\right|. Applying theorem 5 we finally obtain the thesis.  ■\,\blacksquare

Let ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) be a state, a∈Eff(A)a\in{\mathsf{Eff}}({\rm A}) be an effect such that (a∣ρ)=1(a|\rho)=1. Then there exists a transformation C∈Transf(A)\mathscr{C}\in{\mathsf{Transf}}({\rm A}) such that (a∣=(e∣C\left(a\right|=\left(e\right|\mathscr{C} and C=ρI\mathscr{C}=_{\rho}\mathscr{I}.

Proof. Straightforward consequence of lemma 18 and of corollary 8.  ■\,\blacksquare

Finally, we say that an observation-test {ai}i∈X\{a_{i}\}_{i\in{\rm X}} is non-informative upon input of ρ\rho if we have (ai∣=ρpi (e∣\left(a_{i}\right|=_{\rho}p_{i}~\left(e\right| for every i∈Xi\in{\rm X}. This means that the test {ai}i∈X\{a_{i}\}_{i\in{\rm X}} is unable to distinguish the states in the face FρF_{\rho}. As a consequence of theorem 5 we have the following “no disturbance without information” result:

If the test {ai}i∈X\{a_{i}\}_{i\in{\rm X}} is non-informative upon input of ρ\rho then there is a test {Di}i∈X⊂Transf(A)\{\mathscr{D}_{i}\}_{i\in{\rm X}}\subset{\mathsf{Transf}}({\rm A}) that is non-disturbing upon input of ρ\rho and satisfies (e∣Di=(ai∣\left(e\right|\mathscr{D}_{i}=\left(a_{i}\right| for every i∈Xi\in{\rm X}.

Proof. By lemma 18 there exists a test {Ai}⊂Transf(A,B)\{\mathscr{A}_{i}\}\subset{\mathsf{Transf}}({\rm A},{\rm B}) such that each transformation Ai\mathscr{A}_{i} is atomic and (e∣Ai=(ai∣\left(e\right|\mathscr{A}_{i}=\left(a_{i}\right|. By theorem 5, for each Ai\mathscr{A}_{i} there is a correction channel Ci\mathscr{C}_{i} such that CiAi=ρpiIA\mathscr{C}_{i}\mathscr{A}_{i}=_{\rho}p_{i}\mathscr{I}_{\rm A}. Defining Di:=CiAi\mathscr{D}_{i}:=\mathscr{C}_{i}\mathscr{A}_{i} we then obtain the thesis.  ■\,\blacksquare

V Perfectly distinguishable states

In this section we prove some basic facts about perfectly distinguishable states. Let us start from the definition:

The normalized states {ρi}i=1N⊆St1(A)\{\rho_{i}\}_{i=1}^{N}\subseteq{\mathsf{St}}_{1}({\rm A}) are perfectly distinguishable if there exists an observation-test {ai}i=1N\{a_{i}\}_{i=1}^{N} such that (aj∣ρi)=δij(a_{j}|\rho_{i})=\delta_{ij}. The observation-test {ai}i=1N\{a_{i}\}_{i=1}^{N} is called perfectly distinguishing.

From the distinguishability axiom 2 it is clear that every nontrivial system has at least two perfectly distinguishable states:

For every nontrivial system A{\rm A} there are at least two perfectly distinguishable states.

Proof. Let φ\varphi be a pure state of A{\rm A}. Obviously, φ\varphi is not completely mixed (unless the system A{\rm A} has only one state, that is, unless A{\rm A} is trivial). Hence, by axiom 2 there exists at least a state σ\sigma that is perfectly distinguishable from φ\varphi. ■\,\blacksquare

An equivalent condition for perfect distinguishability is the following:

The states {ρi}i=1N⊂St1(A)\{\rho_{i}\}_{i=1}^{N}\subset{\mathsf{St}}_{1}({\rm A}) are perfectly distinguishable if and only if there exists an observation-test {ai}i=1N\{a_{i}\}_{i=1}^{N} such that (ai∣ρi)=1(a_{i}|\rho_{i})=1 for every ii.

Proof. The condition (ai∣ ⁣ρi)=1,∀i=1,…,N\left(a_{i}\right|\left.\!\rho_{i}\right)=1,\quad\forall i=1,\dots,N is clearly necessary. On the other hand, the condition (ai∣ ⁣ρi)=1,∀i=1,…,N\left(a_{i}\right|\left.\!\rho_{i}\right)=1,\quad\forall i=1,\dots,N implies

Since all probabilities are non-negative, we must have (aj∣ρi)=0(a_{j}|\rho_{i})=0 for i≠ji\neq j, and therefore, (aj∣ρi)=δij(a_{j}|\rho_{i})=\delta_{ij}.  ■\,\blacksquare

A very general fact about state discrimination is expressed by the following:

If ρ\rho is perfectly distinguishable from σ\sigma and ρ′\rho^{\prime} (resp. σ′\sigma^{\prime}) belongs to the face identified by ρ\rho (resp. σ\sigma), then ρ′\rho^{\prime} is perfectly distinguishable from σ′\sigma^{\prime}.

Proof. Let {a,e−a}\{a,e-a\} be the binary observation-test that distinguishes perfectly between ρ\rho and σ\sigma. By definition, a∈Eff(A)a\in{\mathsf{Eff}}({\rm A}) is such that (a∣ρ)=1(a|\rho)=1 and (a∣σ)=0(a|\sigma)=0. Now, by lemma 19, (a∣ρ′)=1(a|\rho^{\prime})=1 and (a∣σ′)=0(a|\sigma^{\prime})=0 for all ρ′∈Fρ\rho^{\prime}\in F_{\rho} and σ′∈Fσ\sigma^{\prime}\in F_{\sigma}. ■\,\blacksquare

Thanks to purification and to the local distinguishability axiom 4, we are also in position to show a much stronger result:

Let {ρi}i=1N⊂Fρ\{\rho_{i}\}_{i=1}^{N}\subset F_{\rho} and {ρj}j=N+1N+M⊂Fσ\{\rho_{j}\}_{j=N+1}^{N+M}\subset F_{\sigma} be two sets of perfectly distinguishable states. If ρ\rho is perfectly distinguishable from σ\sigma, then the states {ρi}i=1N+M\{\rho_{i}\}_{i=1}^{N+M} are perfectly distinguishable.

Proof. Let {a,eA−a}\{a,e_{\rm A}-a\} be the observation-test such that (a∣ ⁣ρ)=1\left(a\right|\left.\!\rho\right)=1 and (a∣σ)=0(a|\sigma)=0. Now, by corollary 9 there is a transformation C∈Transf(A)\mathscr{C}\in{\mathsf{Transf}}({\rm A}) such that (eA∣C=(a∣\left(e_{\rm A}\right|\mathscr{C}=\left(a\right| and C=ρIA\mathscr{C}=_{\rho}\mathscr{I}_{\rm A}. Similarly, there exists a transformation C′∈Transf(A)\mathscr{C}^{\prime}\in{\mathsf{Transf}}({\rm A}) such that (eA∣C′=(eA∣−(a∣\left(e_{\rm A}\right|\mathscr{C}^{\prime}=\left(e_{\rm A}\right|-\left(a\right| and C′=σIA\mathscr{C}^{\prime}=_{\sigma}\mathscr{I}_{\rm A}. We can then define the following observation-test

where {ai}i=1N\{a_{i}\}_{i=1}^{N} (resp. {bj}j=N+1N+M\{b_{j}\}_{j=N+1}^{N+M}) is the observation-test that perfectly distinguishes among the states {ρi}i=1N\{\rho_{i}\}_{i=1}^{N} (resp. {ρj}j=N+1N+M\{\rho_{j}\}_{j=N+1}^{N+M}). By corollary 4 [see in particular Eq. (8)], {ci}i=1N+M\{c_{i}\}_{i=1}^{N+M} is indeed an observation-test: each cic_{i} is an effect and one has the normalization

Moreover, since C=ρIA\mathscr{C}=_{\rho}\mathscr{I}_{\rm A} and C′=σIA\mathscr{C}^{\prime}=_{\sigma}\mathscr{I}_{\rm A}, one has (ci∣ ⁣ρi)=1\left(c_{i}\right|\left.\!\rho_{i}\right)=1 for every i=1,…,M+Ni=1,\dots,M+N. By lemma 21, this implies that the states {ρi}i=1N+M\{\rho_{i}\}_{i=1}^{N+M} are perfectly distinguishable.  ■\,\blacksquare

A set of perfectly distinguishable states {ρi}i=1N\{\rho_{i}\}_{i=1}^{N} is maximal if there is no state ρN+1∈St1(A)\rho_{{N+1}}\in{\mathsf{St}}_{1}({\rm A}) such that the states {ρi}i=1N+1\{\rho_{i}\}_{i=1}^{N+1} are perfectly distinguishable

A set of perfectly distinguishable states {ρi}i=1N\{\rho_{i}\}_{i=1}^{N} is maximal if and only if the state ω=∑i=1Nρi/N\omega=\sum_{{i=1}}^{N}\rho_{i}/N is completely mixed.

Proof. We first prove that if ω\omega is completely mixed, then the set {ρi}i=1N\{\rho_{i}\}_{i=1}^{N} must be maximal. Indeed, if there existed a state ρN+1\rho_{N+1} such that {ρi}i=1N+1\{\rho_{i}\}_{i=1}^{N+1} are perfectly distinguishable, then clearly ρN+1\rho_{N+1} would be distinguishable from ω\omega. This is absurd because by proposition 1 no state can be perfectly distinguished from a completely mixed state. Conversely, if {ρi}i=1N\{\rho_{i}\}_{i=1}^{N} is maximal, then ω\omega is completely mixed. If it were not, by the distinguishability axiom 2, ω\omega would be perfectly distinguishable from some state ρN+1\rho_{N+1}. By lemma 23, this would imply that the states {ρi}i=1N+1\{\rho_{i}\}_{i=1}^{N+1} are perfectly distinguishable, in contradiction with the hypothesis that the set {ρi}i=1N\{\rho_{i}\}_{i=1}^{N} is maximal.  ■\,\blacksquare

Every set of perfectly distinguishable pure states can be extended to a maximal set of perfectly distinguishable pure states.

Any pure state belongs to a maximal set of perfectly distinguishable pure states.

We conclude this section with a few elementary facts about how the ideal compression of axiom 3 preserves the distinguishability properties. In the following we will choose a state ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) and E∈Transf(A,C)\mathscr{E}\in{\mathsf{Transf}}({\rm A},{\rm C}) (resp. D∈Transf(C,A)\mathscr{D}\in{\mathsf{Transf}}({\rm C},{\rm A})) will be the encoding (resp. decoding) in the ideal compression scheme for ρ\rho.

If the states {ρi}i=1k⊂Fρ\{\rho_{i}\}_{i=1}^{k}\subset F_{\rho} are perfectly distinguishable, then the states {Eρi}i=1k⊂St1(C)\{\mathscr{E}\rho_{i}\}_{i=1}^{k}\subset{\mathsf{St}}_{1}({\rm C}) are perfectly distinguishable. Conversely, if the states {σi}i=1k⊂St1(C)\{\sigma_{i}\}_{i=1}^{k}\subset{\mathsf{St}}_{1}({\rm C}) are perfectly distinguishable, then the states {Dσi}i=1k⊂Fρ\{\mathscr{D}\sigma_{i}\}_{i=1}^{k}\subset F_{\rho} are perfectly distinguishable.

Proof. Let {ai}i=1k\{a_{i}\}_{i=1}^{k} be the observation-test such that (ai∣ρi)=1(a_{i}|\rho_{i})=1 for every i=1,…,ki=1,\dots,k. Since the compression is lossless, we have DE∣ρi)=∣ρi)\mathscr{D}\mathscr{E}\left|\rho_{i}\right)=\left|\rho_{i}\right) and (ai∣DE∣ρi)=1(a_{i}|\mathscr{D}\mathscr{E}|\rho_{i})=1. Now, consider the test {ci}i=1k\{c_{i}\}_{i=1}^{k} defined by (ci∣=(ai∣D\left(c_{i}\right|=\left(a_{i}\right|\mathscr{D}. Clearly we have (ci∣E∣ρi)=1\left(c_{i}\right|\mathscr{E}\left|\rho_{i}\right)=1 for every i=1,…,ki=1,\dots,k. By lemma 21 this means that the states {Eρi}i=1k\{\mathscr{E}\rho_{i}\}_{i=1}^{k} are perfectly distinguishable. Similarly, let {bi}i=1k\{b_{i}\}_{i=1}^{k} the observation-test that distinguishes the set {σi}i=1k\{\sigma_{i}\}_{i=1}^{k}. Since ED=IC\mathscr{E}\mathscr{D}=\mathscr{I}_{\rm C} (lemma 2), we can conclude by the same argument that the states {Dσi}i=1k\{\mathscr{D}\sigma_{i}\}_{i=1}^{k} are perfectly distinguishable.  ■\,\blacksquare

We say that a set of perfectly distinguishable states {ρi}i=1k⊂Fρ\{\rho_{i}\}_{i=1}^{k}\subset F_{\rho} is maximal in the face FρF_{\rho} if there is no state ρk+1∈Fρ\rho_{k+1}\in F_{\rho} such that the states {ρi}i=1k+1\{\rho_{i}\}_{i=1}^{k+1} are perfectly distinguishable. We then have the following:

If {ρi}i=1k⊂Fρ\{\rho_{i}\}_{i=1}^{k}\subset F_{\rho} is a maximal set of perfectly distinguishable states in the face FρF_{\rho}, then {Eρi}i=1k∈St1(C)\{\mathscr{E}\rho_{i}\}_{i=1}^{k}\in{\mathsf{St}}_{1}({\rm C}) is a maximal set of perfectly distinguishable states. Conversely, if {σi}i=1kSt1(C)\{\sigma_{i}\}_{i=1}^{k}{\mathsf{St}}_{1}({\rm C}) is a maximal set of perfectly distinguishable states, then {Dσi}i=1k\{\mathscr{D}\sigma_{i}\}_{i=1}^{k} is a maximal set of perfectly distinguishable states in the face FρF_{\rho}.

Proof. Distinguishability of the states {Eρi}i=1k\{\mathscr{E}\rho_{i}\}_{i=1}^{k} and {Dσi}i=1k\{\mathscr{D}\sigma_{i}\}_{i=1}^{k} is proved by lemma 25. Let us now prove maximality. By contradiction, suppose that the set {ρi}i=1k\{\rho_{i}\}_{i=1}^{k} is maximal in the face FρF_{\rho} while the set {σi}i=1k\{\sigma_{i}\}_{i=1}^{k}, σi:=Eρi\sigma_{i}:=\mathscr{E}\rho_{i} is not maximal. This means that there exists a state σk+1∈St1(C)\sigma_{k+1}\in{\mathsf{St}}_{1}({\rm C}) such that the states {σi}i=1k+1\{\sigma_{i}\}_{i=1}^{k+1} are perfectly distinguishable. By lemma 25 the states {Dσi}i=1k+1\{\mathscr{D}\sigma_{i}\}_{i=1}^{k+1} are perfectly distinguishable. Since DEρi=ρi\mathscr{D}\mathscr{E}\rho_{i}=\rho_{i} for every i=1,…,ki=1,\dots,k, this means that the states {ρi}i=1k∪{Dσk+1}\{\rho_{i}\}_{i=1}^{k}\cup\{\mathscr{D}\sigma_{k+1}\} are perfectly distinguishable, in contradiction with the fact that {ρi}i=1k\{\rho_{i}\}_{i=1}^{k} is maximal. This proves that the set {Eρi}i=1k\{\mathscr{E}\rho_{i}\}_{i=1}^{k} must be maximal. Conversely, if the set {σi}⊂St1(C)\{\sigma_{i}\}\subset{\mathsf{St}}_{1}({\rm C}) is maximal, using the same argument we can prove that the set {Dσi}i=1k\{\mathscr{D}\sigma_{i}\}_{i=1}^{k} must be maximal in FρF_{\rho}.  ■\,\blacksquare

VI Duality between pure states and atomic effects

We now show the existence of a one-to-one correspondence between states and effects of any system A{\rm A} in the theory. Let us start from a simple observation:

If aa is atomic and (a∣ ⁣ρ)=∥a∥\left(a\right|\left.\!\rho\right)=\|a\| for ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}), then ρ\rho must be pure.

Proof. By lemma 19, the condition (a∣ ⁣ρ)=∥a∥\left(a\right|\left.\!\rho\right)=\|a\| implies a=ρ∥a∥ ea=_{\rho}\|a\|~e. By theorem 1, the condition a=ρ∥a∥ ea=_{\rho}\|a\|~e implies

We are now in position to show that every atomic effect is associated to a unique pure state.

For every atomic effect a∈Eff(A)a\in{\mathsf{Eff}}({\rm A}), there exists a unique pure state φ∈St1(A)\varphi\in{\mathsf{St}}_{1}({\rm A}) such that (a∣ ⁣φ)=∥a∥\left(a\right|\left.\!\varphi\right)=\|a\|.

Proof. Let ρ\rho be a state such that (a∣ ⁣ρ)=∥a∥\left(a\right|\left.\!\rho\right)=\|a\|. By lemma 26, ρ\rho must be pure. Moreover, this pure state must be unique: suppose that φ\varphi and φ′\varphi^{\prime} are pure states such that (a∣ ⁣φ)=(a∣ ⁣φ′)=∥a∥\left(a\right|\left.\!\varphi\right)=\left(a\right|\left.\!\varphi^{\prime}\right)=\|a\|. Then for ω=1/2(φ+φ′)\omega=1/2(\varphi+\varphi^{\prime}) one has (a∣ ⁣ω)=∥a∥\left(a\right|\left.\!\omega\right)=\|a\|. Since ω\omega must be pure, one has φ=φ′\varphi=\varphi^{\prime}.  ■\,\blacksquare

We now show the converse result: for every pure state φ∈St1(A)\varphi\in{\mathsf{St}}_{1}({\rm A}) there exists a unique atomic effect aa such that (a∣ ⁣φ)=1\left(a\right|\left.\!\varphi\right)=1. Let us start from the existence:

Let {φi}i=1N⊂St1(A)\{\varphi_{i}\}_{i=1}^{N}\subset{\mathsf{St}}_{1}({\rm A}) be a maximal set of perfectly distinguishable pure states and let {ai}i=1N\{a_{i}\}_{i=1}^{N} be the observation-test such that (ai∣ ⁣φj)=δij\left(a_{i}\right|\left.\!\varphi_{j}\right)=\delta_{ij}. Then each effect aia_{i} is atomic with ∥ai∥=1\|a_{i}\|=1.

Proof. It is obvious that ∥ai∥=1\|a_{i}\|=1, because of the condition (ai∣ ⁣φi)=1\left(a_{i}\right|\left.\!\varphi_{i}\right)=1. It remains to prove atomicity. Consider the state ω=∑i=1Nφi/N\omega=\sum_{i=1}^{N}\varphi_{i}/N, which is completely mixed by theorem 6. Let Ψω∈St1(AB)\Psi_{\omega}\in{\mathsf{St}}_{1}({\rm A}{\rm B}) be a purification of ω\omega, chosen in such a way that the marginal on system B{\rm B} is completely mixed (theorem 4). As a consequence of purification (lemma 12), there exists an observation-test {bi}i=1N\{b_{i}\}_{i=1}^{N} on system B{\rm B} such that (bi∣B∣Ψω)AB=1/N∣φi)A\left(b_{i}\right|_{\rm B}\left|\Psi_{\omega}\right)_{{\rm A}{\rm B}}=1/N\left|\varphi_{i}\right)_{\rm A}. Since Ψω\Psi_{\omega} is dynamically faithful on system B{\rm B}, each effect bib_{i} must be atomic. Now, define the normalized states {ρi}i=1N⊂St1(B)\{\rho_{i}\}_{i=1}^{N}\subset{\mathsf{St}}_{1}({\rm B}) and the probabilities {pi}i=1N\{p_{i}\}_{i=1}^{N} by

Applying the deterministic effect eBe_{\rm B} on both sides one has pi=(ai∣ ⁣ω)=1/Np_{i}=\left(a_{i}\right|\left.\!\omega\right)=1/N. On the other hand, applying the effect bjb_{j} one has instead 1/N(bj∣ ⁣ρi)B=1/N(ai∣ ⁣φj)=δij/N1/N\left(b_{j}\right|\left.\!\rho_{i}\right)_{\rm B}=1/N\left(a_{i}\right|\left.\!\varphi_{j}\right)=\delta_{ij}/N. This implies (bi∣ ⁣ρi)=1\left(b_{i}\right|\left.\!\rho_{i}\right)=1 for every ii. Since bib_{i} is atomic, lemma 26 forces each ρi\rho_{i} to be pure. Finally, each aia_{i} must be atomic since its Choi state pi∣ρi)B=(ai∣A∣Ψω)ABp_{i}\left|\rho_{i}\right)_{\rm B}=\left(a_{i}\right|_{\rm A}\left|\Psi_{\omega}\right)_{{\rm A}{\rm B}} is pure (theorem 2).  ■\,\blacksquare

As a consequence, we can prove the following existence result:

For every pure state φ∈St1(A)\varphi\in{\mathsf{St}}_{1}({\rm A}) there exists an atomic effect such that (a∣ ⁣φ)=1\left(a\right|\left.\!\varphi\right)=1.

Proof. By corollary 11, every pure state belongs to a maximal set of perfectly distinguishable pure states {φi}i=1N\{\varphi_{i}\}_{i=1}^{N}, say φ=φ1\varphi=\varphi_{1}. The thesis then follows from lemma 27.  ■\,\blacksquare

We now prove that the atomic effect aa such that (a∣φ)=1(a|\varphi)=1 is unique. For this purpose we need two auxiliary lemmas:

Let φ∈St1(A)\varphi\in{\mathsf{St}}_{1}({\rm A}) be an arbitrary pure state and let pφp_{\varphi} be the probability defined by

where χ\chi is the invariant state of system A{\rm A}. Then the value of the probability pφp_{\varphi} is independent of φ\varphi.

Proof. Since for every couple of pure states φ\varphi and ψ\psi one has ψ=Uφ\psi=\mathscr{U}\varphi for some reversible channel U\mathscr{U} (lemma 9), and since χ\chi is invariant, one has χ=pφ+(1−p)σ\chi=p\varphi+(1-p)\sigma if and only if χ=pψ+(1−p)Uσ\chi=p\psi+(1-p)\mathscr{U}\sigma. The maximum probabilities for φ\varphi and ψ\psi are then equal. ■\,\blacksquare

Since pφ=pψp_{\varphi}=p_{\psi} for every couple of pure states, from now on we will write pmax⁡p_{\max} in place of pφp_{\varphi}.

Let φ∈St1(A)\varphi\in{\mathsf{St}}_{1}({\rm A}) be a pure state and a∈Eff(A)a\in{\mathsf{Eff}}({\rm A}) be an atomic effect such that (a∣ ⁣φ)=1\left(a\right|\left.\!\varphi\right)=1. Let ∣Φ)AB\left|\Phi\right)_{{\rm A}{\rm B}} be a purification of the invariant state ∣χ)A\left|\chi\right)_{\rm A}, chosen in such a way that the marginal on system B{\rm B} is completely mixed, and let bb be the unique atomic effect on B{\rm B} such that

[note that bb exists by lemma 12is uniquely defined by Eq. (12) because Φ\Phi is faithful for system B{\rm B}]. Then one has

where ψ\psi is the unique pure state such that (b∣ ⁣ψ)=1\left(b\right|\left.\!\psi\right)=1.

Proof. Define the normalized pure state ψ\psi and the probability qq by

In order to prove the thesis we have to show that q=pmax⁡q=p_{\max} and (b∣ ⁣ψ)=1\left(b\right|\left.\!\psi\right)=1. Applying bb on both sides of Eq. (14) and using Eq. (12) we obtain q(b∣ ⁣ψ)=pmax⁡(a∣ ⁣φ)=pmaxq\left(b\right|\left.\!\psi\right)=p_{\max}\left(a\right|\left.\!\varphi\right)=p_{max}. This implies

with the equality if and only if (b∣ ⁣ψ)=1\left(b\right|\left.\!\psi\right)=1. Let b′b^{\prime} be an atomic effect such that (b′∣ ⁣ψ)=1\left(b^{\prime}\right|\left.\!\psi\right)=1 (such an effect exists because of lemma 28 ). Define the normalized pure state φ′\varphi^{\prime} and the probability p′p^{\prime} by

Applying aa on both sides and using Eq. (14) we obtain p′(a∣ ⁣φ′)=q(b′∣ ⁣ψ)=qp^{\prime}\left(a\right|\left.\!\varphi^{\prime}\right)=q\left(b^{\prime}\right|\left.\!\psi\right)=q, which implies p′≥qp^{\prime}\geq q, with the equality if and only if (a∣ ⁣φ′)=1\left(a\right|\left.\!\varphi^{\prime}\right)=1. Combining this with the inequality (15) we have p′≥q≥pmax⁡p^{\prime}\geq q\geq p_{\max}. On the other hand, by Lemma 29 one has p′≤pmax⁡p^{\prime}\leq p_{\max}, and consequently p′=q=pmax⁡p^{\prime}=q=p_{\max}. This also implies that (b∣ ⁣ψ)=1\left(b\right|\left.\!\psi\right)=1 and (a∣ ⁣φ′)=1\left(a\right|\left.\!\varphi^{\prime}\right)=1. ■\,\blacksquare

For every pure state φ∈St1(A)\varphi\in{\mathsf{St}}_{1}({\rm A}) there is a unique atomic effect a∈Eff(A)a\in{\mathsf{Eff}}({\rm A}) such that (a∣ ⁣φ)=1\left(a\right|\left.\!\varphi\right)=1.

Proof. Existence has been already proved in lemma 28. Let us prove uniqueness: suppose that aa and a′a^{\prime} are two atomic effects such that (a∣ ⁣φ)=(a′∣ ⁣φ)=1\left(a\right|\left.\!\varphi\right)=\left(a^{\prime}\right|\left.\!\varphi\right)=1. Then, applying lemma 30 to aa and a′a^{\prime} we obtain

Since Φ\Phi is dynamically faithful, this implies a=a′a=a^{\prime}. ■\,\blacksquare

Finally, an important consequence of theorem 8 is

If a,a′∈Eff(A)a,a^{\prime}\in{\mathsf{Eff}}({\rm A}) are two atomic effects with ∥a∥=∥a′∥=1\|a\|=\|a^{\prime}\|=1, then there is a reversible channel U∈GA\mathscr{U}\in{\mathbf{G}}_{\rm A} such that (a′∣A=(a∣AU\left(a^{\prime}\right|_{\rm A}=\left(a\right|_{\rm A}\mathscr{U}.

Proof. Let φ\varphi and φ′\varphi^{\prime} be the (unique) normalized states such that (a∣ ⁣φ)=1\left(a\right|\left.\!\varphi\right)=1 and (a′∣ ⁣φ′)=1\left(a^{\prime}\right|\left.\!\varphi^{\prime}\right)=1, respectively. Now, there is a reversible channel U∈GA\mathscr{U}\in{\mathbf{G}}_{\rm A} such that ∣φ)A=U∣φ′)A\left|\varphi\right)_{\rm A}=\mathscr{U}\left|\varphi^{\prime}\right)_{\rm A}. Hence, (a′∣ ⁣φ′)=(a∣ ⁣φ)A=(a∣U∣φ′)\left(a^{\prime}\right|\left.\!\varphi^{\prime}\right)=\left(a\right|\left.\!\varphi\right)_{\rm A}=\left(a\right|\mathscr{U}\left|\varphi^{\prime}\right). By theorem 8, one has (a′∣A=(a∣AU\left(a^{\prime}\right|_{\rm A}=\left(a\right|_{\rm A}\mathscr{U}. ■\,\blacksquare

We conclude this section with an elementary result that will be used later in the paper:

Let E∈Transf(A,C)\mathscr{E}\in{\mathsf{Transf}}({\rm A},{\rm C}) and D∈Transf(C,A)\mathscr{D}\in{\mathsf{Transf}}({\rm C},{\rm A}) be the encoding and the decoding in the ideal compression scheme for ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}). If ∣φ)∈Fρ\left|\varphi\right)\in F_{\rho} is a pure state and (a∣∈Eff(A)\left(a\right|\in{\mathsf{Eff}}({\rm A}) is the atomic effect such that (a∣ ⁣φ)=1\left(a\right|\left.\!\varphi\right)=1, then ∣γ):=E∣φ)∈St1(C)\left|\gamma\right):=\mathscr{E}\left|\varphi\right)\in{\mathsf{St}}_{1}({\rm C}) is a pure state and (c∣:=(a∣D∈Eff(C)\left(c\right|:=\left(a\right|\mathscr{D}\in{\mathsf{Eff}}({\rm C}) is the atomic effect such that (c∣ ⁣γ)=1\left(c\right|\left.\!\gamma\right)=1.

Proof. The state ∣γ):=E∣φ)\left|\gamma\right):=\mathscr{E}\left|\varphi\right) is pure by lemma 4. The effect (c∣:=(a∣D\left(c\right|:=\left(a\right|\mathscr{D} is atomic by lemmas 14 and 16. Since DE=ρIA\mathscr{D}\mathscr{E}=_{\rho}\mathscr{I}_{\rm A}, one has (c∣ ⁣γ)=(a∣DE∣φ)=(a∣φ)=1\left(c\right|\left.\!\gamma\right)=\left(a\right|\mathscr{D}\mathscr{E}|\varphi)=(a|\varphi)=1.  ■\,\blacksquare

VII Dimension

In this section we show that each system in our theory has given informational dimension, defined as the maximum number of perfectly distinguishable pure states available in the system. In the Hilbert space framework, the informational dimension will be the dimension of the Hilbert space.

All maximal sets of perfectly distinguishable pure states have the same number of elements.

Proof. Let {φi}i=1N\{\varphi_{i}\}_{i=1}^{N} be a maximal set of perfectly distinguishable pure states for system A{\rm A}, and let {ai}i=1N\{a_{i}\}_{i=1}^{N} the observation-test such that (ai∣ ⁣φj)=δij\left(a_{i}\right|\left.\!\varphi_{j}\right)=\delta_{ij}. By lemma 27, each aia_{i} is atomic and ∥ai∥=1\|a_{i}\|=1. Then, by corollary 13, one has (ai∣A=(a0∣Ui\left(a_{i}\right|_{\rm A}=\left(a_{0}\right|\mathscr{U}_{i}, where each Ui\mathscr{U}_{i} is a reversible channel and a0a_{0} is a fixed atomic effect with ∥a0∥=1\|a_{0}\|=1. By the invariance of χ\chi we then obtain (ai∣χA)=(a0∣Ui∣χA)=(a0∣χA)(a_{i}|\chi_{\rm A})=(a_{0}|\mathscr{U}_{i}|\chi_{\rm A})=(a_{0}|\chi_{\rm A}). On the other hand, one has ∑i=1N(ai∣ ⁣χA)=1\sum_{i=1}^{N}\left(a_{i}\right|\left.\!\chi_{\rm A}\right)=1, which implies N=1/(a0∣ ⁣χA)N=1/\left(a_{0}\right|\left.\!\chi_{\rm A}\right). Since a0a_{0} is arbitrary, NN is independent of the choice of the set {φi}i=1N\{\varphi_{i}\}_{i=1}^{N}.  ■\,\blacksquare

An immediate consequence of the proof of lemma 32 is

For every atomic effect aa with ∥a∥=1\|a\|=1 one has (a∣ ⁣χA)=1/dA\left(a\right|\left.\!\chi_{\rm A}\right)=1/d_{\rm A}.

This simple fact has two very important consequences. The first is that the dimension of a composite system is the product of the dimensions of the components:

The dimension of the composite system AB{\rm A}{\rm B} is the product of the dimensions of A{\rm A} and B{\rm B}, namely dAB=dAdBd_{{\rm A}{\rm B}}=d_{\rm A}d_{\rm B}.

Proof. From lemma 10 we know that χA⊗χB\chi_{\rm A}\otimes\chi_{\rm B} is the unique invariant state of system AB{\rm A}{\rm B}. Now, if a∈Eff(A)a\in{\mathsf{Eff}}({\rm A}) and b∈Eff(B)b\in{\mathsf{Eff}}({\rm B}) are such that ∥a∥=∥b∥=1\|a\|=\|b\|=1, then a⊗ba\otimes b is such that ∥a⊗b∥=1\|a\otimes b\|=1. Hence we have 1/dAB=(a⊗b∣ ⁣χA⊗χB)=(a∣ ⁣χA)(b∣ ⁣χB)=1/(dAdB)1/d_{{\rm A}{\rm B}}=\left(a\otimes b\right|\left.\!\chi_{\rm A}\otimes\chi_{\rm B}\right)=\left(a\right|\left.\!\chi_{\rm A}\right)\left(b\right|\left.\!\chi_{\rm B}\right)=1/(d_{\rm A}d_{\rm B}).  ■\,\blacksquare

The second consequence is the relation between the dimension and the maximum probability of a pure state in the convex decomposition of the invariant state ∣χ)A\left|\chi\right)_{\rm A}:

For every system A{\rm A}, the maximum probability of a pure state in the convex decomposition of the invariant state is pmax⁡=1/dAp_{\max}=1/d_{\rm A}.

Proof. Let Φ∈St1(AB)\Phi\in{\mathsf{St}}_{1}({\rm A}{\rm B}) be a purification of the invariant state ∣χA)\left|\chi_{\rm A}\right), chosen in such a way that the marginal on system B{\rm B} is completely mixed. Let a∈Eff(A)a\in{\mathsf{Eff}}({\rm A}) be an atomic effect with ∥a∥=1\|a\|=1. Then, equation (13) becomes

where ψ\psi is some normalized pure state of system B{\rm B}. Applying the deterministic effect ee on system B{\rm B} on both sides we obtain (a∣ ⁣χA)=pmax⁡\left(a\right|\left.\!\chi_{{\rm A}}\right)=p_{\max}. Finally, corollary 14 states (a∣ ⁣χA)=1/dA\left(a\right|\left.\!\chi_{\rm A}\right)=1/d_{\rm A}. By comparison, we obtain pmax⁡=1/dAp_{\max}=1/d_{\rm A}.  ■\,\blacksquare

Thanks to the compression axiom 3, the notion of dimension can be applied not only to the whole state space St1(A){\mathsf{St}}_{1}({\rm A}) but also to its faces. With face FF of the convex set St1(A){\mathsf{St}}_{1}({\rm A}) we always mean the face FρF_{\rho} identified by some state ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}).

Let FF be a face of the convex set St1(A){\mathsf{St}}_{1}({\rm A}). Every maximal set {φi}i=1k\{\varphi_{i}\}_{i=1}^{k} of perfectly distinguishable pure states in FF has the same cardinality kk. Precisely, if FF is the face identified by ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) and E∈Transf(A,C)\mathscr{E}\in{\mathsf{Transf}}({\rm A},{\rm C}) is the encoding in the ideal compression for ρ\rho, then we have k=dCk=d_{\rm C}

Proof. The set {Eφi}i=1k⊂St1(C)\{\mathscr{E}\varphi_{i}\}_{i=1}^{k}\subset{\mathsf{St}}_{1}({\rm C}) is perfectly distinguishable by lemma 25, and it is maximal by corollary 12. Moreover, the states {Eφi}i=1k\{\mathscr{E}\varphi_{i}\}_{i=1}^{k} are pure by lemma 4. Hence, the cardinality kk of the set {φi}i=1k\{\varphi_{i}\}_{i=1}^{k} must be k=dCk=d_{\rm C}. ■\,\blacksquare

From now on the maximum number of perfectly distinguishable states in the face FF will be called the dimension of the face FF and will be denoted by ∣F∣|F|.

VIII Decomposition into perfectly distinguishable pure states

In this section we show that in a theory satisfying our principles any state can be written as a convex combination of perfectly distinguishable pure states. In quantum theory, this corresponds to the diagonalization of the density matrix.

To prove this result we need first a sufficient condition for the distinguishability of states, given in the following

Let {ρi}i=1N⊂St1(A)\{\rho_{i}\}_{i=1}^{N}\subset{\mathsf{St}}_{1}({\rm A}) be a set of states. If there exists a set of effects {bi}i=1N⊂Eff(A)\{b_{i}\}_{i=1}^{N}\subset{\mathsf{Eff}}({\rm A}) (not necessarily an observation-test) such that (bi∣ρj)=δij(b_{i}|\rho_{j})=\delta_{ij}, then the states {ρi}i=1N\{\rho_{i}\}_{i=1}^{N} are perfectly distinguishable.

Proof. For each i=1,…,Ni=1,\dots,N consider the binary test {bi,e−bi}\{b_{i},e-b_{i}\}. Since by hypothesis (bi∣ ⁣ρj)=δij\left(b_{i}\right|\left.\!\rho_{j}\right)=\delta_{ij}, the test {bi,e−bi}\{b_{i},e-b_{i}\} can perfectly distinguish ρi\rho_{i} from any mixture of the states {ρj}j≠i\{\rho_{j}\}_{j\not=i}. In particular, this means that, for every M<NM<N, ρM+1\rho_{M+1} can be perfectly distinguished from the mixture ωM=∑j=1Mρj/M\omega_{M}=\sum_{j=1}^{M}\rho_{j}/M. Note that, by definition, the states {ρi}i=1M\{\rho_{i}\}_{i=1}^{M} belong to the face FωMF_{\omega_{M}}. We now prove by induction on MM that the states {ρi}i=1M\{\rho_{i}\}_{i=1}^{M} are perfectly distinguishable. This is true for M=1M=1. Now, suppose that the states {ρi}i=1M\{\rho_{i}\}_{i=1}^{M} are perfectly distinguishable. Since the state ρM+1\rho_{M+1} is perfectly distinguishable from ωM\omega_{M}, by lemma 23 we have that the states {ρi}i=1M+1\{\rho_{i}\}_{i=1}^{M+1} are perfectly distinguishable. Taking M=N−1M=N-1 the thesis follows. ■\,\blacksquare

We now show that the invariant state χ\chi is a mixture of perfectly distinguishable pure states.

For every maximal set of perfectly distinguishable pure states {φi}i=1dA⊂St1(A)\{\varphi_{i}\}_{i=1}^{d_{\rm A}}\subset{\mathsf{St}}_{1}({\rm A}) one has

Proof. Let {ai}i=1dA\{a_{i}\}_{i=1}^{d_{\rm A}} be the observation test such that (ai∣ ⁣φj)=δij\left(a_{i}\right|\left.\!\varphi_{j}\right)=\delta_{ij}, and Φ∈St1(AB)\Phi\in{\mathsf{St}}_{1}({\rm A}{\rm B}) be a purification of χ\chi, chosen in such a way that the marginal on system B{\rm B} is completely mixed (theorem 4). Let {ψi}i=1dA⊂St1(B)\{\psi_{i}\}_{i=1}^{d_{\rm A}}\subset{\mathsf{St}}_{1}({\rm B}) be the pure states defined by

and, for each ii, let bib_{i} be the atomic effect such that

(here we used lemma 30 and the fact that pmax⁡=1/dAp_{\max}=1/d_{\rm A}). Then we have

By lemma 35, this implies that the states {ψi}i=1dA\{\psi_{i}\}_{i=1}^{d_{\rm A}} are perfectly distinguishable. Now, since the marginal of ∣Φ)AB\left|\Phi\right)_{{\rm A}{\rm B}} on system B{\rm B} is completely mixed, theorem 6 states that the set {ψi}i=1dA\{\psi_{i}\}_{i=1}^{d_{\rm A}} is maximal. Let {bi′}i=1dA\{b^{\prime}_{i}\}_{i=1}^{d_{\rm A}} the observation test such that (bi′∣ ⁣ψj)=δij\left(b_{i}^{\prime}\right|\left.\!\psi_{j}\right)=\delta_{ij}. By lemma 27, each bi′b_{i}^{\prime} must be atomic. On the other hand, there is a unique atomic effect bib_{i} such that (bi∣ ⁣ψi)=1\left(b_{i}\right|\left.\!\psi_{i}\right)=1 (theorem 8). Therefore, bi′=bib_{i}^{\prime}=b_{i}. This means that the effects {bi}i=1dA\{b_{i}\}_{i=1}^{d_{\rm A}} form an observation test. Once this fact has been proved, using Eq. (16) we obtain

The distance between the invariant state χA\chi_{\rm A} and an arbitrary pure state φ∈St1(A)\varphi\in{\mathsf{St}}_{1}({\rm A}) is

Proof. Take a maximal set of perfectly distinguishable pure states {φi}i=1dA\{\varphi_{i}\}_{i=1}^{d_{\rm A}} such that φ1=φ\varphi_{1}=\varphi (corollary 11). Since χ=∑i=1dAφi/dA\chi=\sum_{i=1}^{d_{\rm A}}\varphi_{i}/d_{\rm A} one has χ−φ=(dA−1)dA(σ−φ1)\chi-\varphi=\frac{(d_{\rm A}-1)}{d_{\rm A}}(\sigma-\varphi_{1}), where σ=∑i=2dAφi/(dA−1)\sigma=\sum_{i=2}^{d_{\rm A}}\varphi_{i}/(d_{\rm A}-1). Hence, one has ∥χ−φ∥=(dA−1)dA∥σ−φ1∥=2(dA−1)dA\|\chi-\varphi\|=\frac{(d_{\rm A}-1)}{d_{\rm A}}\|\sigma-\varphi_{1}\|=\frac{2(d_{\rm A}-1)}{d_{\rm A}}, having used that σ\sigma and φ1\varphi_{1} are perfectly distinguishable and therefore ∥σ−φ1∥=2\|\sigma-\varphi_{1}\|=2 (see subsection II-I in Ref. purification).  ■\,\blacksquare

We can now prove the following strong result:

For every system A{\rm A}, every mixed state can be written as a convex combination of perfectly distinguishable pure states.

for some state σt0\sigma_{t_{0}} on the border of St1(A){\mathsf{St}}_{1}({\rm A}), that is, for some state that is not completely mixed. But we know from the discussion of point (1) that the state σt0\sigma_{t_{0}} is a mixture of perfectly distinguishable pure states, say σt0=∑i=1kpiφi\sigma_{t_{0}}=\sum_{i=1}^{k}p_{i}\varphi_{i}. By lemma 24, this set can be extended to a maximal set of perfectly distinguishable pure states {φi}i=1dA\{\varphi_{i}\}_{i=1}^{d_{\rm A}}. On the other hand, theorem 9 states that χ=∑i=1dAφi/dA\chi=\sum_{i=1}^{d_{\rm A}}\varphi_{i}/d_{{\rm A}}. This implies the desired decomposition

where qi=piq_{i}=p_{i} for 1≤i≤k1\leq i\leq k, and qi=0q_{i}=0 otherwise.  ■\,\blacksquare

It is easy to show that the marginals of a pure bipartite state have the same spectral decomposition:

Proof. Let {ai}i=1dA\{a_{i}\}_{i=1}^{d_{\rm A}} be the observation-test such that (ai∣ ⁣φj)=δij\left(a_{i}\right|\left.\!\varphi_{j}\right)=\delta_{ij}, {bi}i=1r\{b_{i}\}_{i=1}^{r} be the observation test such that (bi∣B∣Ψ)AB=pi∣φi)A\left(b_{i}\right|_{\rm B}\left|\Psi\right)_{{\rm A}{\rm B}}=p_{i}\left|\varphi_{i}\right)_{\rm A} for every i≤ri\leq r. For i≤ri\leq r, define the pure state ψi∈St1(B)\psi_{i}\in{\mathsf{St}}_{1}({\rm B}) and the probability qiq_{i} via the relation

[Note that ψi\psi_{i} is pure due to the pure conditioning axiom] By definition, we have

The above relation implies qi=∑j=1dAqi(bj∣ ⁣ψi)=∑jpiδij=piq_{i}=\sum_{j=1}^{d_{\rm A}}q_{i}\left(b_{j}\right|\left.\!\psi_{i}\right)=\sum_{j}p_{i}\delta_{ij}=p_{i} and (bj∣ ⁣ψi)=δij\left(b_{j}\right|\left.\!\psi_{i}\right)=\delta_{ij}. Hence, the states {ψi}i=1r\{\psi_{i}\}_{i=1}^{r} are perfectly distinguishable. On the other hand, we have (ai⊗eB∣ ⁣Ψ)=(ai∣ ⁣ρ)=0∀i>r,\left(a_{i}\otimes e_{\rm B}\right|\left.\!\Psi\right)=\left(a_{i}\right|\left.\!\rho\right)=0\quad\forall i>r, which implies (ai∣A∣Ψ)AB=0\left(a_{i}\right|_{{\rm A}}\left|\Psi\right)_{{\rm A}{\rm B}}=0, ∀i>r\forall i>r . Therefore, we obtained

which is the desired spectral decomposition.  ■\,\blacksquare

The spectral decomposition of states has many consequences. Here we just discuss the simplest ones, which are needed for the purpose of the derivation of quantum theory.

A first consequence is the following lemma:

Let φ∈St1(A)\varphi\in{\mathsf{St}}_{1}({\rm A}) be a pure state and let a∈Eff(A)a\in{\mathsf{Eff}}({\rm A}) be the unique atomic effect such that (a∣ ⁣φ)=1\left(a\right|\left.\!\varphi\right)=1. If φ\varphi is perfectly distinguishable from ρ\rho, then (a∣ ⁣ρ)=0\left(a\right|\left.\!\rho\right)=0.

Proof. Let us write ρ=∑i=1kpiφi\rho=\sum_{i=1}^{k}p_{i}\varphi_{i}, with {φi}i=1k\{\varphi_{i}\}_{i=1}^{k} perfectly distinguishable pure states and pi>0p_{i}>0 for each ii. Now, by lemma 23 the states {φ1,…,φk,φ}\{\varphi_{1},\dots,\varphi_{k},\varphi\} are perfectly distinguishable, and by lemma 24 this set can be extended to a maximal set of perfectly distinguishable pure states {γm}m=1dA\{\gamma_{m}\}_{m=1}^{d_{\rm A}}, with γi=φi\gamma_{i}=\varphi_{i} for i≤ki\leq k and γk+1=φ\gamma_{k+1}=\varphi. Denote by {cm}m=1dA\{c_{m}\}_{m=1}^{d_{\rm A}} the observation test that perfectly distinguishes between the states {γm}\{\gamma_{m}\}. Note that, by definition, (ck+1∣ ⁣φ)=1\left(c_{k+1}\right|\left.\!\varphi\right)=1 and (ck+1∣ ⁣φj)=0\left(c_{k+1}\right|\left.\!\varphi_{j}\right)=0 for every j≠k+1j\not=k+1. Also, recall that ck+1c_{k+1} is atomic (lemma 27). By the duality of theorem 8 we have a=ck+1a=c_{k+1}, and, therefore, (a∣ ⁣ρ)=∑i=1kpi(ck+1∣ ⁣ψi)=0\left(a\right|\left.\!\rho\right)=\sum_{i=1}^{k}p_{i}\left(c_{k+1}\right|\left.\!\psi_{i}\right)=0. ■\,\blacksquare

Another consequence of theorem 10 is the following characterization of the completely mixed states as full rank states:

(Characterization of completely mixed states) A state ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}), written as a mixture ρ=∑i=1dApiφi\rho=\sum_{i=1}^{d_{\rm A}}p_{i}\varphi_{i} of a maximal set of perfectly distinguishable pure states {φi}i=1dA\{\varphi_{i}\}_{i=1}^{d_{\rm A}}, is completely mixed if and only if pi>0p_{i}>0 for every i=1,…,dAi=1,\dots,d_{\rm A}.

Proof. Necessity: If pi=0p_{i}=0 for some ii, then ρ\rho is perfectly distinguishable from φi\varphi_{i}. Hence, it cannot be completely mixed. Sufficiency: let pmin⁡=min⁡{pi,i=1,…,dA}p_{\min}=\min\{p_{i},i=1,\dots,d_{\rm A}\}. Then we have ρ=pmin⁡χ+(1−pmin⁡)σ\rho=p_{\min}\chi+(1-p_{\min})\sigma, where σ\sigma is the state defined by σ=1/(1−pmin⁡)∑i=1dA(pi−pmin⁡/dA)φi\sigma=1/(1-p_{\min})\sum_{i=1}^{d_{\rm A}}(p_{i}-p_{\min}/d_{\rm A})\varphi_{i}. Since ρ\rho contains χ\chi in its convex decomposition, and since χ\chi is completely mixed, we conclude that ρ\rho is completely mixed.  ■\,\blacksquare

In particular, for two-dimensional systems we have the result:

For dA=2d_{\rm A}=2 any state on the border of St1(A){\mathsf{St}}_{1}({\rm A}) is pure.

Proof. Write ξ\xi as ξ=c+ρ−c−σ\xi=c_{+}\rho-c_{-}\sigma, where c+,c−≥0c_{+},c_{-}\geq 0 and ρ\rho and σ\sigma are normalized states. If c−=0c_{-}=0 there is nothing to prove, because ξ\xi is proportional to a state. Then, suppose that c−>0c_{-}>0. Write σ\sigma as σ=∑ipiψi\sigma=\sum_{i}p_{i}\psi_{i} where {ψi}\{\psi_{i}\} are perfectly distinguishable and define k=max⁡{pi}k=\max\{p_{i}\}. Then one has χ+1/(c−kdA)ξ=(χ−1/(kdA)σ)+c+/(c−kdA)ρ\chi+1/(c_{-}kd_{\rm A})\xi=(\chi-1/(kd_{\rm A})\sigma)+c_{+}/(c_{-}kd_{\rm A})\rho. Now, by definition χ−1/(kdA)σ\chi-1/(kd_{\rm A})\sigma is proportional to a state: indeed we have (χ−1/(kdA)σ)=1/dA∑i(1−pi/k)ψi(\chi-1/(kd_{\rm A})\sigma)=1/d_{\rm A}\sum_{i}(1-p_{i}/k)\psi_{i}, and, by definition 1−pi/k≥01-p_{i}/k\geq 0. Therefore χ+1/(c−kdA)ξ\chi+1/(c_{-}kd_{\rm A})\xi is proportional to a state, say χ+1/(c−kdA)ξ=tτ\chi+1/(c_{-}kd_{\rm A})\xi=t\tau, with t>0t>0. Writing τ\tau as τ=∑iqiφi\tau=\sum_{i}q_{i}\varphi_{i}, where {φi}i=1dA\{\varphi_{i}\}_{i=1}^{d_{\rm A}} is a maximal set of perfectly distinguishable pure states, we then obtain ξ=(c−kdA)(tτ−χ)=(c−kdA)∑i(tqi−1/dA)φi\xi=(c_{-}kd_{\rm A})(t\tau-\chi)=(c_{-}kd_{\rm A})\sum_{i}(tq_{i}-1/d_{\rm A})\varphi_{i}, which is the desired decomposition.  ■\,\blacksquare

In quantum theory, corollary 21 is equivalent to the fact that every Hermitian matrix is diagonal in a suitable orthonormal basis. A simple consequence of corollary 21 is the following

For every system A{\rm A} with dA=2d_{\rm A}=2 there is a continuous set of pure states.

We conclude this section with the dual result to the “spectral decomposition” of corollary 21:

Since Φ\Phi is dynamically faithful, this implies (x∣=∑idi(ai∣\left(x\right|=\sum_{i}d_{i}\left(a_{i}\right|, where di:=cidAd_{i}:=c_{i}d_{\rm A}.  ■\,\blacksquare

IX Teleportation revisited

We start by showing a probabilistic teleportation scheme that achieves success probability pA=1/dAp_{\rm A}=1/d_{\rm A} for every system A{\rm A}:

For every system A{\rm A}, probabilistic teleportation can be achieved with probability pA=1/dA2p_{\rm A}=1/{d_{\rm A}^{2}}.

and, since Φ\Phi is dynamically faithful,

as can be verified applying both members of Eq. (19) to Φ\Phi, thus obtaining Eq. (18).  ■\,\blacksquare

IX.2 Isotropic states and effects

Let us start from the definition of the transpose:

[note that the transposition is defined with respect to the given state Φ\Phi]

namely V=Uτ\mathscr{V}=\mathscr{U}^{\tau}. ■\,\blacksquare

The conjugate is just defined as the inverse of the transpose:

We can now give the definition of isotropic pure state (isotropic atomic effect):

An example of isotropic state is Φ\Phi: indeed, by definition of conjugate we have, for every U∈GA\mathscr{U}\in{\mathbf{G}}_{\rm A},

As a consequence, the teleportation effect EE is isotropic: indeed one has

which implies (E∣(U∗⊗U)=(E∣\left(E\right|(\mathscr{U}^{*}\otimes\mathscr{U})=\left(E\right|, since the state Φ⊗Φ\Phi\otimes\Phi is dynamically faithful.

We now show that all isotropic pure states (isotropic atomic effects) are connected to the state Φ\Phi (to the effect EE) through a local reversible transformation.

Since Φ\Phi is dynamically faithful, the above equation implies UVU−1=V\mathscr{U}\mathscr{V}\mathscr{U}^{-1}=\mathscr{V} for every U∈GA\mathscr{U}\in{\mathbf{G}}_{\rm A}.  ■\,\blacksquare

By the duality between states and effects, it is easy to obtain the following:

As a consequence, every isotropic effect is connected to the teleportation effect by a local reversible transformation:

Proof. Since (E∣\left(E\right| and (F∣\left(F\right| are both isotropic, lemma 39 implies that they are both connected to (A∣\left(A\right| through a local reversible transformation, say V\mathscr{V} and W\mathscr{W}, respectively. Therefore, they are connected to each other through the transformation WV−1\mathscr{W}\mathscr{V}^{-1}.  ■\,\blacksquare

IX.3 Dimension of the state space

Due to local distinguishability, any bipartite state Ψ∈St(AB)\Psi\in{\mathsf{St}}({\rm A}{\rm B}) can be written as

with (αl∗∣ ⁣αi)=δil\left(\alpha_{l}^{*}\right|\left.\!\alpha_{i}\right)=\delta_{il} and (βk∗∣ ⁣βj)=δjk\left(\beta^{*}_{k}\right|\left.\!\beta_{j}\right)=\delta_{jk}. Finally, a transformation C\mathscr{C} from A{\rm A} to B{\rm B} can be written as

In this matrix representation, the teleportation diagram of Eq. (19) becomes

where IDAI_{D_{\rm A}} is the identity matrix in dimension DAD_{\rm A}. On the other hand, we also have

We now show that one has the equality, using the following standard lemma:

where OUO_{\mathscr{U}} is an orthogonal (DA−1)×(DA−1)(D_{\rm A}-1)\times(D_{\rm A}-1) matrix.

where dU{\rm d}\mathscr{U} is the Haar measure on the compact group GA{\mathbf{G}}_{\rm A} (see corollary 30 of Ref. purification for the proof of compactness) and ATA^{T} denotes the transpose of AA. By definition, one has PT=PP^{T}=P and OUTPOU=PO^{T}_{\mathscr{U}}PO_{\mathscr{U}}=P for every U∈GA\mathscr{U}\in{\mathbf{G}}_{\rm A}. Let us now define the new representation

obtained from OUO_{\mathscr{U}} by a change of basis in the subspace spanned by {ξi}i=2DA\{\xi_{i}\}_{i=2}^{D_{\rm A}}. With this choice, each matrix OU′O^{\prime}_{\mathscr{U}} is orthogonal:

having used Eq. (23) for the last equality. Using the inequality Tr[MV]≤Tr[IDA]{\rm Tr}[M_{\mathscr{V}}]\leq{\rm Tr}[I_{D_{\rm A}}], that holds for every orthogonal DA×DAD_{\rm A}\times D_{\rm A} matrix, we then obtain

ans, therefore (E∣ ⁣Φ)=1\left(E\right|\left.\!\Phi\right)=1  ■\,\blacksquare

The dimension DAD_{\rm A} of the vector space generated by the states in St(A){\mathsf{St}}({\rm A}) is DA=dA2D_{\rm A}=d_{\rm A}^{2}.

Proof. Using lemma 41 and Eq. (23) we obtain 1=(E∣ ⁣Φ)=Tr[EΦ]=Tr[IDA]/dA2=DA/dA21=\left(E\right|\left.\!\Phi\right)={\rm Tr}[E\Phi]={\rm Tr}[I_{D_{\rm A}}]/d_{\rm A}^{2}=D_{\rm A}/d_{\rm A}^{2}. Hence, DA=dA2D_{\rm A}=d_{\rm A}^{2} ■\,\blacksquare

An interesting consequence of the relation (E∣ ⁣Φ)=1\left(E\right|\left.\!\Phi\right)=1 is the following

Let us write an arbitrary state ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) as ρ=χA+ξ\rho=\chi_{\rm A}+\xi, with (e∣ ⁣ξ)=0\left(e\right|\left.\!\xi\right)=0. Then, the linear map N\mathscr{N} defined by N(ρ)=χA−ξ\mathscr{N}(\rho)=\chi_{\rm A}-\xi is not a physical transformation.

Since this quantity is negative for every dA>1d_{\rm A}>1, the map N\mathscr{N} cannot be a physical transformation.  ■\,\blacksquare

cannot represent a physical transformation of system A{\rm A}.

X Derivation of the qubit

The first step is to prove that the set of normalized states St1(A){\mathsf{St}}_{1}({\rm A}) is a sphere. The idea of the proof is a simple geometric observation: in the ordinary three-dimensional space the sphere is the only compact convex set that has an infinite number of pure states connected by orthogonal transformations. The complete proof is given in the following

Since the convex set of density matrices on a two-dimensional Hilbert space is a sphere, we can represent the states in St1(A){\mathsf{St}}_{1}({\rm A}) as density matrices. Precisely, we can choose three orthogonal axes passing through the center of the sphere and call them x,y,zx,y,z axes, take φ+,k,φ−,k\varphi_{+,k},\varphi_{-,k}, k=x,y,zk=x,y,z to be the two perfectly distinguishable pure states in the direction of the kk-axis and define σk:=φk,+−φk,−\sigma_{k}:=\varphi_{k,+}-\varphi_{k,-}. From the geometry of the sphere we know that any state ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) can be written as

where the pure states are those for which ∑k=x,y,znk2=1\sum_{k=x,y,z}n_{k}^{2}=1. The Bloch representation SρS_{\rho} of quantum state ρ\rho is then obtained by associating the basis vectors χ,σx,σy,σz\chi,\sigma_{x},\sigma_{y},\sigma_{z} to the matrices

Proof. Clearly the matrix EaE_{a} must be positive for every effect aa, since we have Tr[EaSρ]=(a∣ ⁣ρ)≥0{\rm Tr}[E_{a}S_{\rho}]=\left(a\right|\left.\!\rho\right)\geq 0 for every density matrix SρS_{\rho}. Moreover, since we have Tr[EaSρ]=(a∣ ⁣ρ)≤1{\rm Tr}[E_{a}S_{\rho}]=\left(a\right|\left.\!\rho\right)\leq 1 for every density matrix SρS_{\rho}, we must have Ea≤IE_{a}\leq I. Finally, we know that for every couple of perfectly distinguishable pure states φ,φ⊥\varphi,\varphi_{\perp} there exists an atomic effect aa such that (a∣ ⁣φ)=1\left(a\right|\left.\!\varphi\right)=1 and (a∣ ⁣φ⊥)=0\left(a\right|\left.\!\varphi_{\perp}\right)=0. Since the two pure states φ,φ⊥\varphi,\varphi_{\perp} are represented by orthogonal rank-one projectors SφS_{\varphi} and Sφ⊥S_{\varphi_{\perp}}, we must have Ea=SφE_{a}=S_{\varphi}. This proves that the atomic effects are the whole set of positive rank-one projectors. As a consequence, also every positive matrix PP with P≤IP\leq I must represent some effect aa.  ■\,\blacksquare

Note that we proved that all two-dimensional systems A{\rm A} and B{\rm B} in our theory have the same states (St1(A)≃St1(B){\mathsf{St}}_{1}({\rm A})\simeq{\mathsf{St}}_{1}({\rm B})), the same effects (Eff(A)≃Eff(B){\mathsf{Eff}}({\rm A})\simeq{\mathsf{Eff}}({\rm B})), and the same reversible transformations (GA≃GB{\mathbf{G}}_{\rm A}\simeq{\mathbf{G}}_{\rm B}), but we did not show that A{\rm A} and B{\rm B} are operationally equivalent. For example, A{\rm A} and B{\rm B} could be different when we compose them with a third system C{\rm C}: the set of states St1(AC){\mathsf{St}}_{1}({\rm A}{\rm C}) and St1(BC){\mathsf{St}}_{1}({\rm B}{\rm C}) could be non-isomorphic. The fact that every couple of two-dimensional systems A{\rm A} and B{\rm B} are operationally equivalent will be proved later (cf. corollary 40).

We conclude this section with a simple fact that will be very useful later:

(Superposition principle for qubits) Let {φ1,φ2}⊂St1(A)\{\varphi_{1},\varphi_{2}\}\subset{\mathsf{St}}_{1}({\rm A}) be two perfectly distinguishable pure states of a system A{\rm A} with dA=2d_{\rm A}=2. Let {a1,a2}\{a_{1},a_{2}\} be the observation-test such that (ai∣ ⁣φj)=δij\left(a_{i}\right|\left.\!\varphi_{j}\right)=\delta_{ij}. Then, for every probability 0≤p≤10\leq p\leq 1 there exists a pure state ψp∈St1(A)\psi_{p}\in{\mathsf{St}}_{1}({\rm A}) such that

Precisely, the set of pure states ψp∈St1(A)\psi_{p}\in{\mathsf{St}}_{1}({\rm A}) satisfying Eq. (28) is a circle in the Bloch sphere.

Proof. Elementary property of density matrices.  ■\,\blacksquare

XI Projections

In this section we define the projection on a face FF of the convex set St1(A){\mathsf{St}}_{1}({\rm A}) and we prove several properties of projections. The projection on the face FF will be defined as an atomic operation ΠF∈Transf(A)\Pi_{F}\in{\mathsf{Transf}}({\rm A}) that acts as the identity on states in the face FF and that annihilates the states on the orthogonal face F⊥F^{\perp}. In the following we first introduce the concept of orthogonal face, then prove the existence and uniqueness of projections, and finally give some useful results on the projection of a pure state on two orthogonal faces.

In order to introduce the notion of orthogonal face we need first a few elementary results. We start by showing that there is a canonical way to associate a state ωF\omega_{F} to a face FF:

Let FF be a face of the convex set St1(A){\mathsf{St}}_{1}({\rm A}) and let {φi}i=1∣F∣\{\varphi_{i}\}_{i=1}^{|F|} be a maximal set of perfectly distinguishable pure states in FF. Then the state ωF:=1∣F∣∑i=1∣F∣φi\omega_{F}:=\frac{1}{|F|}\sum_{i=1}^{|F|}\varphi_{i} depends only on the face FF and not on the particular set {φi}i=1∣F∣\{\varphi_{i}\}_{i=1}^{|F|}. Morever, FF is the face identified by ωF\omega_{F}

Proof. Suppose that FF is the face identified by ρ\rho and let E∈Transf(A,C)\mathscr{E}\in{\mathsf{Transf}}({\rm A},{\rm C}) (resp. D∈Transf(C,A)\mathscr{D}\in{\mathsf{Transf}}({\rm C},{\rm A})) be the encoding (resp. decoding) in the ideal compression for ρ\rho. By lemma 4 and corollary 12, {Eφi}i=1∣F∣\{\mathscr{E}\varphi_{i}\}_{i=1}^{|F|} is a maximal set of perfectly distinguishable pure states of C{\rm C} and by theorem 9 one has χC=1∣F∣∑i=1∣F∣Eφi\chi_{\rm C}=\frac{1}{|F|}\sum_{i=1}^{|F|}\mathscr{E}\varphi_{i}. Hence, ωF=1∣F∣∑i=1∣F∣φi=1∣F∣∑i=1∣F∣DEφi=DχC\omega_{F}=\frac{1}{|F|}\sum_{i=1}^{|F|}\varphi_{i}=\frac{1}{|F|}\sum_{i=1}^{|F|}\mathscr{D}\mathscr{E}\varphi_{i}=\mathscr{D}\chi_{\rm C}. Since the right-hand side of the equality is independent of the particular set {φi}i=1∣F∣\{\varphi_{i}\}_{i=1}^{|F|}, the state ωF\omega_{F} in the left-hand side is independent too. To prove that FF is the face identified by ωF\omega_{F} it is enough to observe that ωF\omega_{F} is completely mixed relative to FF: this fact follows from the relation ωF=DχC\omega_{F}=\mathscr{D}\chi_{{\rm C}} and from lemma 5  ■\,\blacksquare

We now define the orthogonal complement of the state ωF\omega_{F}:

The orthogonal complement of the state ωF\omega_{F} is the state ωF⊥∈St1(A)∪{0}\omega_{F}^{\perp}\in{\mathsf{St}}_{1}({\rm A})\cup\{0\} defined as follows:

if ∣F∣=dA|F|=d_{\rm A}, then ωF⊥=0\omega_{F}^{\perp}=0

if F<dAF<d_{\rm A}, then ωF⊥\omega_{F}^{\perp} is defined by the relation

An easy way to write the orthogonal complement is

Take a maximal set {φi}i=1∣F∣\{\varphi_{i}\}_{i=1}^{|F|} of perfectly distinguishable pure states in FF and extend it to a maximal set {φi}i=1dA\{\varphi_{i}\}_{i=1}^{d_{\rm A}} of perfectly distinguishable pure states in St1(A){\mathsf{St}}_{1}({\rm A}), then for ∣F∣<dA|F|<d_{\rm A} we have

Proof. By definition, for ∣F∣<dA|F|<d_{\rm A} we have ωF⊥=1dA−∣F∣(dAχA−∣F∣ωF)\omega_{F}^{\perp}=\frac{1}{d_{\rm A}-|F|}(d_{\rm A}\chi_{\rm A}-|F|\omega_{F}). Substituting the expressions χA=1dA∑i=1dAφi\chi_{\rm A}=\frac{1}{d_{\rm A}}\sum_{i=1}^{d_{\rm A}}\varphi_{i} and ωF=1∣F∣∑i=1∣F∣φi\omega_{F}=\frac{1}{|F|}\sum_{i=1}^{|F|}\varphi_{i} we then obtain the thesis.  ■\,\blacksquare

Note, however, that by definition the orthogonal complement ωF⊥\omega_{F}^{\perp} depends only on the face FF and not on the choice of the maximal set in lemma 43.

The states ωF\omega_{F} and ωF⊥\omega_{F}^{\perp} are perfectly distinguishable.

Proof. Take a maximal set {φi}i=1∣F∣\{\varphi_{i}\}_{i=1}^{|F|} of perfectly distinguishable pure states in FF, extend it to a maximal set {φi}i=1dA\{\varphi_{i}\}_{i=1}^{d_{\rm A}}, and take the observation-test such that (ai∣ ⁣φj)=δij\left(a_{i}\right|\left.\!\varphi_{j}\right)=\delta_{ij}. Then the binary test {aF,e−aF}\{a_{F},e-a_{F}\}, defined by aF:=∑i=1∣F∣aia_{F}:=\sum_{i=1}^{|F|}a_{i} distinguishes perfectly between ωF\omega_{F} and ωF⊥\omega_{F}^{\perp}.  ■\,\blacksquare

We say that a state τ∈St1(A)\tau\in{\mathsf{St}}_{1}({\rm A}) is perfectly distinguishable from the face FF if τ\tau is perfectly distinguishable from every state σ\sigma in the face FF. With this definition we have the following

τ\tau is perfectly distinguishable from the face FF

τ\tau is perfectly distinguishable from ωF\omega_{F}

τ\tau belongs to the face identified by ωF⊥\omega_{F}^{\perp}, i.e. τ∈FωF⊥\tau\in F_{\omega_{F}^{\perp}}.

Proof. (1⇔21\Leftrightarrow 2) τ\tau is perfectly distinguishable from ωF\omega_{F} if and only if then there exists a binary test {a,e−a}\{a,e-a\} such that (a∣ ⁣τ)=1\left(a\right|\left.\!\tau\right)=1 and (a∣ ⁣ωF)=0\left(a\right|\left.\!\omega_{F}\right)=0. By lemma 19 this is equivalent to the condition (a∣ ⁣τ)=1\left(a\right|\left.\!\tau\right)=1 and a=ωF0a=_{\omega_{F}}0, that is, τ\tau is distinguishable from any state σ\sigma in the face identified by ωF\omega_{F}, which by definition is FF. (2⇒32\Rightarrow 3) Let {φi}i=1∣F∣\{\varphi_{i}\}_{i=1}^{|F|} be a maximal set of perfectly distinguishable states in FF, ωF=1∣F∣∑i=1∣F∣φi\omega_{F}=\frac{1}{|F|}\sum_{i=1}^{|F|}\varphi_{i}, and let {φi}i=∣F∣+1k\{\varphi_{i}\}_{i=|F|+1}^{k} be the maximal set of perfectly distinguishable pure states in the spectral decomposition τ=∑i=∣F∣+1kpiφi\tau=\sum_{i=|F|+1}^{k}p_{i}\varphi_{i}, with pi>0p_{i}>0 for every i=∣F∣+1,…,ki=|F|+1,\dots,k. Since τ\tau is perfectly distinguishable from ωF\omega_{F}, by lemma 23 we have that the states {φi}i=1k\{\varphi_{i}\}_{i=1}^{k} are all perfectly distinguishable. Let us extend this set to a maximal set {φi}i=1dA\{\varphi_{i}\}_{i=1}^{d_{\rm A}}. By lemma 43 have ωF⊥=1dA−∣F∣∑i=∣F∣+1dAφi\omega_{F}^{\perp}=\frac{1}{d_{\rm A}-|F|}\sum_{i=|F|+1}^{d_{\rm A}}\varphi_{i}. Hence, all the states {φi}i=∣F∣+1dA\{\varphi_{i}\}_{i=|F|+1}^{d_{\rm A}} are in the face FωF⊥F_{\omega_{F}^{\perp}}. Since τ\tau is a mixture of these states, it also belongs to the face FωF⊥F_{\omega_{F}^{\perp}}. (3⇒23\Rightarrow 2) Since ωF\omega_{F} and ωF⊥\omega_{F}^{\perp} are perfectly distinguishable, if τ\tau belongs to the face identified by ωF⊥\omega_{F}^{\perp}, then by lemma 22 τ\tau is perfectly distinguishable from ωF\omega_{F}.  ■\,\blacksquare

If ρ\rho is perfectly distinguishable from σ\sigma and from τ\tau, then ρ\rho is perfectly distinguishable from any convex mixture of σ\sigma and τ\tau.

Proof. Let FF be the face identified by ρ\rho. Then by lemma 44 we have σ,τ∈FωF⊥\sigma,\tau\in F_{\omega_{F}^{\perp}}. Since FωF⊥F_{\omega_{F}^{\perp}} is a convex set, any mixture of σ\sigma and τ\tau belongs to it. By lemma 44, this means that any mixture of σ\sigma and τ\tau is perfectly distinguishable from ρ\rho.  ■\,\blacksquare

We are now ready to give the definition of orthogonal face:

The orthogonal face F⊥F^{\perp} is the set of all states that are perfectly distinguishable from the face FF.

By lemma 44 it is clear that F⊥F^{\perp} is the face identified by ωF⊥\omega_{F}^{\perp}, that is F⊥=FωF⊥F^{\perp}=F_{\omega_{F}^{\perp}}.

In the following we list few elementary facts about orthogonal faces:

χA=∣F∣dAωF+∣F⊥∣dAωF⊥\chi_{\rm A}=\frac{|F|}{d_{\rm A}}\omega_{F}+\frac{|F^{\perp}|}{d_{\rm A}}\omega_{F^{\perp}}

Proof. Item 1. If ∣F∣=dA|F|=d_{\rm A} the thesis is obvious. If ∣F∣<dA|F|<d_{\rm A}, take a maximal set {φi}i=1∣F∣\{\varphi_{i}\}_{i=1}^{|F|} (resp. {φj}j=∣F∣+1∣F∣+∣F⊥∣\{\varphi_{j}\}_{j=|F|+1}^{|F|+|F^{\perp}|}) of perfectly distinguishable pure states in FF (resp. F⊥F^{\perp}). Hence we have

By corollary 32 the states ωF\omega_{F} and ωF⊥\omega_{F^{\perp}} are perfectly distinguishable. Hence, the states {φi}i=1∣F∣+∣F⊥∣\{\varphi_{i}\}_{i=1}^{|F|+|F^{\perp}|} are perfectly distinguishable jointly (lemma 23). Now, we must have ∣F∣+∣F⊥∣=dA|F|+|F^{\perp}|=d_{\rm A}, otherwise there would be a pure state ψ\psi that is perfectly distinguishable from the states {φi}i=1∣F∣+∣F⊥∣\{\varphi_{i}\}_{i=1}^{|F|+|F^{\perp}|}. This implies that ψ\psi belongs to F⊥F^{\perp} and that states {ψ}∪{φj}j=∣F∣+1∣F∣+∣F⊥∣\{\psi\}\cup\{\varphi_{j}\}_{j=|F|+1}^{|F|+|F^{\perp}|} are perfectly distinguishable in F⊥F^{\perp}, in contradiction with the hypotheses that the set {φj}j=∣F∣+1∣F∣+∣F⊥∣\{\varphi_{j}\}_{j=|F|+1}^{|F|+|F^{\perp}|} is maximal in F⊥F^{\perp}. Item 2 Immediate from item 1 and definition 9. Item 3 and 4 Both items follow by comparison of item 2 with Eq. 29. Item 5 By condition 3 of lemma 44, (F⊥)⊥\left(F^{\perp}\right)^{\perp} is the face identified by the state ωF⊥⊥\omega_{F^{\perp}}^{\perp}, which, by item 4, is ωF\omega_{F}. Since the face identified by ωF\omega_{F} is FF, we have (F⊥)⊥=F\left(F^{\perp}\right)^{\perp}=F.  ■\,\blacksquare

We now show that there is a canonical way to associate an effect aFa_{F} to a face FF:

We say that aF∈Eff(A)a_{F}\in{\mathsf{Eff}}({\rm A}) is the effect associated to the face F⊆St1(A)F\subseteq{\mathsf{St}}_{1}({\rm A}) if and only if aF=ωFea_{F}=_{\omega_{F}}e and aF=ωF⊥0a_{F}=_{\omega_{F}^{\perp}}0.

In other words, the definition imposes that (aF∣ ⁣ρ)=1\left(a_{F}\right|\left.\!\rho\right)=1 for every ρ∈F\rho\in F and (aF∣ ⁣σ)=0\left(a_{F}\right|\left.\!\sigma\right)=0 for every σ∈F⊥\sigma\in F^{\perp}.

A state ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) belongs to the face FF if and only if (aF∣ρ)=1(a_{F}|\rho)=1.

Proof. By definition, if ρ\rho belongs to FF, then (aF∣ρ)=1(a_{F}|\rho)=1. Conversely, if (aF∣ρ)=1(a_{F}|\rho)=1, then ρ\rho is perfectly distinguishable from ωF⊥\omega^{\perp}_{F}, because (aF∣ ⁣ωF⊥)=0\left(a_{F}\right|\left.\!\omega_{F}^{\perp}\right)=0. Now, we know that ωF⊥\omega_{F}^{\perp} is equal to ωF⊥\omega_{F^{\perp}} (item 4 of lemma 45). By item 2 of lemma 44 the fact that ρ\rho is perfectly distinguishable from ωF⊥\omega_{F^{\perp}} implies that ρ\rho belongs to (F⊥)⊥\left(F^{\perp}\right)^{\perp}, which is just FF (item 5 of lemma 45).  ■\,\blacksquare

We now show that the effect aFa_{F} associated to the face FF exists and is unique. A preliminary result needed to this purpose is the following:

The effect aFa_{F} must have the form aF=∑i=1∣F∣aia_{F}=\sum_{i=1}^{|F|}a_{i}, where aia_{i} is the atomic effect such that (ai∣ ⁣φi)=1\left(a_{i}\right|\left.\!\varphi_{i}\right)=1 and {φi}i=1∣F∣\{\varphi_{i}\}_{i=1}^{|F|} is a maximal set of perfectly distinguishable pure states in FF.

Proof. By corollary 23 we have that aFa_{F} can be written as (aF∣=∑idi(ai∣\left(a_{F}\right|=\sum_{i}d_{i}\left(a_{i}\right| where {ai}i=1dA\{a_{i}\}_{i=1}^{d_{\rm A}} is a perfectly distinguishing test. Moreover, since aFa_{F} is an effect, we must have di≥0d_{i}\geq 0 forall i=1,…,dAi=1,\dots,d_{\rm A}. Now, by definition we have (aF∣ ⁣ωF⊥)=0\left(a_{F}\right|\left.\!\omega_{F}^{\perp}\right)=0, which implies di(ai∣ ⁣ωF⊥)=0d_{i}\left(a_{i}\right|\left.\!\omega_{F}^{\perp}\right)=0 for every i=1,…,dAi=1,\dots,d_{\rm A}, that is, (ai∣ ⁣ωF⊥)=0\left(a_{i}\right|\left.\!\omega_{F}^{\perp}\right)=0 whenever di≠0d_{i}\not=0. Let us focus on the values of ii for which di≠0d_{i}\not=0. Let φi\varphi_{i} be the pure state such that (ai∣ ⁣φi)=1\left(a_{i}\right|\left.\!\varphi_{i}\right)=1. The condition (ai∣ ⁣ωF⊥)=0\left(a_{i}\right|\left.\!\omega_{F}^{\perp}\right)=0 implies that φi\varphi_{i} is perfectly distinguishable from ωF⊥\omega_{F}^{\perp}. Therefore, φi\varphi_{i} belongs to (F⊥)⊥(F^{\perp})^{\perp}, which is FF. Since by definition we must have (aF∣ ⁣φi)=1\left(a_{F}\right|\left.\!\varphi_{i}\right)=1, this also implies that di=1d_{i}=1. In summary, we proved that aF=∑i′aia_{F}=\sum^{\prime}_{i}a_{i} where the prime means that the sum is restricted to those values of ii such that φi∈F\varphi_{i}\in F. The condition aF=ωFea_{F}=_{\omega_{F}}e also implies that the number of terms in the sum must be exactly ∣F∣|F|. The thesis is then proved by suitably relabelling the effects {ai}i=1dA\{a_{i}\}_{i=1}^{d_{\rm A}}, in such a way that φi\varphi_{i} belongs to FF for every i=1,…,∣F∣i=1,\dots,|F|.  ■\,\blacksquare

The effect aFa_{F} associated to the face FF is unique.

Proof. Suppose that aF=∑i=1∣F∣aia_{F}=\sum_{i=1}^{|F|}a_{i} and aF′=∑i=1∣F∣ai′a_{F}^{\prime}=\sum_{i=1}^{|F|}a_{i}^{\prime} are two effects associated to the face FF, both written as in lemma 47. Let {φi}i=1∣F∣\{\varphi_{i}\}_{i=1}^{|F|} (resp. {φi′}i=1∣F∣\{\varphi_{i}^{\prime}\}_{i=1}^{|F|}) be the maximal set of perfectly distinguishable pure states in FF such that (ai∣ ⁣φi)=1\left(a_{i}\right|\left.\!\varphi_{i}\right)=1 for every i=1,…,∣F∣i=1,\dots,|F| (resp. (ai′∣ ⁣φi′)=1\left(a^{\prime}_{i}\right|\left.\!\varphi^{\prime}_{i}\right)=1 for every i=1,…,∣F∣i=1,\dots,|F|), and let {ψj}j=1∣F⊥∣\{\psi_{j}\}_{j=1}^{|F^{\perp}|} be a maximal set of perfectly distinguishable pure states in F⊥F^{\perp}. Since ωF\omega_{F} and ωF⊥\omega_{F}^{\perp} are perfectly distinguishable, the states {φi}i=1∣F∣∪{ψj}j=1∣F⊥∣\{\varphi_{i}\}_{i=1}^{|F|}\cup\{\psi_{j}\}_{j=1}^{|F^{\perp}|} (resp. {φi′}i=1∣F∣∪{ψj}j=1∣F⊥∣\{\varphi^{\prime}_{i}\}_{i=1}^{|F|}\cup\{\psi_{j}\}_{j=1}^{|F^{\perp}|} are perfectly distinguishable (lemma 23). Moreover, the set is maximal since ∣F∣+∣F⊥∣=dA|F|+|F^{\perp}|=d_{\rm A}. Let bjb_{j} be the atomic effect such that (bj∣ ⁣ψj)=1\left(b_{j}\right|\left.\!\psi_{j}\right)=1. Then, the test that distinguishes the states {φi}i=1∣F∣∪{ψj}j=1∣F⊥∣\{\varphi_{i}\}_{i=1}^{|F|}\cup\{\psi_{j}\}_{j=1}^{|F^{\perp}|} (resp. {φi′}i=1∣F∣∪{ψj}j=1∣F⊥∣\{\varphi^{\prime}_{i}\}_{i=1}^{|F|}\cup\{\psi_{j}\}_{j=1}^{|F^{\perp}|} is given by {ai}i=1∣F∣∪{bj}j=1∣F⊥∣\{a_{i}\}_{i=1}^{|F|}\cup\{b_{j}\}_{j=1}^{|F^{\perp}|} (resp. {ai′}i=1∣F∣∪{bj}j=1∣F⊥∣\{a^{\prime}_{i}\}_{i=1}^{|F|}\cup\{b_{j}\}_{j=1}^{|F^{\perp}|} and its normalization reads

By comparison we obtain aF=aF′a_{F}=a_{F}^{\prime}.  ■\,\blacksquare

XI.2 Projections

We are now in position to define the projection on a face:

Let FF be a face of St1(A){\mathsf{St}}_{1}({\rm A}). A projection on the face FF is an atomic transformation ΠF\Pi_{F} such that

ΠF=ωFIA\Pi_{F}=_{\omega_{F}}\mathscr{I}_{\rm A}

When FF is the face identified by a pure state φ∈St1(A)\varphi\in{\mathsf{St}}_{1}({\rm A}), we have F={φ}F=\{\varphi\} and call Π{φ}\Pi_{\{\varphi\}} a projection on the pure state φ\varphi.

The first condition in definition 12 means that the projection ΠF\Pi_{F} does not disturb the states in the face FF. The second condition means that ΠF\Pi_{F} annihilates all states in the orthogonal face F⊥F^{\perp}. As a notation, we will indicate with ΠF⊥\Pi_{F}^{\perp} the projection on the face F⊥F^{\perp}, that is, we will use the definition ΠF⊥:=ΠF⊥\Pi_{F}^{\perp}:=\Pi_{F^{\perp}}.

An equivalent condition for ΠF\Pi_{F} to be a projection on the face FF is the following:

Let {φi}i=1dA\{\varphi_{i}\}_{i=1}^{d_{\rm A}} be a maximal set of perfectly distinguishable pure states for system A{\rm A}. The transformation ΠF\Pi_{F} in Transf(A){\mathsf{Transf}}({\rm A}) is a projection on the face generated by the subset {φi}i=1∣F∣\{\varphi_{i}\}_{i=1}^{|F|} if and only if

ΠF=ωFIA\Pi_{F}=_{\omega_{F}}\mathscr{I}_{\rm A}

Proof. The condition is clearly necessary, since by Definition 12 ΠF∣φl)=0\Pi_{F}|\varphi_{l})=0 for l>∣F∣l>|F|. On the other hand, if ΠF∣φl)=0\Pi_{F}|\varphi_{l})=0 for l>∣F∣l>|F| then by definition of ωF⊥\omega^{\perp}_{F} we have ΠF∣ωF⊥)=0\Pi_{F}|\omega^{\perp}_{F})=0, and, therefore ΠF=ωF⊥0\Pi_{F}=_{\omega_{F}^{\perp}}0. ■\,\blacksquare

Proof. ΠF⊗IB\Pi_{F}\otimes\mathscr{I}_{\rm B} is atomic, being the product of two atomic transformations. We now show that ΠF⊗IB=ωF⊗χBIA⊗IB\Pi_{F}\otimes\mathscr{I}_{\rm B}=_{\omega_{F}\otimes\chi_{\rm B}}\mathscr{I}_{\rm A}\otimes\mathscr{I}_{\rm B}: Indeed, by the local tomography axiom it is easy to see that every state σ∈FωF⊗χB\sigma\in F_{\omega_{F}\otimes\chi_{\rm B}} can be written as ∣σ)=∑i=1r∑j=1dBσij∣αi)∣βj)\left|\sigma\right)=\sum_{i=1}^{r}\sum_{j=1}^{d_{\rm B}}\sigma_{ij}\left|\alpha_{i}\right)\left|\beta_{j}\right), where {αi}i=1r\{\alpha_{i}\}_{i=1}^{r} is a basis for Span(F)\mathsf{Span}(F) and {βj}j=1dB\{\beta_{j}\}_{j=1}^{d_{\rm B}} is a basis for St1(B){\mathsf{St}}_{1}({\rm B}). Since ΠF=ωFIA\Pi_{F}=_{\omega_{F}}\mathscr{I}_{\rm A}, we have

In the following we will show that for every face FF there exists a unique projection ΠF\Pi_{F} and we will prove several properties of projections. Let us start from an elementary observation:

Let φ\varphi be a pure state in the face F⊆St1(A)F\subseteq{\mathsf{St}}_{1}({\rm A}) and let a∈Eff(A)a\in{\mathsf{Eff}}({\rm A}) be the atomic effect such that (a∣ ⁣φ)=1\left(a\right|\left.\!\varphi\right)=1. If A∈Transf(A)\mathscr{A}\in{\mathsf{Transf}}({\rm A}) is an atomic transformation such that A=ωFIA\mathscr{A}=_{\omega_{F}}\mathscr{I}_{\rm A}, then (a∣A=(a∣\left(a\right|\mathscr{A}=\left(a\right|. Moreover, if aFa_{F} is the effect associated to the face FF, then we have (aF∣A=(aF∣\left(a_{F}\right|\mathscr{A}=\left(a_{F}\right|.

Proof. By lemma 16, the effect (a∣A\left(a\right|\mathscr{A} is atomic. Now, since A∣φ)=∣φ)\mathscr{A}|\varphi)=|\varphi), we have (a∣A∣φ)=(a∣ ⁣φ)=1(a|\mathscr{A}|\varphi)=\left(a\right|\left.\!\varphi\right)=1. However, by theorem 8 (a∣\left(a\right| is the unique atomic effect such that (a∣ ⁣φ)=1\left(a\right|\left.\!\varphi\right)=1. Hence, (a∣A=(a∣(a|\mathscr{A}=(a|. Moreover, writing aFa_{F} as aF=∑i=1∣F∣aia_{F}=\sum_{i=1}^{|F|}a_{i} with (ai∣ ⁣φi)=1\left(a_{i}\right|\left.\!\varphi_{i}\right)=1, φi∈F\varphi_{i}\in F (lemma 47), we obtain (aF∣A=∑i=1∣F∣(ai∣A=∑i=1∣F∣(ai∣=(aF∣\left(a_{F}\right|\mathscr{A}=\sum_{i=1}^{|F|}\left(a_{i}\right|\mathscr{A}=\sum_{i=1}^{|F|}\left(a_{i}\right|=\left(a_{F}\right|. ■\,\blacksquare

When applied to the case of projections, the above lemma gives the following

Let φ\varphi be a pure state in the face F⊆St1(A)F\subseteq{\mathsf{St}}_{1}({\rm A}) and let a∈Eff(A)a\in{\mathsf{Eff}}({\rm A}) be the atomic effect such that (a∣ ⁣φ)=1\left(a\right|\left.\!\varphi\right)=1. Then, we have (a∣ΠF=(a∣(a|\Pi_{F}=(a|. Moreover, if aFa_{F} is the effect associated to the face FF, then we have (aF∣=(aF∣ΠF\left(a_{F}\right|=\left(a_{F}\right|\Pi_{F}.

The counterpart of corollary 34 is given as follows:

Let ψ\psi be a pure state in the face F⊥F^{\perp} and let bb be the atomic effect such that (b∣ ⁣ψ)=1\left(b\right|\left.\!\psi\right)=1. Then, we have (b∣ΠF=0(b|\Pi_{F}=0. Moreover, if aF⊥a^{\perp}_{F} is the effect associated to the face F⊥F^{\perp}, then we have (aF⊥∣ΠF=0\left(a_{F}^{\perp}\right|\Pi_{F}=0.

Proof. By lemma 16, the effect (b∣ΠF\left(b\right|\Pi_{F} is atomic. Hence, (b∣ΠF\left(b\right|\Pi_{F} must be proportional to an atomic effect b′b^{\prime} with ∣ ⁣∣b′∣ ⁣∣=1|\!|b^{\prime}|\!|=1, for some proportionality constant λ∈\lambda\in, that is (b∣ΠF=λ(b′∣\left(b\right|\Pi_{F}=\lambda\left(b^{\prime}\right|. We want to prove that λ\lambda is zero. By contradiction, suppose that λ≠0\lambda\not=0. Let ψ′\psi^{\prime} be the pure state such that (b′∣ ⁣ψ′)=1\left(b^{\prime}\right|\left.\!\psi^{\prime}\right)=1. Now, since ΠF∣ωF⊥)=0\Pi_{F}\left|\omega_{F}^{\perp}\right)=0, we have 0=(b∣ΠF∣ωF⊥)=λ(b′∣ωF⊥)0=(b|\Pi_{F}|\omega^{\perp}_{F})=\lambda(b^{\prime}|\omega_{F}^{\perp}), which implies (b′∣ωF⊥)=0(b^{\prime}|\omega_{F}^{\perp})=0. Hence, ψ′\psi^{\prime} is perfectly distinguishable from ωF⊥\omega_{F}^{\perp}, which in turn implies that ψ′\psi^{\prime} belongs to (F⊥)⊥=F\left(F^{\perp}\right)^{\perp}=F. We then have λ=(b∣ΠF∣ψ′)=(b∣ ⁣ψ′)=0\lambda=\left(b\right|\Pi_{F}\left|\psi^{\prime}\right)=\left(b\right|\left.\!\psi^{\prime}\right)=0 (the last equality follows from the fact that ψ\psi and ψ′\psi^{\prime} belong to F⊥F^{\perp} and FF, respectively, and hence are perfectly distinguishable). This is in contradiction with the assumption λ≠0\lambda\not=0, thus concluding the proof that (b∣ΠF=0\left(b\right|\Pi_{F}=0. Moreover, writing aF⊥a_{F}^{\perp} as aF⊥=∑i=1∣F⊥∣bia_{F}^{\perp}=\sum_{i=1}^{|F^{\perp}|}b_{i} with (bi∣ ⁣ψi)=1\left(b_{i}\right|\left.\!\psi_{i}\right)=1, ψi∈F⊥\psi_{i}\in F^{\perp}, we obtain (aF⊥∣ΠF=∑i=1∣F⊥∣(bi∣ΠF=0\left(a_{F}^{\perp}\right|\Pi_{F}=\sum_{i=1}^{|F^{\perp}|}\left(b_{i}\right|\Pi_{F}=0.  ■\,\blacksquare

Combining corollary 34 and lemma 52 we obtain an important property of projections, expressed by the following:

If ΠF\Pi_{F} is a projection on the face FF, then one has (eA∣ΠF=(aF∣\left(e_{\rm A}\right|\Pi_{F}=\left(a_{F}\right|.

Proof. The thesis follows from corollary 34 and lemma 52 and from the fact that aF+aF⊥=ea_{F}+a_{F}^{\perp}=e.  ■\,\blacksquare

In the following we will see that for every face FF there exists a unique projection. To prove that, let us start from the existence:

For every face FF of St1(A){\mathsf{St}}_{1}({\rm A}) there exists a projection ΠF\Pi_{F}.

Proof. By lemma 18, there exists a system B{\rm B} and an atomic transformation A∈Transf(A,B)\mathscr{A}\in{\mathsf{Transf}}({\rm A},{\rm B}) with (e∣BA=(aF∣(e|_{\rm B}\mathscr{A}=(a_{F}|. Then, if ΨωF∈St(AC)\Psi_{\omega_{F}}\in{\mathsf{St}}({\rm A}{\rm C}) is a purification of ωF\omega_{F}, we can define the state ∣Σ)BC:=(A⊗IC)∣ΨωF)AC|\Sigma)_{{\rm B}{\rm C}}:=(\mathscr{A}\otimes\mathscr{I}_{\rm C})|\Psi_{\omega_{F}})_{{\rm A}{\rm C}}. By lemma 16, Σ\Sigma is a pure state. Moreover, the pure states Σ\Sigma and ΨωF\Psi_{\omega_{F}} have the same marginal on system C{\rm C}: indeed, we have (eB∣∣Σ)=[(eB∣A]∣ΨωF)=(aF∣∣ΨωF)\left(e_{\rm B}\right|\left|\Sigma\right)=[\left(e_{\rm B}\right|\mathscr{A}]\left|\Psi_{\omega_{F}}\right)=\left(a_{F}\right|\left|\Psi_{\omega_{F}}\right) and, by definition, aF=ωFeAa_{F}=_{\omega_{F}}e_{\rm A}, which by theorem 1 implies (aF∣∣ΨωF)=(eA∣∣ΨωF)\left(a_{F}\right|\left|\Psi_{\omega_{F}}\right)=\left(e_{\rm A}\right|\left|\Psi_{\omega_{F}}\right). If φ0\varphi_{0} and ψ0\psi_{0} are two arbitrary pure states of A{\rm A} and B{\rm B}, respectively, the uniqueness of purification stated by Postulate 1 implies that there exists a reversible channel U∈GAB\mathscr{U}\in{\mathbf{G}}_{{\rm A}{\rm B}} such that

Now, take the atomic effect b∈Eff(B)b\in{\mathsf{Eff}}({\rm B}) such that (b∣ψ0)=1(b|\psi_{0})=1, and define the transformation ΠF∈Transf(A)\Pi_{F}\in{\mathsf{Transf}}({\rm A}) as

Applying bb on both sides of Eq. (30) we then obtain

and, therefore, ΠF=ωFIA\Pi_{F}=_{\omega_{F}}\mathscr{I}_{\rm A}. Moreover, the transformation ΠF\Pi_{F} is atomic, being the composition of atomic transformations (lemma 16). Finally, we have ΠF=ωF⊥0\Pi_{F}=_{\omega_{F}^{\perp}}0: indeed, by construction of ΠF\Pi_{F} we have

This implies (eA∣ΠF∣ωF⊥)=(aF∣ ⁣ωF⊥)0\left(e_{\rm A}\right|\Pi_{F}\left|\omega_{F}^{\perp}\right)=\left(a_{F}\right|\left.\!\omega_{F}^{\perp}\right)0, and, therefore, ΠF=ωF⊥0\Pi_{F}=_{\omega_{F}^{\perp}}0. In conclusion, ΠF\Pi_{F} is the desired projection.  ■\,\blacksquare

To prove the uniqueness of the projection ΠF\Pi_{F} we need two auxiliary lemmas, given in the following.

Proof. The state ΦF\Phi_{F} is pure by lemma 16. Let us choose a maximal set of perfectly distinguishable pure states {φi}i=1dA\{\varphi_{i}\}_{i=1}^{d_{\rm A}} such that {φi}i=1∣F∣\{\varphi_{i}\}_{i=1}^{|F|} is maximal in FF. Now, we have

having used that χA=∑i=1dAφi/dA\chi_{\rm A}=\sum_{i=1}^{d_{\rm A}}\varphi_{i}/d_{\rm A} (theorem 9), and the definition of ΠF\Pi_{F}.  ■\,\blacksquare

Let ΠF∈Transf(A)\Pi_{F}\in{\mathsf{Transf}}({\rm A}) be a projection. A transformation C∈Transf(A)\mathscr{C}\in{\mathsf{Transf}}({\rm A}) satisfies C=ωFIA\mathscr{C}=_{\omega_{F}}\mathscr{I}_{\rm A} if and only if

Proof. Let ΦF\Phi_{F} be the purification of ωF\omega_{F} defined in lemma 54. Since C=ωFIA\mathscr{C}=_{\omega_{F}}\mathscr{I}_{\rm A}, we have (C⊗I)∣ΦF)=∣ΦF)(\mathscr{C}\otimes\mathscr{I})\left|\Phi_{F}\right)=\left|\Phi_{F}\right). In other words, we have (CΠF⊗I)∣Φ)=(ΠF⊗I)∣Φ)(\mathscr{C}\Pi_{F}\otimes\mathscr{I})\left|\Phi\right)=(\Pi_{F}\otimes\mathscr{I})\left|\Phi\right). Since Φ\Phi is dynamically faithful, this implies that CΠF=ΠF\mathscr{C}\Pi_{F}=\Pi_{F}. Conversely, Eq. (32) implies that for σ∈FωF\sigma\in F_{\omega_{F}}, C∣σ)=CΠF∣σ)=ΠF∣σ)=∣σ)\mathscr{C}|\sigma)=\mathscr{C}\Pi_{F}|\sigma)=\Pi_{F}|\sigma)=|\sigma), namely C=ωFIA\mathscr{C}=_{\omega_{F}}\mathscr{I}_{{\rm A}}. ■\,\blacksquare

The projection ΠF\Pi_{F} satisfying Definition 12 is unique.

We now show a few simple properties of projections. In the following, given a maximal set of perfectly distinguishable pure states {φi}i=1dA\{\varphi_{i}\}_{i=1}^{d_{\rm A}} and any subset V⊆{1,…,dA}V\subseteq\{1,\dots,d_{\rm A}\} we define (with a slight abuse of notation) ωV:=∑i∈Vφi/∣V∣\omega_{V}:=\sum_{i\in V}\varphi_{i}/|V|, and ΠV\Pi_{V} as the projection on the face FV:=FωVF_{V}:=F_{\omega_{V}}. We will refer to FVF_{V} as the face generated by VV.

For two arbitrary subsets V,W⊆{1,…,dA}V,W\subseteq\{1,\dots,d_{\rm A}\} one has

In particular, if V∩W=∅V\cap W=\emptyset one has ΠVΠW=0\Pi_{V}\Pi_{W}=0.

Proof. First of all, ΠVΠW\Pi_{V}\Pi_{W} is atomic, being the product of two atomic transformations. Moreover, since the face FV∩WF_{V\cap W} is contained in the faces FVF_{V} and FWF_{W}, we have ΠVΠW∣ρ)=ΠV∣ρ)=∣ρ)\Pi_{V}\Pi_{W}\left|\rho\right)=\Pi_{V}\left|\rho\right)=\left|\rho\right) for every ρ∈FV∩W\rho\in F_{V\cap W}. In other words, ΠVΠW=ωV∩WIA\Pi_{V}\Pi_{W}=_{\omega_{V\cap W}}\mathscr{I}_{\rm A}. Moreover, if l∉V∩Wl\not\in V\cap W we have ΠVΠW∣φl)=0\Pi_{V}\Pi_{W}\left|\varphi_{l}\right)=0. By lemma 49 and and by the uniqueness of projections (theorem 14) we then obtain that ΠVΠW\Pi_{V}\Pi_{W} is the projection on the face generated by V∩WV\cap W.  ■\,\blacksquare

Every projection ΠF\Pi_{F} satisfies the identity ΠF2=ΠF\Pi_{F}^{2}=\Pi_{F}.

Proof. Consider a maximal set of perfectly distinguishable pure states {φi}i=1dA\{\varphi_{i}\}_{i=1}^{d_{\rm A}} such that {φi}i∈V\{\varphi_{i}\}_{i\in V} is maximal in FF. In this way FF is the face generated by VV, and, therefore ΠF=ΠV\Pi_{F}=\Pi_{V}. The thesis follows by taking V=WV=W in lemma 56.  ■\,\blacksquare

For every state ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) such that ρ∉F⊥\rho\not\in F^{\perp}, the normalized state ρ′\rho^{\prime} defined by

Proof. By corollary 35, we have (e∣ΠF=(aF∣\left(e\right|\Pi_{F}=(a_{F}|. Since ρ∉F⊥\rho\not\in F^{\perp}, we must have (e∣ΠF∣ρ)=(aF∣ρ)>0(e|\Pi_{F}|\rho)=(a_{F}|\rho)>0, and, therefore, the state ρ′\rho^{\prime} in Eq. (33) is well defined. Moreover, using the definition of ρ′\rho^{\prime} we obtain

having used corollaries 34 and 35 for the last equality. Finally, lemma 46 implies that ρ′\rho^{\prime} belongs to the face FF.  ■\,\blacksquare

Let Π{φ}\Pi_{\{\varphi\}} be the projection on the pure state φ∈St1(A)\varphi\in{\mathsf{St}}_{1}({\rm A}) and aa be the atomic effect such that (a∣ ⁣φ)=1\left(a\right|\left.\!\varphi\right)=1. Then for every state ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) one has Π{φ}∣ρ)=p∣φ)\Pi_{\{\varphi\}}\left|\rho\right)=p\left|\varphi\right) where p=(a∣ ⁣ρ)p=\left(a\right|\left.\!\rho\right).

Proof. Recall that, by corollary 35, we have (a∣=(e∣Π{φ}\left(a\right|=\left(e\right|\Pi_{\{\varphi\}}. If (a∣ ⁣ρ)=0\left(a\right|\left.\!\rho\right)=0 then clearly Π{φ}∣ρ)=0\Pi_{\{\varphi\}}|\rho)=0. Otherwise, the proof is a straightforward application of corollary 37.  ■\,\blacksquare

We conclude the present subsection with a result that will be useful in the next subsection.

An atomic transformation A∈Transf(A)\mathscr{A}\in{\mathsf{Transf}}({\rm A}) satisfies A=ωFIA\mathscr{A}=_{\omega_{F}}\mathscr{I}_{\rm A} if and only if

Conversely, suppose that Eq. (34) is satisfied. Let φ∈F\varphi\in F be a pure state in FF and aa be the atomic effect such that (a∣ ⁣φ)=1\left(a\right|\left.\!\varphi\right)=1. Then, we have

having used the relation (a∣ΠF=(a∣\left(a\right|\Pi_{F}=\left(a\right| (corollary 34). Then, by theorem 7, Aφ=φ\mathscr{A}\varphi=\varphi. Since φ∈F\varphi\in F is arbitrary, this implies A=ωFIA\mathscr{A}=_{\omega_{F}}\mathscr{I}_{\rm A}.  ■\,\blacksquare

XI.3 Projection of a pure state on two orthogonal faces

In Section X we proved a number of results concerning two-dimensional systems. Some properties of two-dimensional systems will be extended to the case of generic systems using the following lemma:

Consider a pure state φ∈St1(A)\varphi\in{\mathsf{St}}_{1}({\rm A}) and two complementary projections ΠF\Pi_{F} and ΠF⊥\Pi^{\perp}_{F}. Then, φ\varphi belongs to the face identified by the state ∣θ):=(ΠF+ΠF⊥)∣φ)\left|\theta\right):=(\Pi_{F}+\Pi^{\perp}_{F})\left|\varphi\right).

Proof. If ΠF∣φ)=0\Pi_{F}\left|\varphi\right)=0 (resp. ΠF⊥∣φ)=0\Pi_{F}^{\perp}\left|\varphi\right)=0), then there is nothing to prove: this means that ΠF⊥∣φ)=∣φ)\Pi_{F}^{\perp}\left|\varphi\right)=\left|\varphi\right) (resp. ΠF∣φ)=∣φ)\Pi_{F}\left|\varphi\right)=\left|\varphi\right)) and the thesis is trivially true. Suppose now that ΠF∣φ)≠0\Pi_{F}\left|\varphi\right)\not=0 and ΠF⊥∣φ)≠0\Pi_{F}^{\perp}\left|\varphi\right)\not=0. Using the notation Π1:=ΠF\Pi_{1}:=\Pi_{F}, Π2:=ΠF⊥\Pi_{2}:=\Pi_{F}^{\perp}, we can define the two pure states ∣φi):=Πi∣φ)/(e∣Πi∣φ)|\varphi_{i}):=\Pi_{i}|\varphi)/\left(e\right|\Pi_{i}|\varphi), i=1,2i=1,2, and the probabilities pi=(e∣Πi∣φ)p_{i}=\left(e\right|\Pi_{i}\left|\varphi\right). In this way we have Πi∣φ)=pi∣φi)\Pi_{i}\left|\varphi\right)=p_{i}\left|\varphi_{i}\right) for i=1,2i=1,2 and θ=p1φ1+p2φ2\theta=p_{1}\varphi_{1}+p_{2}\varphi_{2}. Taking the atomic effect (ai∣\left(a_{i}\right| such that (ai∣ ⁣φi)=1\left(a_{i}\right|\left.\!\varphi_{i}\right)=1 we have aFθ=a1+a2a_{F_{\theta}}=a_{1}+a_{2}, where aFθa_{F_{\theta}} is the effect associated to the face FθF_{\theta}. Recalling that (ai∣Πi=(ai∣\left(a_{i}\right|\Pi_{i}=\left(a_{i}\right| for i=1,2i=1,2 (corollary 34), we then conclude the following

Finally, lemma 46 yields φ∈Fθ\varphi\in F_{\theta}. ■\,\blacksquare

A consequence of lemma 58 is the following

Let φ∈St1(A)\varphi\in{\mathsf{St}}_{1}({\rm A}) be a pure state, a∈Eff(A)a\in{\mathsf{Eff}}({\rm A}) be the unique atomic effect such that (a∣ ⁣φ)=1\left(a\right|\left.\!\varphi\right)=1, and FF be a face in St1(A){\mathsf{St}}_{1}({\rm A}). If ρ\rho is perfectly distinguishable from ΠF∣φ)\Pi_{F}\left|\varphi\right) and from ΠF⊥∣φ)\Pi^{\perp}_{F}\left|\varphi\right) then ρ\rho is perfectly distinguishable from ∣φ)\left|\varphi\right). In particular, one has (a∣ ⁣ρ)=0\left(a\right|\left.\!\rho\right)=0.

Proof. Since ρ\rho is perfectly distinguishable from ΠF∣φ)\Pi_{F}\left|\varphi\right) and ΠF⊥∣φ)\Pi^{\perp}_{F}\left|\varphi\right), it is also perfectly distinguishable from any convex combination of them (corollary 33). Equivalently, ρ\rho is perfectly distinguishable from the face FθF_{\theta} identified by ∣θ):=ΠF∣φ)+ΠF⊥∣φ)\left|\theta\right):=\Pi_{F}\left|\varphi\right)+\Pi_{F}^{\perp}\left|\varphi\right). In particular, it must be perfectly distinguishable from φ\varphi, which belongs to FθF_{\theta} by virtue of lemma 58. If aa is the atomic effect such that (a∣ ⁣φ)=1\left(a\right|\left.\!\varphi\right)=1, then by lemma 36 we have (a∣ ⁣ρ)=0\left(a\right|\left.\!\rho\right)=0.  ■\,\blacksquare

A technical result that will be useful in the following is:

Let φ∈St1(A)\varphi\in{\mathsf{St}}_{1}({\rm A}) be a pure state such that ΠF∣φ)≠0\Pi_{F}\left|\varphi\right)\not=0 and ΠF⊥∣φ)≠0\Pi_{F}^{\perp}\left|\varphi\right)\not=0. Define the pure states ∣φ1):=ΠF∣φ)/(e∣ΠF∣φ)\left|\varphi_{1}\right):=\Pi_{F}\left|\varphi\right)/\left(e\right|\Pi_{F}\left|\varphi\right) and ∣φ2):=ΠF⊥∣φ)/(e∣ΠF⊥∣φ)\left|\varphi_{2}\right):=\Pi_{F}^{\perp}\left|\varphi\right)/\left(e\right|\Pi^{\perp}_{F}\left|\varphi\right) and the mixed state ∣θ):=(ΠF+ΠF⊥)∣φ)\left|\theta\right):=(\Pi_{F}+\Pi_{F}^{\perp})\left|\varphi\right). Then, we have

Proof. Let {ψi}i=1∣F∣\{\psi_{i}\}_{i=1}^{|F|} be a maximal set of perfectly distinguishable pure states in FF, chosen in such a way that ψ1=φ1\psi_{1}=\varphi_{1}, and let {ψi}i=∣F∣+1dA\{\psi_{i}\}_{i=|F|+1}^{d_{\rm A}} be a maximal set of perfectly distinguishable pure states in F⊥F^{\perp}, chosen in such a way that ψ∣F∣+1=φ2\psi_{|F|+1}=\varphi_{2}. Defining the sets V:={1,…,∣F∣}V:=\{1,\dots,|F|\}, W:={∣F∣+1,…,dA}W:=\{|F|+1,\dots,d_{\rm A}\}, and U:={1,∣F∣+1}U:=\{1,|F|+1\} we then have ΠV=ΠF\Pi_{V}=\Pi_{F}, ΠW=ΠF⊥\Pi_{W}=\Pi_{F}^{\perp} and ΠU=ΠFθ\Pi_{U}=\Pi_{F_{\theta}}. Using lemma 56 we obtain

We conclude this subsection with an important observation about the group of reversible transformations that act as the identity on two orthogonal faces FF and F⊥F^{\perp}. If FF is a face of St1(A){\mathsf{St}}_{1}({\rm A}), let us define GF,F⊥{\mathbf{G}}_{F,F^{\perp}} as the group of all reversible transformations U∈GA\mathscr{U}\in{\mathbf{G}}_{\rm A} such that

For every face F⊂St1(A)F\subset{\mathsf{St}}_{1}({\rm A}) such that F≠{0}F\not=\{0\} and F≠St1(A)F\not={\mathsf{St}}_{1}({\rm A}), the group GF,F⊥{\mathbf{G}}_{F,F^{\perp}} is topologically equivalent to a circle.

We now prove that in fact they are the whole circle. Let Ψ\Psi be a state in Cζ\mathsf{C}_{\zeta}. Since ∣Ψ)\left|\Psi\right) belongs to the face FθF_{\theta}, we obtain

XII The superposition principle

The validity of the superposition principle, proved for two-dimensional systems using the geometry of the Bloch sphere (corollary 31), can be now extended to arbitrary systems thanks to lemma 58.

(Superposition principle for general systems) Let {φi}i=1dA⊆St1(A)\{\varphi_{i}\}_{i=1}^{d_{\rm A}}\subseteq{\mathsf{St}}_{1}({\rm A}) be a maximal set of perfectly distinguishable pure states and {ai}i=1dA\{a_{i}\}_{i=1}^{d_{\rm A}} be the observation-test such that (ai∣ ⁣φj)=δij\left(a_{i}\right|\left.\!\varphi_{j}\right)=\delta_{ij}. Then, for every choice of probabilities {pi}i=1dA\{p_{i}\}_{i=1}^{d_{\rm A}}, pi≥0,∑i=1dApi=1p_{i}\geq 0,\sum_{i=1}^{d_{\rm A}}p_{i}=1 there exists at least one pure state φp∈St1(A)\varphi_{\mathbf{p}}\in{\mathsf{St}}_{1}({\rm A}) such that

where Π{φi}\Pi_{\{\varphi_{i}\}} is the projection on φi\varphi_{i}.

Proof. Let us first prove the equivalence between Eqs. (35) and (36). From Eq. (36) we obtain Eq. (35) using the relation (e∣Π{i}=(ai∣\left(e\right|\Pi_{\{i\}}=\left(a_{i}\right|, which follows from corollary 35. Conversely, from Eq. (35) we obtain Eq. (36) using corollary 38 . Now, we will prove Eq. (35) by induction. The statement for N=2N=2 is proved by corollary 31. Assume that the statement holds for every system B{\rm B} of dimension dB=Nd_{\rm B}=N and suppose that dA=N+1d_{\rm A}=N+1. Let FF be the face identified by ωF=1/N∑i=1Nφi\omega_{F}=1/N\sum_{i=1}^{N}\varphi_{i} and F⊥F^{\perp} be the orthogonal face, identified by the state φN+1\varphi_{N+1}. Now there are two cases: either pN+1=1p_{N+1}=1 or pN+1≠1p_{N+1}\not=1. If pN+1=1p_{N+1}=1, then there is nothing to prove: the desired state is φN+1\varphi_{N+1}. Then, suppose that pN+1≠1p_{N+1}\not=1. Using the induction hypothesis and the compression axiom 3 we can find a state ψq∈F\psi_{\mathbf{q}}\in F such that (ai∣ ⁣ψq)=qi\left(a_{i}\right|\left.\!\psi_{\mathbf{q}}\right)=q_{i}, with qi=pi/(1−pN+1)q_{i}=p_{i}/(1-p_{N+1}), i=1,…,Ni=1,\dots,N. Let us then define a new maximal set of perfectly distinguishable pure states {φi′}i=1N+1\{\varphi^{\prime}_{i}\}_{i=1}^{N+1} , with φ1′=ψq\varphi^{\prime}_{1}=\psi_{\mathbf{q}} and φN+1′=φN+1\varphi^{\prime}_{N+1}=\varphi_{N+1}. Note that one has ωF=1/N∑i=1Nφi′\omega_{F}=1/N\sum_{i=1}^{N}\varphi^{\prime}_{i}, that is, FF is the face generated by the states {φi′}i=1N\{\varphi_{i}^{\prime}\}_{i=1}^{N}. Now consider the two-dimensional face F′F^{\prime} identified by θ=1/2(φ1′+φN+1′)\theta=1/2(\varphi_{1}^{\prime}+\varphi_{N+1}^{\prime}). By corollary 31 (superposition principle for qubits) we know that there exists a pure state φ∈F′\varphi\in F^{\prime} with (a1′∣φ)=1−pN+1(a^{\prime}_{1}|\varphi)=1-p_{N+1} and (aN+1′∣φ)=pN+1(a^{\prime}_{N+1}|\varphi)=p_{N+1}. Let us define V:={1,…,N}V:=\{1,\dots,N\} and W:={1,N+1}W:=\{1,N+1\}. Then, we have ΠF=ΠV\Pi_{F}=\Pi_{V} and ΠF′=ΠW\Pi_{F^{\prime}}=\Pi_{W}, and by lemma 56,

having used corollary 38 for the last equality. Finally, for i=1,…,Ni=1,\dots,N we have

On the other hand we have (aN+1∣φ)=(aN+1′∣φ)=pN+1(a_{N+1}|\varphi)=(a^{\prime}_{N+1}|\varphi)=p_{N+1}.  ■\,\blacksquare

Using the superposition principle and the spectral decomposition of theorem 10 we can now show that every state of system A{\rm A} has a purification in AB{\rm A}{\rm B} provided dB≥dAd_{\rm B}\geq d_{\rm A}:

For every state ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) and for every system B{\rm B} with dB≥dAd_{\rm B}\geq d_{\rm A} there exists a purification of ρ\rho in St1(AB){\mathsf{St}}_{1}({\rm A}{\rm B}).

Proof. Take the spectral decomposition of ρ\rho, given by ρ=∑i=1dApiφi\rho=\sum_{i=1}^{d_{\rm A}}p_{i}\varphi_{i}, where {pi}\{p_{i}\} are probabilities and {φi}i=1dA⊂St1(A)\{\varphi_{i}\}_{i=1}^{d_{\rm A}}\subset{\mathsf{St}}_{1}({\rm A}) is a maximal set of perfectly distinguishable pure states. Let {ψi}i=1dB\{\psi_{i}\}_{i=1}^{d_{\rm B}} be a maximal set of perfectly distinguishable pure states and {ai}i=1dA⊂Eff(A)\{a_{i}\}_{i=1}^{d_{\rm A}}\subset{\mathsf{Eff}}({\rm A}) (resp. {bi}i=1dB⊂Eff(B)\{b_{i}\}_{i=1}^{d_{\rm B}}\subset{\mathsf{Eff}}({\rm B})) be the test such that (ai∣ ⁣φj)=δij\left(a_{i}\right|\left.\!\varphi_{j}\right)=\delta_{ij} (resp. (bi∣ ⁣ψj)=δij\left(b_{i}\right|\left.\!\psi_{j}\right)=\delta_{ij}). Clearly, {φi⊗ψj}\{\varphi_{i}\otimes\psi_{j}\} is a maximal set of perfectly distinguishable pure states for AB{\rm A}{\rm B}. Then, by the superposition principle (theorem 16) there exists a pure state Ψρ\Psi_{\rho} such that (ai⊗bj∣ ⁣Ψρ)=piδij\left(a_{i}\otimes b_{j}\right|\left.\!\Psi_{\rho}\right)=p_{i}\delta_{ij}. Equivalently, we have (bi∣B∣Ψρ)AB=pi ∣φi)A\left(b_{i}\right|_{\rm B}\left|\Psi_{\rho}\right)_{{\rm A}{\rm B}}=p_{i}~\left|\varphi_{i}\right)_{\rm A} for every i=1,…,dAi=1,\dots,d_{\rm A} and (bi∣B∣Ψρ)AB=0\left(b_{i}\right|_{\rm B}\left|\Psi_{\rho}\right)_{{\rm A}{\rm B}}=0 for i>dAi>d_{\rm A}. Summing over ii we then obtain (e∣B∣Ψρ)AB=∑i=1dB(bi∣B∣Ψρ)AB=∑i=1dApi∣φi)A=∣ρ)A\left(e\right|_{\rm B}\left|\Psi_{\rho}\right)_{{\rm A}{\rm B}}=\sum_{i=1}^{d_{\rm B}}\left(b_{i}\right|_{\rm B}\left|\Psi_{\rho}\right)_{{\rm A}{\rm B}}=\sum_{i=1}^{d_{\rm A}}p_{i}\left|\varphi_{i}\right)_{\rm A}=\left|\rho\right)_{\rm A}.  ■\,\blacksquare

In the terminology of Ref. purification, lemma 61 states that a system B{\rm B} with dB≥dAd_{\rm B}\geq d_{\rm A} is complete for the purification of system A{\rm A}.

As a consequence of lemma 61 we have the following:

XII.2 Equivalence of systems with equal dimension

We are now in position to prove that two systems A{\rm A} and B{\rm B} with the same dimension are operationally equivalent, namely that there is a reversible transformation from A{\rm A} to B{\rm B}. In other words, we prove that the informational dimension classifies the systems of our theory up to operational equivalence. The fact that this property is derived from the principles, rather than being assumed from the start, is one of the important differences of our work with respect to Refs. Har01; DakBru09; Mas10. Another difference is that here the equivalence of systems with the same dimension is proved after the derivation of the qubit, whereas in Refs. Har01; DakBru09; Mas10 the derivation of the qubit requires the equivalence of systems with the same dimension.

(Operational equivalence of systems with equal dimension) Every two systems A{\rm A} an B{\rm B} with dA=dBd_{\rm A}=d_{\rm B} are operationally equivalent.

XII.3 Reversible operations of perfectly distinguishable pure states

An important consequence of the superposition principle is the possibility of transforming an arbitrary maximal set of perfectly distinguishable pure states into another via a reversible transformation:

Let A{\rm A} and B{\rm B} be two systems with dA=dB=:dd_{\rm A}=d_{\rm B}=:d and let {φi}i=1d\{\varphi_{i}\}_{i=1}^{d} (resp. {ψi}i=1d\{\psi_{i}\}_{i=1}^{d}) be a maximal set of perfectly distinguishable pure states in A{\rm A} (resp. B{\rm B}). Then, there exists a reversible transformation U∈Transf(A,B)\mathscr{U}\in{\mathsf{Transf}}({\rm A},{\rm B}) such that U∣φi)=∣ψi)\mathscr{U}\left|\varphi_{i}\right)=\left|\psi_{i}\right).

Combining Eqs. (37), (38), (39) we finally obtain

that is, U∣φi)=∣ψi)\mathscr{U}\left|\varphi_{i}\right)=\left|\psi_{i}\right) for every i=1,…,di=1,\dots,d.  ■\,\blacksquare

XIII Derivation of the density matrix formalism

The goal of this section is to show that our set of axioms implies that

the set of effects is the set of positive matrices bounded by the identity

the pairing between a state and an effect is given by the trace of the product of the corresponding matrices.

Using the result of theorem 3, we will then obtain that all the physical transformations in our theory are exactly the physical transformations allowed in quantum mechanics. This will conclude our derivation of quantum theory.

Let φx,+mn,φx,−mn∈Fmn\varphi_{x,+}^{mn},\varphi_{x,-}^{mn}\in F_{mn} (φy,+mn,φy,−mn∈Fmn\varphi_{y,+}^{mn},\varphi_{y,-}^{mn}\in F_{mn}) be the two perfectly distinguishable states in the direction of the xx-axis (yy-axis) and define

An immediate observation is the following:

Proof. Linear independence is evident from the geometry of the Bloch sphere. Moreover, for l∉{m,n}l\not\in\{m,n\} the states φk,±mn\varphi_{k,\pm}^{mn} are perfectly distinguishable from φl\varphi_{l}, and, therefore (al∣ ⁣σkmn)=0\left(a_{l}\right|\left.\!\sigma_{k}^{mn}\right)=0. If l∈{m,n}l\in\{m,n\}, since the states φk,±mn\varphi_{k,\pm}^{mn}, k=x,yk=x,y lie on the equator of the Bloch sphere, we know that (al∣ ⁣φk,±mn)=1/2\left(a_{l}\right|\left.\!\varphi_{k,\pm}^{mn}\right)=1/2 for k=x,yk=x,y. Hence, (al∣ ⁣σkmn)=12−120\left(a_{l}\right|\left.\!\sigma_{k}^{mn}\right)=\frac{1}{2}-\frac{1}{2}0.  ■\,\blacksquare

Let V⊂{1,…,dA}V\subset\{1,\dots,d_{\rm A}\}, and consider the projection ΠV\Pi_{V}. Then, for m∈Vm\in V and n∉Vn\not\in V, one has ΠV∣σkmn)=0\Pi_{V}|\sigma_{k}^{mn})=0 for k=x,yk=x,y.

Proof. Using lemma 56 and corollary 38 we obtain

Since the face FmnF_{mn} is isomorphic to the Bloch sphere and the state since φk±mn\varphi_{k\pm}^{mn}, k=x,yk=x,y lie on the equator of the Bloch sphere, we know that (am∣φk±mn)=12(a_{m}|\varphi_{k\pm}^{mn})=\frac{1}{2}. This implies

Proof. Since the number of vectors is exactly dA2d_{\rm A}^{2}, to prove that they form a basis it is enough to show that they are linearly independent. Suppose that there exists a vector of coefficients {cm}∪{ckmn}\{c_{m}\}\cup\{c_{k}^{mn}\} such that

Applying the projection Π{m,n}\Pi_{\{m,n\}} on both sides and using lemma 63 we obtain

However, we know that the vectors {φm,φn,σxmn,σymn}\{\varphi_{m},\varphi_{n},\sigma_{x}^{mn},\sigma_{y}^{mn}\} are linearly independent. Consequently, cm=cn=cmnk=0c_{m}=c_{n}=c^{k}_{mn}=0 for all m,n,km,n,k.  ■\,\blacksquare

XIII.2 The matrices

Since the state space St(A){\mathsf{St}}({\rm A}) for system A{\rm A} spans a real vector space of dimension DA=dA2D_{\rm A}=d_{\rm A}^{2}, we can decide to represent the vectors {φm}m=1dA∪{σkmn}n>m=1,…,dA k=x,y\{\varphi_{m}\}_{m=1}^{d_{\rm A}}\cup\{\sigma_{k}^{mn}\}_{n>m=1,\dots,d_{\rm A}\ k=x,y} as Hermitian dA×dAd_{\rm A}\times d_{\rm A} matrices. Precisely, we associate the vector φm\varphi_{m} to the matrix SφmS_{\varphi_{m}} defined by

the vector σxmn\sigma^{mn}_{x} to the matrix

and the vector σymn\sigma^{mn}_{y} to the matrix

where λ\lambda can take the values +1+1 or −1-1. The freedom in the choice of λ\lambda will be useful in subsection XIII.3, where we will introduce the representation of composite systems of two qubits. However, this choice of sign plays no role in the present subsection, and for simplicity we will take the positive sign.

Recall that in principle any orthogonal direction in the plane orthogonal to the zz-axis can be chosen to be the xx-axis. In general, the other possible choices for the xx-axis will lead to matrices of the form

and the corresponding choice for the yy-axis will lead to a matrices of the form

and the expansion coefficients {ρm}m=1dA∪{ρkmn}n>m=1,…,dA; k=x,y\{\rho_{m}\}_{m=1}^{d_{\rm A}}\cup\{\rho^{mn}_{k}\}_{n>m=1,\dots,d_{\rm A};\ k=x,y} are all real. Hence, each state ρ\rho is in one-to-one correspondence with a Hermitian matrix, given by

Since effects are linear functionals on states, they are also represented by Hermitian matrices. We will indicate with EaE_{a} the Hermitian matrix associated to the effect a∈Eff(A)a\in{\mathsf{Eff}}({\rm A}). The matrix EaE_{a} is uniquely defined by the relation:

In the rest of the section we show that the set of matrices {Sρ ∣ ρ∈St1(A)}\{S_{\rho}~|~\rho\in{\mathsf{St}}_{1}({\rm A})\} is the whole set of positive Hermitian matrices with unit trace and that the set of matrices {Ea ∣ a∈Eff(A)}\{E_{a}~|~a\in{\mathsf{Eff}}({\rm A})\} is the set of positive Hermitian matrices bounded by the identity.

The invariant state χA\chi_{\rm A} has matrix representation SχA=IdAdAS_{\chi_{\rm A}}=\frac{I_{d_{\rm A}}}{d_{\rm A}}, where IdAI_{d_{\rm A}} is the identity matrix in dimension dAd_{\rm A}.

Proof. Obvious from the expression χA=1d∑mφm\chi_{\rm A}=\frac{1}{d}\sum_{m}\varphi_{m} and from the matrix representation of the states {φm}m=1dA\{\varphi_{m}\}_{m=1}^{d_{\rm A}} in Eq. (41). ■\,\blacksquare

Let am∈Eff(A)a_{m}\in{\mathsf{Eff}}({\rm A}) be the atomic effect such that (am∣ ⁣φm)=1\left(a_{m}\right|\left.\!\varphi_{m}\right)=1. Then, the effect ama_{m} has matrix representation EamE_{a_{m}} such that Eam=SφmE_{a_{m}}=S_{\varphi_{m}}.

Proof. Let ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) be an arbitrary state. Expanding ρ\rho as in Eq. (46) and using lemma 62 we obtain (am∣ ⁣ρ)=ρm\left(a_{m}\right|\left.\!\rho\right)=\rho_{m}. On the other hand, by Eq. (47) we have that ρm\rho_{m} is the mm-th diagonal element of the matrix SρS_{\rho}: by definition of SφmS_{\varphi_{m}} [eq. (41)], this implies ρm=Tr[SφmSρ]\rho_{m}={\rm Tr}[S_{\varphi_{m}}S_{\rho}]. Now, by construction we have Tr[EamSρ]=(am∣ ⁣ρ)=rhom=Tr[SφmSρ]{\rm Tr}[E_{a_{m}}S_{\rho}]=\left(a_{m}\right|\left.\!\rho\right)=rho_{m}={\rm Tr}[S_{\varphi_{m}}S_{\rho}] for every ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}). Hence, Eam=SφmE_{a_{m}}=S_{\varphi_{m}}.  ■\,\blacksquare

The deterministic effect e∈Eff(A)e\in{\mathsf{Eff}}({\rm A}) has matrix representation Ee=IdAE_{e}=I_{d_{\rm A}}.

Proof. Obvious from the expression e=∑mame=\sum_{m}a_{m}, combined with lemma 66 and Eq. (41). ■\,\blacksquare

For every state ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) one has

Proof. Tr[Sρ]=Tr[EeSρ]=(e∣ ⁣ρ)=1{\rm Tr}[S_{\rho}]={\rm Tr}[E_{e}S_{\rho}]=\left(e\right|\left.\!\rho\right)=1.  ■\,\blacksquare

The matrix elements of SφS_{\varphi} for a pure state φ∈St1(A)\varphi\in{\mathsf{St}}_{1}({\rm A}) are (Sφ)mn=pmpneiθmn(S_{\varphi})_{mn}=\sqrt{p_{m}p_{n}}e^{i\theta_{mn}}, with ∑m=1dApm=1\sum_{m=1}^{d_{\rm A}}p_{m}=1, θmn∈[0,2π)\theta_{mn}\in[0,2\pi), θmn=0\theta_{mn}=0 and θmn=−θnm\theta_{mn}=-\theta_{nm}.

Proof. First of all, the diagonal elements of SφS_{\varphi} are given by [Sφ]mm=(am∣ ⁣φ)[S_{\varphi}]_{mm}=\left(a_{m}\right|\left.\!\varphi\right) [cf. Eqs. (46) and (47)]. Denoting the mm-th element by pmp_{m}, we clearly have ∑m=1dApm=(e∣ ⁣φ)=1\sum_{m=1}^{d_{\rm A}}p_{m}=\left(e\right|\left.\!\varphi\right)=1. Now, the projection Π{m,n}∣φ)\Pi_{\{m,n\}}|\varphi) is a state in the face FmnF_{mn}, and, by our choice of representation, the corresponding matrix SΠ{m,n}∣φ)S_{\Pi_{\{m,n\}}\left|\varphi\right)} is proportional to a pure qubit state (non-negative rank-one matrix). On the other hand, it is easy to see from Eqs. (46) and (47) that SΠ{m,n}∣φ)S_{\Pi_{\{m,n\}}|\varphi)} is the matrix with the same elements as SφS_{\varphi} in the block corresponding to the qubit (m,n)(m,n) and 0 elsewhere. In order to be positive and rank-one the corresponding 2×22\times 2 sub-matrix must have the off-diagonal elements (Sφ)mn=pmpneiθmn(S_{\varphi})_{mn}=\sqrt{p_{m}p_{n}}e^{i\theta_{mn}}, for some θmn∈[0,2π)\theta_{mn}\in[0,2\pi) with θnm=−θmn\theta_{nm}=-\theta_{mn}. Repeating the same argument for all choices of indices m,nm,n, the thesis follows. ■\,\blacksquare

For a pure state φ∈St1(A)\varphi\in{\mathsf{St}}_{1}({\rm A}), the corresponding atomic effect aφa_{\varphi} such that (aφ∣φ)=1(a_{\varphi}|\varphi)=1 has a matrix representation EφE_{\varphi} with the property that Eφ=SφE_{\varphi}=S_{\varphi}.

Proof. We already know that the statement holds for dA=2d_{\rm A}=2, where we proved the Bloch sphere representation, equivalent to the fact that states and effects are represented as 2×22\times 2 positive complex matrices, with the set of pure states identified with the set of all rank-one projectors. Let us now consider a generic system A{\rm A}. For every m<nm<n, the face FmnF_{mn} generated by {φm,φn}\{\varphi_{m},\varphi_{n}\} can be encoded in a two-dimensional system. Therefore, the matrices SΠ{m,n}∣φ)S_{\Pi_{\{m,n\}}\left|\varphi\right)} and E(a∣Π{m,n}E_{\left(a\right|\Pi_{\{m,n\}}} are positive (also, recall that all matrix elements outside the (m,n)(m,n) block are zero). Let φ⊥(mn)\varphi^{(mn)}_{\perp} be the pure state in the face FmnF_{mn} that is perfectly distinguishable from Π{m,n}∣φ)\Pi_{\{m,n\}}\left|\varphi\right). Note that, since φ⊥(mn)\varphi^{(mn)}_{\perp} belongs to the face FmnF_{mn}, it is also perfectly distinguishable from Π{1,…,dA}∖{m,n}∣φ)\Pi_{\{1,\dots,d_{\rm A}\}\setminus\{m,n\}}\left|\varphi\right). Hence, φ⊥(mn)\varphi^{(mn)}_{\perp} is perfectly distinguishable from φ\varphi and, in particular, (a∣φ⊥(mn))=0(a|\varphi^{(mn)}_{\perp})=0 (lemma 59). This implies the relation

Now, since the matrix E(a∣Π{m,n}E_{\left(a\right|\Pi_{\{m,n\}}} is positive, the above relation implies E(a∣Π{m,n}=cmnSΠ{m,n}∣φ)E_{\left(a\right|\Pi_{\{m,n\}}}=c_{mn}S_{\Pi_{\{m,n\}}\left|\varphi\right)}, where cmn≥0c_{mn}\geq 0. Finally, repeating the argument for all possible values of (m,n)(m,n), we obtain that cmn=cc_{mn}=c for every m,nm,n, that is, Ea=cSφE_{a}=cS_{\varphi}. Taking the trace on both sides we obtain Tr[Ea]=c{\rm Tr}[E_{a}]=c. To prove that c=1c=1, we use the relation Tr[Ea]/dA=(a∣ ⁣χA)=1/dA{\rm Tr}[E_{a}]/d_{{\rm A}}=\left(a\right|\left.\!\chi_{\rm A}\right)=1/d_{\rm A}.  ■\,\blacksquare

We conclude with a simple corollary that will be used in the next subsection:

Let φ∈St1(A)\varphi\in{\mathsf{St}}_{1}({\rm A}) be a pure state and let {γi}i=1r⊂St1(A)\{\gamma_{i}\}_{i=1}^{r}\subset{\mathsf{St}}_{1}({\rm A}) be a set of pure states. If the state φ\varphi can be written as

for some real coefficients {xi}i=1r\{x_{i}\}_{i=1}^{r}, then the atomic effect aa such that (a∣ ⁣φ)=1\left(a\right|\left.\!\varphi\right)=1 is given by

where cic_{i} is the atomic effect such that (ci∣ ⁣γi)=1\left(c_{i}\right|\left.\!\gamma_{i}\right)=1.

Proof. For every ρ∈St(A)\rho\in{\mathsf{St}}({\rm A}) by theorem 18 one has

thus implying the thesis. ■\,\blacksquare

XIII.3 Choice of axes for a two-qubit system

If A{\rm A} and B{\rm B} are two systems with dA=dB=2d_{\rm A}=d_{\rm B}=2, then we can use two different types of matrix representations for the states of the composite systemAB{\rm A}{\rm B}:

The first type of representation is the representation SφS_{\varphi} introduced through lemma 64: here we will refer to it as the standard representation. Note that there are many different representations of this type because for every pair (m,n)(m,n) there is freedom in choice of the xx- and yy-axis [cf. Eqs. (44 ) and (45)]

The second type of representation is the tensor product representation TφT_{\varphi}, defined by the tensor product of matrices representing states of systems A{\rm A} and B{\rm B}: for a state ∣ρ)=∑i,jρij∣αi)∣βj)\left|\rho\right)=\sum_{i,j}\rho_{ij}\left|\alpha_{i}\right)\left|\beta_{j}\right), with αi∈St(A),βj∈St(B)\alpha_{i}\in{\mathsf{St}}({\rm A}),\beta_{j}\in{\mathsf{St}}({\rm B}), we have

We now show a few properties of the tensor representation. Let FAF_{A} denote the matrix corresponding to the effect A∈Eff(AB)A\in{\mathsf{Eff}}({\rm A}{\rm B}) in the tensor representation, that is, the matrix defined by

It is easy to show that the matrix representation for effects must satisfy the analogue of Eq. (48):

Let A∈Eff(AB)A\in{\mathsf{Eff}}({\rm A}{\rm B}) be a bipartite effect, written as (A∣=∑i,jAij(ai∣(bj∣\left(A\right|=\sum_{i,j}A_{ij}\left(a_{i}\right|\left(b_{j}\right|. Then one has

where EaiAE^{\rm A}_{a_{i}} (resp. EbjBE^{\rm B}_{b_{j}}) is the matrix representing the single-qubit effect aia_{i} (resp. bjb_{j}) in the standard representation for qubit A{\rm A} (resp. B{\rm B}).

Proof. For every bipartite state ∣ρ)=∑k,lρkl∣αk)∣βl)\left|\rho\right)=\sum_{k,l}\rho_{kl}\left|\alpha_{k}\right)\left|\beta_{l}\right) one has

which implies the thesis.  ■\,\blacksquare

Let Ψ∈St1(AB)\Psi\in{\mathsf{St}}_{1}({\rm A}{\rm B}) be a pure state and let A∈Eff(AB)A\in{\mathsf{Eff}}({\rm A}{\rm B}) be the atomic effect such that (A∣ ⁣Ψ)=1\left(A\right|\left.\!\Psi\right)=1. Then one has FA=TΨF_{A}=T_{\Psi}.

For every bipartite state ρ∈St1(AB)\rho\in{\mathsf{St}}_{1}({\rm A}{\rm B}), dA=dB=2d_{\rm A}=d_{\rm B}=2 one has Tr[Tρ]=1{\rm Tr}[T_{\rho}]=1.

Hence, EeAA=EeBB=IE^{\rm A}_{e_{\rm A}}=E^{\rm B}_{e_{\rm B}}=I, where II is the 2×22\times 2 identity matrix. By lemma 68, we then have FeA⊗eB=I⊗IF_{e_{\rm A}\otimes e_{\rm B}}=I\otimes I and, therefore Tr[Tρ]=Tr[FeA⊗eBTρ]=(eA⊗eB∣ ⁣ρ)=1{\rm Tr}[T_{\rho}]={\rm Tr}[F_{e_{\rm A}\otimes e_{\rm B}}T_{\rho}]=\left(e_{\rm A}\otimes e_{\rm B}\right|\left.\!\rho\right)=1.  ■\,\blacksquare

Finally, an immediate consequence of local distinguishability is the following:

Then, we have T(U⊗V)τ=(U⊗V)Tτ(U†⊗V†)T_{(\mathscr{U}\otimes\mathscr{V})\tau}=(U\otimes V)T_{\tau}(U^{\dagger}\otimes V^{\dagger}) for every τ∈St1(AB)\tau\in{\mathsf{St}}_{1}({\rm A}{\rm B}).

The rest of this subsection is aimed at showing that, with a suitable choice of matrix representation for system B{\rm B}, the standard representation coincides with the tensor representation, that is, Sρ=TρS_{\rho}=T_{\rho} for every ρ∈St(AB)\rho\in{\mathsf{St}}({\rm A}{\rm B}). This technical result is important because some properties used in our derivation are easily proved in the standard representation, while the property expressed by lemma 69 is easily proved in the tensor representation: it is then essential to show that we can construct a representation that enjoys both properties.

The four states {φm⊗φn}m,n=12\{\varphi_{m}\otimes\varphi_{n}\}_{m,n=1}^{2} are clearly a maximal set of perfectly distinguishable pure states in AB{\rm A}{\rm B}. In the following we will construct the standard representation starting from this set.

For a composite system AB{\rm A}{\rm B} with dA=dB=2d_{\rm A}=d_{\rm B}=2 one can choose the standard representation in such a way that the following equalities hold

Proof. Let us choose single-qubit representations SAS^{\rm A} and SBS^{\rm B} that satisfy Eqs. (41), (42), and (43). On the other hand, choosing tho states {φn⊗φn}\{\varphi_{n}\otimes\varphi_{n}\} in lexicographic order as the four distinguishable states for the standard representation, we have

With this choice, we get Sφm⊗φn=SφmA⊗SφnB=Tφm⊗φnS_{\varphi_{m}\otimes\varphi_{n}}=S^{\rm A}_{\varphi_{m}}\otimes S^{\rm B}_{\varphi_{n}}=T_{\varphi_{m}\otimes\varphi_{n}} for every m,n=1,2m,n=1,2. This proves Eq. (51). Let us now prove Eqs. (52) and (53). Consider the two-dimensional face F11,12F_{11,12}, generated by the states φ1⊗φ1\varphi_{1}\otimes\varphi_{1} and φ1⊗φ2\varphi_{1}\otimes\varphi_{2}. This face is the face identified by the state ω11,12:=φ1⊗χB\omega_{11,12}:=\varphi_{1}\otimes\chi_{\rm B}, and we have F11,12≃{φ1}⊗St1(B)F_{11,12}\simeq\{\varphi_{1}\}\otimes{\mathsf{St}}_{1}({\rm B}). Therefore we can choose the vectors σk11,12\sigma_{k}^{11,12}, k=x,yk=x,y to satisfy the relation σk11,12:=φ1⊗σk\sigma_{k}^{11,12}:=\varphi_{1}\otimes\sigma_{k}, k=x,yk=x,y. Now, in the standard representation we have

[cf. Eqs. (42) and (43)]. This implies Sσk11,12=Sφ11A⊗SσkB=Tφ11⊗σkS_{\sigma^{11,12}_{k}}=S^{\rm A}_{\varphi_{11}}\otimes S^{\rm B}_{\sigma_{k}}=T_{\varphi_{11}\otimes\sigma_{k}}, for k=x,yk=x,y. Repeating the same argument for the face F22,21F_{22,21}, F11,21F_{11,21}, and F21,22F_{21,22} we obtain the proof of Eqs. (52) and (53).  ■\,\blacksquare

In order to prove that, with a suitable choice of axes, the standard representation coincides with the tensor representation—i.e. Sρ=TρS_{\rho}=T_{\rho} for every ρ∈St(AB)\rho\in{\mathsf{St}}({\rm A}{\rm B})—it remains to find a choice of axes such that Sσk⊗σl=Tσk⊗σlS_{\sigma_{k}\otimes\sigma_{l}}=T_{\sigma_{k}\otimes\sigma_{l}}, k=x,yk=x,y. This will be proved in the following.

Let Φ∈St1(AB)\Phi\in{\mathsf{St}}_{1}({\rm A}{\rm B}) be a pure state such that (a1⊗a1∣ ⁣Φ)=(a2⊗a2∣ ⁣Φ)=1/2\left(a_{1}\otimes a_{1}\right|\left.\!\Phi\right)=\left(a_{2}\otimes a_{2}\right|\left.\!\Phi\right)=\frac{1}{/}2 [such a state exists due to the superposition principle]. With a suitable choice of the matrix representation SBS^{\rm B}, the state Φ\Phi is represented by the matrix

where II is the 2×22\times 2 identity matrix. The most general form for TΦT_{\Phi} is then the following

having defined α:=x1\alpha:=x_{1} and β:=(x0−x1)/2\beta:=(x_{0}-x_{1})/2.

Now, by construction the state Φ\Phi satisfies the condition

By definition of the tensor representation, the conditional states (am∣A∣Φ)AB\left(a_{m}\right|_{\rm A}\left|\Phi\right)_{{\rm A}{\rm B}} are described by the diagonal blocks of the matrix TΦT_{\Phi}:

Since the states φ1\varphi_{1} and φ2\varphi_{2} are pure, the above matrices must be be rank-one. Moreover, their trace must be equal to (am⊗eB∣ ⁣Φ)=1/2(eB∣ ⁣φm)=12\left(a_{m}\otimes e_{\rm B}\right|\left.\!\Phi\right)=1/2\left(e_{\rm B}\right|\left.\!\varphi_{m}\right)=\frac{1}{2}, m=1,2m=1,2. Then we have two possibilities. Either i) α=0\alpha=0 and β=12\beta=\frac{1}{2} or ii) α=−β=12\alpha=-\beta=\frac{1}{2}. In the case i), Eq. 54 holds. In the case ii), to prove Eq. (54) we need to change our choice of matrix representation for the qubit B{\rm B}. Precisely, we make the following change:

where σz:=φ1−φ2\sigma_{z}:=\varphi_{1}-\varphi_{2}. Note that the inversion of the axes, sending σk\sigma_{k} to −σk-\sigma_{k} for every k=x,y,zk=x,y,z is not an allowed physical transformation, but this is not a problem here, because Eq. (58) is just a new choice of matrix representation, in which the set of states of system B{\rm B} is still represented by the Bloch sphere.

More concisely, the change of matrix representation SB↦S~BS^{\rm B}\mapsto\widetilde{S}^{\rm B} can be expressed as

Note that in the new representation S~B\widetilde{S}^{{\rm B}} the physical transformation U∗\mathscr{U}^{*} is still represented as S~UρB=U∗S~ρBUT\widetilde{S}^{{\rm B}}_{\mathscr{U}\rho}=U^{*}\widetilde{S}^{{\rm B}}_{\rho}U^{T}: indeed we have

Let us now prove Eq. (55). Using the fact that by definition Tρ⊗τ=(SρA⊗SτB)T_{\rho\otimes\tau}=(S^{{\rm A}}_{\rho}\otimes S^{{\rm B}}_{\tau}) one can directly verify the relation

This is precisely the matrix version of Eq. (55).  ■\,\blacksquare

Note that the choice of SBS^{\rm B} needed in Eq. (54) is compatible with the choice of SBS^{\rm B} needed in lemma 70: indeed, to prove compatibility we only have to show that the representation SBS_{\rm B} used in Eq. (54) has the property [SφmB]rs=δmrδms[S^{\rm B}_{\varphi_{m}}]_{rs}=\delta_{mr}\delta_{ms}, m=1,2m=1,2. This property is automatically guaranteed by the relation (am∣A∣Φ)AB=1/2∣φm)\left(a_{m}\right|_{\rm A}\left|\Phi\right)_{{\rm A}{\rm B}}=1/2\left|\varphi_{m}\right), m=1,2m=1,2 and by Eq. (57) with α=0\alpha=0 and β=1/2\beta=1/2.

In the standard representation the state Φ∈St1(AB)\Phi\in{\mathsf{St}}_{1}({\rm A}{\rm B}) is represented by the matrix

Proof. The thesis follows from theorem 17 and lemma 70.  ■\,\blacksquare

We now define the reversible transformations Ux,π\mathscr{U}_{x,\pi} and Uz,π2\mathscr{U}_{z,\frac{\pi}{2}} as follows

Also, we define the states Ψ,Φz,π2\Psi,{\Phi_{z,\frac{\pi}{2}}}, and Ψz,π2\Psi_{z,\frac{\pi}{2}} as

The states Ψ,Φz,π2\Psi,\Phi_{z,\frac{\pi}{2}}, and Ψz,π2\Psi_{z,\frac{\pi}{2}} have the following tensor representation

Proof. Eq. (72) is obtained from Eq. (54) by explicit calculation using lemma 69 and Eq. (60). Then, the validity of Eq. (72) is easily obtained from Eq. (55) using the relations

The states Ψ,Φz,π2\Psi,\Phi_{z,\frac{\pi}{2}}, and Ψz,π2\Psi_{z,\frac{\pi}{2}} have a standard representation of the form

with θ\theta as in corollary 59, γ∈[0,2π)\gamma\in[0,2\pi) and λ,μ∈{−1,1}\lambda,\mu\in\{-1,1\}.

Proof. Let us start from Ψ\Psi. First, from Eq.(72) it is immediate to obtain (a1⊗a1∣ ⁣Ψ)=(a2⊗a2∣ ⁣Ψ)=0\left(a_{1}\otimes a_{1}\right|\left.\!\Psi\right)=\left(a_{2}\otimes a_{2}\right|\left.\!\Psi\right)=0 and (a1⊗a2∣ ⁣Ψ)=(a2⊗a1∣ ⁣Ψ)=1/2\left(a_{1}\otimes a_{2}\right|\left.\!\Psi\right)=\left(a_{2}\otimes a_{1}\right|\left.\!\Psi\right)=1/2. This gives the diagonal elements of SΨS_{\Psi}. Then, using theorem 17 we obtain that SΨS_{\Psi} must be as in Eq. (63), for some value of γ\gamma. Let us now consider Φz,π2\Phi_{z,\frac{\pi}{2}}. Again, the diagonal elements of the matrix SΦz,π2S_{\Phi_{z,\frac{\pi}{2}}} are obtained from Eq. (72), which in this case yields (a1⊗a1∣ ⁣Φz,π2)=(a2⊗a2∣ ⁣Φz,π2)=1/2\left(a_{1}\otimes a_{1}\right|\left.\!\Phi_{z,\frac{\pi}{2}}\right)=\left(a_{2}\otimes a_{2}\right|\left.\!\Phi_{z,\frac{\pi}{2}}\right)=1/2 and (a1⊗a2∣ ⁣Φz,π2)=(a2⊗a1∣ ⁣Φz,π2)=0\left(a_{1}\otimes a_{2}\right|\left.\!\Phi_{z,\frac{\pi}{2}}\right)=\left(a_{2}\otimes a_{1}\right|\left.\!\Phi_{z,\frac{\pi}{2}}\right)=0. Hence, by theorem 17 we must have

for some value of λ∈[0,2π)\lambda\in[0,2\pi). Now, denote by AA the effect such that (A∣ ⁣Φ)=1\left(A\right|\left.\!\Phi\right)=1. We then have

having used theorem 18, corollary 44, and Eq. (72). Hence, we have Tr[SΦSΦz,π2]=1/2{\rm Tr}[S_{\Phi}S_{\Phi_{z,\frac{\pi}{2}}}]=1/2, which implies λ=θ±π2\lambda=\theta\pm\frac{\pi}{2}, as in Eq. (63). Finally, the same arguments can be used for Ψz,π2\Psi_{z,\frac{\pi}{2}}: The diagonal elements of SΨz,π2S_{\Psi_{z,\frac{\pi}{2}}} are obtained from the relations (a1⊗a1∣ ⁣Ψz,π2)=(a2⊗a2∣ ⁣Ψz,π2)=0\left(a_{1}\otimes a_{1}\right|\left.\!\Psi_{z,\frac{\pi}{2}}\right)=\left(a_{2}\otimes a_{2}\right|\left.\!\Psi_{z,\frac{\pi}{2}}\right)=0 and (a1⊗a2∣ ⁣Ψz,π2)=(a2⊗a1∣ ⁣Ψz,π2)=1/2\left(a_{1}\otimes a_{2}\right|\left.\!\Psi_{z,\frac{\pi}{2}}\right)=\left(a_{2}\otimes a_{1}\right|\left.\!\Psi_{z,\frac{\pi}{2}}\right)=1/2, which follow from Eq. (72). This implies that the matrix SΨz,π2S_{\Psi_{z,\frac{\pi}{2}}} has the form

for some μ∈[0,2π)\mu\in[0,2\pi). The relation Tr[SΨSΨz,π2]=Tr[TΨTΨz,π2]=1/2{\rm Tr}[S_{\Psi}S_{\Psi_{z,\frac{\pi}{2}}}]={\rm Tr}[T_{\Psi}T_{\Psi_{z,\frac{\pi}{2}}}]=1/2 then implies μ=γ±π2\mu=\gamma\pm\frac{\pi}{2}.  ■\,\blacksquare

Let us now consider the four vectors Σx(11,22),Σy(11,22),Σx(12,21),Σy(12,21)\Sigma_{x}^{(11,22)},\Sigma_{y}^{(11,22)},\Sigma_{x}^{(12,21)},\Sigma_{y}^{(12,21)} defined as follows

By the previous results, it is immediate to obtain the matrix representations of these vectors. In the tensor representation, using Eqs. (54) and (72) we obtain

while in the standard representation, using Eqs. (59) and (63), we obtain

Comparing the two matrix representations we are now in position to prove the desired result:

With a suitable choice of axes, one has Sσk⊗σl=Tσk⊗σlS_{\sigma_{k}\otimes\sigma_{l}}=T_{\sigma_{k}\otimes\sigma_{l}} for every k,l=x,yk,l=x,y.

Proof. For the face (11,22)(11,22), using the freedom coming from Eqs. (43) and (44), we redefine the xx and yy axes so that σx(11,22):=Σx(11,22)\sigma^{(11,22)}_{x}:=\Sigma^{(11,22)}_{x} and λσy(11,22):=Σy(11,22)\lambda\sigma^{(11,22)}_{y}:=\Sigma^{(11,22)}_{y}. In this way we have

Likewise, for the face (12,21)(12,21) we redefine the xx and yy axes so that σx(12,21):=Σx(12,21)\sigma^{(12,21)}_{x}:=\Sigma^{(12,21)}_{x} and μσy(12,21):=Σy(12,21)\mu\sigma^{(12,21)}_{y}:=\Sigma^{(12,21)}_{y}, so that we have

Finally, using Eqs. (55), (72), and (64) we have the relations

Since SS and TT coincide on the right-hand side of each equality, they must also coincide on the left-hand side. ■\,\blacksquare

With a suitable choice of axes, the standard representation coincides with the tensor representation, that is, Sρ=TρS_{\rho}=T_{\rho} for every ρ∈St(AB)\rho\in{\mathsf{St}}({\rm A}{\rm B}).

Proof. Combining lemma 70 with lemma 74 we obtain that SS and TT coincide on the tensor products basis B×B\mathcal{B}\times\mathcal{B}, where B={φ1,φ2,σx,σy}\mathcal{B}=\{\varphi_{1},\varphi_{2},\sigma_{x},\sigma_{y}\}. By linearity, SS and TT coincide on every state.  ■\,\blacksquare

From now on, whenever we will consider a composite system AB{\rm A}{\rm B} where A{\rm A} and B{\rm B} are two-dimensional we will adopt the choice that guarantees that the standard representation coincides with the tensor representation.

XIII.4 Positivity of the matrices

In this paragraph we show that the states in our theory can be represented by positive matrices. This amounts to prove that for every system A{\rm A}, the set of states St1(A){\mathsf{St}}_{1}({\rm A}) can be represented as a subset of the set of density matrices in dimension dAd_{\rm A}. This result will be completed in subsection XIII.5, where we will see that, in fact, every density matrix in dimension dAd_{\rm A} corresponds to some state of St1(A){\mathsf{St}}_{1}({\rm A}).

The starting point to prove positivity is the following:

Let A{\rm A} and B{\rm B} be two-dimensional systems. Then, for every pure state Ψ∈St(AB)\Psi\in{\mathsf{St}}({\rm A}{\rm B}) one has SΨ≥0S_{\Psi}\geq 0.

where U\mathscr{U} and V\mathscr{V} are the reversible transformations defined by SUρ=USρU†S_{\mathscr{U}\rho}=US_{\rho}U^{\dagger} and SVρ=VSρV†S_{\mathscr{V}\rho}=VS_{\rho}V^{\dagger}, respectively (U\mathscr{U} and V\mathscr{V} are physical transformations by virtue of corollary 30). Here we used the fact that the standard two-qubit representation coincides with the tensor representation and, therefore, S(U⊗V)Ψ=(U⊗V)SΨ(U⊗V)†S_{(\mathscr{U}\otimes\mathscr{V})\Psi}=(U\otimes V)S_{\Psi}(U\otimes V)^{\dagger}. Denoting the pure state (U⊗V)∣Ψ)(\mathscr{U}\otimes\mathscr{V})\left|\Psi\right) by ∣Ψ′)\left|\Psi^{\prime}\right) we then have

Since by theorem 17 we have [SΨ′]11,22=[SΨ′]11,11[SΨ′]22,22eiθ[S_{\Psi^{\prime}}]_{11,22}=\sqrt{[S_{\Psi^{\prime}}]_{11,11}[S_{\Psi^{\prime}}]_{22,22}}e^{i\theta}, we conclude

Let C{\rm C} be a system of dimension dC=4d_{\rm C}=4. Then, with a suitable choice of matrix representation the pure states of C{\rm C} are represented by positive matrices.

Proof. The system C{\rm C} is operationally equivalent to the composite system AB{\rm A}{\rm B}, where dA=dB=2d_{\rm A}=d_{\rm B}=2. Let U∈Transf(AB,C)\mathscr{U}\in{\mathsf{Transf}}({\rm A}{\rm B},{\rm C}) be the reversible transformation implementing the equivalence. Now, we know that the states of AB{\rm A}{\rm B} are represented by positive matrices. If we define the basis vectors for C{\rm C} by applying U\mathscr{U} to the basis for AB{\rm A}{\rm B}, then we obtain that the states of C{\rm C} are represented by the same matrices representing the states of AB{\rm A}{\rm B}.  ■\,\blacksquare

Let A{\rm A} be a system with dA=3d_{\rm A}=3. With a suitable choice of matrix representation, the matrix SφS_{\varphi} is positive for every pure state φ∈St(A)\varphi\in{\mathsf{St}}({\rm A}).

Proof. Let C{\rm C} be a system with dC=4d_{\rm C}=4. By corollary 47 the states of C{\rm C} are represented by positive matrices. Define the state ω:=13(φ1+φ2+φ3)\omega:=\frac{1}{3}(\varphi_{1}+\varphi_{2}+\varphi_{3}), where {φm}m=14\{\varphi_{m}\}_{m=1}^{4} are four perfectly distinguishable pure states. By the compression axiom, the face FωF_{\omega} can be encoded in a three-dimensional system D{\rm D} (corollary 40). In fact, since D{\rm D} is operationally equivalent to A{\rm A}, the face FωF_{\omega} can be encoded in A{\rm A}. Let E∈Transf(D,A)\mathscr{E}\in{\mathsf{Transf}}({\rm D},{\rm A}) and D∈Transf(A,D)\mathscr{D}\in{\mathsf{Transf}}({\rm A},{\rm D}) be the encoding and decoding operation, respectively. If we define the basis vectors for A{\rm A} by applying E\mathscr{E} to the basis vectors for the face FωF_{\omega}, then we obtain that the states of A{\rm A} are represented by the same matrices representing the states in the face FωF_{\omega}. Since these matrices are positive, the thesis follows.  ■\,\blacksquare

From now on, for every three-dimensional system A{\rm A} we will choose the xx and yy axes so that SρS_{\rho} is positive for every ρ∈St(A)\rho\in{\mathsf{St}}({\rm A}).

Let φ∈St1(A)\varphi\in{\mathsf{St}}_{1}({\rm A}) be a pure state with dA=3d_{\rm A}=3. Then, the corresponding matrix SφS_{\varphi}, given by

Proof. The relation can be trivially satisfied when pi=0p_{i}=0 for some i∈{1,2,3}i\in\{1,2,3\}. Hence, let us assume p1,p2,p3>0p_{1},p_{2},p_{3}>0. Computing the determinant of SφS_{\varphi} one obtains det⁡(Sφ)=2p1p2p3[cos⁡(θ12+θ23−θ13)−1]\det(S_{\varphi})=2p_{1}p_{2}p_{3}[\cos(\theta_{12}+\theta_{23}-\theta_{13})-1]. Since SφS_{\varphi} is positive, we must have det⁡(Sφ)≥0\det(S_{\varphi})\geq 0. If p1,p2,p3>0p_{1},p_{2},p_{3}>0 the only possibility is θ13=θ12+θ23mod  2π\theta_{13}=\theta_{12}+\theta_{23}\mod 2\pi.  ■\,\blacksquare

Corollary 49 can be easily extended to systems of arbitrary dimension. To this purpose, we choose the x−x- and y−y-axes in such a way that the projection of every state ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) on a three-dimensional face is represented by a positive matrix.

Proof. Consider a triple V={p,q,r}⊆{1,…,N}V=\{p,q,r\}\subseteq\{1,\dots,N\}. Then the state ΠV∣φ)\Pi_{V}|\varphi) is proportional to a pure state of a three dimensional system, whose representation SΠVφS_{\Pi_{V}\varphi} is the 3×33\times 3 square sub-matrix of SφS_{\varphi} with elements [Sφ]kl=pkpkeiθkl[S_{\varphi}]_{kl}=\sqrt{p_{k}p_{k}}e^{i\theta_{kl}}, (k,l)∈V×V(k,l)\in V\times V. Now, corollary 49 forces the relation eiθpr=ei(θpq+θqr)e^{i\theta_{pr}}=e^{i(\theta_{pq}+\theta_{qr})}. Since this relation must hold for every choice of the triple V={p,q,r}V=\{p,q,r\}, if we define αp:=θp1\alpha_{p}:=\theta_{p1}, then we have eiθpq=ei(θp1+θ1q)=ei(θp1−θq1)=ei(αp−αq)e^{i\theta_{pq}}=e^{i(\theta_{p1}+\theta_{1q})}=e^{i(\theta_{p1}-\theta_{q1})}=e^{i(\alpha_{p}-\alpha_{q})}. It is then immediate to verify that Sφ=∣v⟩⟨v∣S_{\varphi}=|v\rangle\langle v|, where v=(p1,p2e−iα2,…,pNe−iαN)Tv=(\sqrt{p_{1}},\sqrt{p_{2}}e^{-i\alpha_{2}},\dots,\sqrt{p_{N}}e^{-i\alpha_{N}})^{T}.  ■\,\blacksquare

For every system A{\rm A}, the state space St1(A){\mathsf{St}}_{1}({\rm A}) can be represented as a subset of the set of density matrices in dimension dAd_{\rm A}.

Proof. For every state ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) the matrix SρS_{\rho} is Hermitian by construction, with unit trace by corollary 42, and positive since it is a convex mixture of positive matrices.  ■\,\blacksquare

XIII.5 Quantum theory in finite dimensions

Here we conclude our derivation of quantum theory by showing that every density matrix in dimension dAd_{\rm A} corresponds to some state ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}).

We already know from the superposition principle (lemma 35) that for every choice probabilities {pi}i=1dA\{p_{i}\}_{i=1}^{d_{\rm A}} there is a pure state φ∈St1(A)\varphi\in{\mathsf{St}}_{1}({\rm A}) such that {pi}i=1dA\{p_{i}\}_{i=1}^{d_{\rm A}} are the diagonal elements of SφS_{\varphi}. Thus, the set of density matrices corresponding to pure states contains at least one matrix of the form Sφ=∣v⟩⟨v∣S_{\varphi}=|v\rangle\langle v|, with ∣v⟩=(p1,p2e−iβ2,…,pdAe−iβdA)|v\rangle=(\sqrt{p_{1}},\sqrt{p_{2}}e^{-i\beta_{2}},\dots,\sqrt{p_{d_{\rm A}}}e^{-i\beta_{d_{\rm A}}}). It only remains to prove that every possible choice of phases βi∈[0,2π)\beta_{i}\in[0,2\pi) corresponds to some pure state.

Recall that for a face F⊆St1(A)F\subseteq{\mathsf{St}}_{1}({\rm A}) we defined the group GF,F⊥{\mathbf{G}}_{F,F^{\perp}} to be the group of reversible transformations U∈GA\mathscr{U}\in{\mathbf{G}}_{\rm A} such that U=ωFIA\mathscr{U}=_{\omega_{F}}\mathscr{I}_{\rm A} and U=ωF⊥IA\mathscr{U}=_{\omega_{F}^{\perp}}\mathscr{I}_{\rm A}. We then have the following

Consider a system A{\rm A} with dA=Nd_{\rm A}=N. Let {φi}i=1N⊂St1(A)\{\varphi_{i}\}_{i=1}^{N}\subset{\mathsf{St}}_{1}({\rm A}) be a maximal set of perfectly distinguishable pure states, FF be the face identified by ωF=1/(N−1)∑i=1N−1φi\omega_{F}=1/(N-1)\sum_{i=1}^{N-1}\varphi_{i} and F⊥F^{\perp} its orthogonal face, identified by the state φN\varphi_{N}. If U\mathscr{U} is a reversible transformation in GF,F⊥{\mathbf{G}}_{F,F^{\perp}}, then the action of U\mathscr{U} is given by

where IN−1I_{N-1} is the (N−1)×(N−1)(N-1)\times(N-1) identity matrix and β∈[0,2π)\beta\in[0,2\pi).

Proof. Consider an arbitrary state ρ∈St1(A)\rho\in{\mathsf{St}}_{1}({\rm A}) and its matrix representation

Let us start from the case N=3N=3. Since U∣φi)=∣φi)∀i=1,2,3\mathscr{U}\left|\varphi_{i}\right)=\left|\varphi_{i}\right)\quad\forall i=1,2,3, we have (ai∣U=(ai∣∀i=1,2,3\left(a_{i}\right|\mathscr{U}=\left(a_{i}\right|\quad\forall i=1,2,3 (lemma 51). This implies that U\mathscr{U} sends states in the face F13F_{13} to states in the face F13F_{13}: indeed, for every ρ∈F13\rho\in F_{13} one has (a13∣U∣ρ)=(a13∣ ⁣ρ)=1\left(a_{13}\right|\mathscr{U}\left|\rho\right)=\left(a_{13}\right|\left.\!\rho\right)=1, which implies Uρ∈F13\mathscr{U}\rho\in F_{13} (lemma 46). In other words, the restriction of U\mathscr{U} to the face F13F_{13} is a reversible qubit transformation. Therefore, the action of U\mathscr{U} on a state ρ∈F13\rho\in F_{13} must be given by

for some β∈[0,2π)\beta\in[0,2\pi). Similarly, we can see that U\mathscr{U} sends states in the face F23F_{23} to states in the face F23F_{23}. Hence, for every σ∈F23\sigma\in F_{23} we have

for some β′∈[0,2π)\beta^{\prime}\in[0,2\pi). We now show that eiβ′=eiβe^{i\beta^{\prime}}=e^{i\beta}. To see that, consider a generic state φ∈St1(A)\varphi\in{\mathsf{St}}_{1}({\rm A}), with the property that pi=(ai∣ ⁣φ)>0p_{i}=\left(a_{i}\right|\left.\!\varphi\right)>0 for every i=1,2,3i=1,2,3 (such state exists due to the superposition principle of theorem 16). Writing SφS_{\varphi} as in Eq. (65) we then have

Now, since φ\varphi and Uφ\mathscr{U}\varphi are pure states, by corollary 49 we must have

By comparison we obtain eiβ=eiβ′e^{i\beta}=e^{i\beta^{\prime}}. This proves Eq. (66) for N=3N=3. The proof for N>3N>3 is then immediate: for every three-dimensional face FpqNF_{pqN} the action of U\mathscr{U} is given Eq. (66) for some βpq\beta_{pq}. However, since the two faces FpqNF_{pqN} and Fpq′NF_{pq^{\prime}N} overlap on φp\varphi_{p} we must have βpq=βpq′\beta_{pq}=\beta_{pq^{\prime}}. Similarly βpq=βp′q\beta_{pq}=\beta_{p^{\prime}q}. We conclude that βpq=β\beta_{pq}=\beta for every p,qp,q. This proves Eq. (66) in the general case.  ■\,\blacksquare

We now show that every possible phase shift in Eq. (66) corresponds to a physical transformation:

A transformation U\mathscr{U} of the form of Eq. (66) is a reversible transformation for every β∈[0,2π)\beta\in[0,2\pi).

Proof. By lemma 77, the group GF,F⊥{\mathbf{G}}_{F,F^{\perp}} is a subgroup of U(1)U(1). Now, there are two possibilities: either GF,F⊥{\mathbf{G}}_{F,F\perp} is a (finite) cyclic group or GF,F⊥{\mathbf{G}}_{F,F^{\perp}} coincides with U(1)U(1). However, we know from theorem 15 that GF,F⊥{\mathbf{G}}_{F,F^{\perp}} has a continuum of elements. Hence, GF,F⊥≃U(1){\mathbf{G}}_{F,F^{\perp}}\simeq U(1) and β\beta can take every value in [0,2π)[0,2\pi).  ■\,\blacksquare

An obvious corollary of the previous lemmas is the following

The transformation Uβ\mathscr{U}_{{\boldsymbol{\beta}}} defined by

where UU is the diagonal matrix with diagonal elements (1,eiβ1,…,eiβN−1)(1,e^{i\beta_{1}},\dots,e^{i\beta_{N-1}}) is a reversible transformation for every vector β:=(β2,…,βN)∈[0,2π)×⋯×[0,2π)\boldsymbol{\beta}:=(\beta_{2},\dots,\beta_{N})\in[0,2\pi)\times\dots\times[0,2\pi).

This leads directly to the conclusion of our derivation:

Proof. Let N=dAN=d_{\rm A}. For every choice of probabilities p=(p1,…,pN){\mathbf{p}}=(p_{1},\dots,p_{N}) there exists at least one pure state φp\varphi_{\mathbf{p}} such that pk=(ak∣ ⁣φp)p_{k}=\left(a_{k}\right|\left.\!\varphi_{\mathbf{p}}\right) for every k=1,…,Nk=1,\dots,N(lemma 35 ). This state is represented by the matrix Sφp=∣vp⟩⟨vp∣S_{\varphi_{\mathbf{p}}}=|v_{\mathbf{p}}\rangle\langle v_{\mathbf{p}}| with ∣vp⟩=(p1,p2e−iα2,…,pNe−iαN)T|v_{\mathbf{p}}\rangle=(\sqrt{p_{1}},\sqrt{p_{2}}e^{-i\alpha_{2}},\dots,\sqrt{p_{N}}e^{-i\alpha_{N}})^{T} (lemma 76). Finally, we can transform φp\varphi_{\mathbf{p}} with every reversible transformation Uβ\mathscr{U}_{\boldsymbol{\beta}} defined in Eq. (67), thus obtaining SUβφp=Uβ∣vp⟩⟨vp∣Uβ†S_{\mathscr{U}_{\boldsymbol{\beta}}\varphi_{\mathbf{p}}}=U_{\boldsymbol{\beta}}|v_{\mathbf{p}}\rangle\langle v_{\mathbf{p}}|U_{\boldsymbol{\beta}}^{\dagger} where Uβ∣vp⟩=(p1,p2e−i(α2+β2),…pNe−i(αN+βN))TU_{\boldsymbol{\beta}}|v_{\mathbf{p}}\rangle=(\sqrt{p_{1}},\sqrt{p_{2}}e^{-i(\alpha_{2}+\beta_{2})},\dots\sqrt{p_{N}}e^{-i(\alpha_{N}+\beta_{N})})^{T}. Since p\mathbf{p} and β\boldsymbol{\beta} are arbitrary, this means that every rank-one density matrix corresponds to some pure state. Taking the possible convex mixtures we obtain that every N×NN\times N density matrix corresponds to some state of system A{\rm A}.  ■\,\blacksquare

Choosing a suitable representation ρ↦Sρ\rho\mapsto S_{\rho}, we proved that for every system A{\rm A} the set of normalized states St1(A){\mathsf{St}}_{1}({\rm A}) is the whole set of density matrices in dimension dAd_{\rm A}. Thanks to the purification postulate, this is enough to prove that all the effects Eff(A){\mathsf{Eff}}({\rm A}) and all the transformations Transf(A,B){\mathsf{Transf}}({\rm A},{\rm B}) allowed in our theory are exactly the effects and the transformations allowed in quantum theory. Precisely we have the following

Proof. We proved that our theory has the same normalized states of quantum theory. On the other hand, quantum theory is a theory with purification and in quantum theory the possible physical transformations are quantum operations, i.e. completely positive trace-preserving maps. The thesis then follows from the fact that two theories with purification that have the same set of normalized states are necessarily the same (theorem 3).  ■\,\blacksquare

XIV Conclusion

Quantum theory can be derived from purely informational principles. In particular, it belongs to a broad class of theories of information-processing that includes classical and quantum information theory as special cases. Within this class, quantum theory is identified uniquely by the purification postulate, stating that the ignorance about a part is always compatible with the maximal knowledge of the whole in an essentially unique way. This postulate appears as the origin of the key features of quantum information processing, such as no-cloning, teleportation, and error correction (see also Ref. purification). The general vision underlying the present work is that the main primitives of quantum information processing should be derived directly from the principles, without the abstract mathematics of Hilbert spaces, in order to make the revolutionary aspects of quantum information immediately accessible and to place them in the broader context of the fundamental laws of physics.

Finally, we would like to comment on possible generalizations of our work. As in any axiomatic construction, one can ask how the results change when the principles are modified. For example, one may be interested in relaxing the local distinguishability axiom and in considering theories, like quantum theory on real Hilbert spaces, where global measurements are essential to characterize the state of a composite system. In this direction, the results of Ref. purification suggest that also quantum theory on real Hilbert spaces can be derived from the purification principle, after that the local distinguishability requirement has been suitably relaxed. A possible way to weaken the local distinguishability requirement is to assume only the property of local distinguishability from pure states proposed in Ref. purification: this property states that the probability of distinguishing two states by local measurements is larger than 1/21/2 whenever one of the two states is pure. A different way to relax local distinguishability would be to assume the property of 2-local tomography proposed in Ref. hw, which requires that the state of a multipartite system can be completely characterized using only measurements on bipartite subsystems. This property is equivalent to 2-local distinguishability, defined as the requirement that two different states of a multipartite system can be distinguished with probability of success larger than 1/21/2 using only local measurements or measurements on bipartite subsystems.

A more radical generalization of our work would be to relax the assumption of causality. This would be particularly important for the discussion of quantum gravity scenarios, where the causal structure is not given a priori but is part of the dynamical variables of the theory. In this respect, the contribution of our work is twofold. First, it makes evident how fundamental is the assumption of causality in the ordinary formulation of quantum theory: the whole formalism of quantum states as density matrices with unit trace, quantum measurements as resolutions of the identity, and quantum channels as trace-preserving maps is crucially based on it. Technically speaking, the fact that the normalization of a state is given by a single linear functional (the trace, in quantum theory) is the signature of causality. This partly explains the troubles and paradoxes encountered when trying to combine the formalism of density matrices with non-causal evolutions, as in Deutsch’s model for close timelike curves deutsch; bennett. Moreover, given that the usual notion of normalization has to be abandoned in the non-causal scenario, and that the ordinary quantum formalism becomes inadequate, one may ask in what sense a theory of quantum gravity would be “quantum”. The suggestion coming from our work is that a “quantum” theory is a theory satisfying the purification principle, which can be suitably formulated even in the absence of causality inprep. The discussion of theories with purification in the non-causal scenario is an exciting avenue of future research.

References