Tight uniform continuity bounds for quantum entropies: conditional entropy, relative entropy distance and energy constraints

Andreas Winter

I Introduction

On finite dimensional systems, the von Neumann entropy S(ρ)=−Tr⁡ρlog⁡ρS(\rho)=-\operatorname{Tr}\rho\log\rho is continuous, but this becomes useful only once one has explicit continuity bounds, most significantly the one due to Fannes Fannes , the sharpest form of which is the following:

For states ρ\rho and σ\sigma on a Hilbert space AA of dimension d=∣A∣<∞d=|A|<\infty, if 12∥ρ−σ∥1≤ϵ≤1\frac{1}{2}\|\rho-\sigma\|_{1}\leq\epsilon\leq 1, then

with h(x)=H(x,1−x)=−xlog⁡x−(1−x)log⁡(1−x)h(x)=H(x,1-x)=-x\log x-(1-x)\log(1-x) the binary entropy. A simplified, but universal bound reads

We include a short proof for self-containedness, and also because it deserves to be known better. It seems that it was first found by Petz (Petz:book, , Thm. 3.8), who credits Csiszár for the classical case; the latter seems to have appeared first in Zhang’s paper Zhang (see also Sason ).

We only have to treat the case ϵ≤1−1d\epsilon\leq 1-\frac{1}{d}. We begin with the classical case of two probability distributions pp and qq on the same ground set of dd elements. It is well known, and in fact elementary to confirm, that one can find two jointly distributed random variables, X∼pX\sim p and Y∼qY\sim q (meaning XX is distributed according to the probability law pp, and YY according to qq), with Pr⁡{X≠Y}=12∥p−q∥1≤ϵ\Pr\{X\neq Y\}=\frac{1}{2}\|p-q\|_{1}\leq\epsilon. The crucial idea is to let Pr⁡{X=Y=x}=min⁡(px,qx)\Pr\{X=Y=x\}=\min(p_{x},q_{x}) and to distribute the remaining probability weight suitably off the diagonal. (This is also the minimum probability over all such coupled random variables Zhang . For the reader with a taste for the sophisticated, this is the Kantorovich-Rubinshtein dual formula for the Wasserstein distance in the case of the trivial metric d(x,y)=1d(x,y)=1 for all x≠yx\neq y and d(x,x)=0d(x,x)=0, cf. the broad survey Wasserstein-survey .) Then, by the monotonicity of the Shannon entropy under taking marginals and Fano’s inequality (see CoverThomas ),

and likewise for H(Y)−H(X)H(Y)-H(X). [For the simplified bound, we use H(X∣Y)≤ϵlog⁡d+h(ϵ)H(X|Y)\leq\epsilon\log d+h(\epsilon).]

Next, we reduce the quantum case to the classical one: W.l.o.g. S(ρ)≤S(σ)S(\rho)\leq S(\sigma), and consider the dephasing operation EE in the eigenbasis of ρ\rho, which maps ρ\rho to itself, a diagonal matrix with a probability distribution pp along the diagonal, and σ\sigma to E(σ)E(\sigma), a diagonal matrix with a probability distribution qq along the diagonal. Hence

At the same time, ∥p−q∥1=∥E(ρ)−E(σ)∥1≤∥ρ−σ∥1\|p-q\|_{1}=\|E(\rho)-E(\sigma)\|_{1}\leq\|\rho-\sigma\|_{1}, and so, using the classical case,

Note that the inequality is tight for all ϵ\epsilon and dd, e.g. by σ=∣0⟩ ⁣⟨0∣\sigma=|0\rangle\!\langle 0| and ρ=(1−ϵ)∣0⟩ ⁣⟨0∣+ϵd−1(\openone−∣0⟩ ⁣⟨0∣)\rho=(1-\epsilon)|0\rangle\!\langle 0|+\frac{\epsilon}{d-1}({\openone}-|0\rangle\!\langle 0|). ∎

We are interested in bounds of the above form, i.e. only referring to the trace distance of the states and some general global parameter specifying the system, for a number of entropic quantities, starting with the conditional von Neumann entropy, relative entropy distances from certain sets, etc, which have numerous applications in quantum information theory and quantum statistical physics. Furthermore, and perhaps even more urgently, in situations of infinite dimensional Hilbert spaces, where the above form of the Fannes inequality becomes trivial.

The rest of the paper is structured as follows: in Section II we present and prove an almost tight version of Lemma 1 for the conditional entropy (originally due to Alicki and Fannes AlickiFannes ), then in Section III we generalize the principle behind our proof to a family of relative entropy distance measures from a convex set; in these two sections we also present some illustrative applications of the conditional entropy bounds to two entanglement measures, ERE_{R} and EFE_{F}, as well as their regularizations. In Section IV we expand the methodology of the first part of the paper to infinite dimensional systems, where Fannes-type continuity bounds are obtained under an energy constraint for a broad class of Hamiltonians, and specifically for quantum harmonic oscillators. All entropy continuity bounds are stated as Lemmas, while the applications appear as Corollaries, and two auxiliary results (on “quantum coupling” of density matrices) as Propositions. The absence of Theorems is meant to encourage readers to apply the results presented here.

II Conditional entropy

Alicki and Fannes AlickiFannes proved an extension of the Fannes inequality for the conditional entropy

defined for states ρ\rho on a bipartite (tensor product) Hilbert space A⊗BA\otimes B. While a double application of Lemma 1 would yield such a bound involving both the dimensions of AA and BB, Alicki and Fannes show that if ∥ρ−σ∥1≤ϵ≤1\|\rho-\sigma\|_{1}\leq\epsilon\leq 1, then

In particular, this form is independent of the dimension of BB, which might even be infinite. Note that for classical, Shannon, conditional entropy, an inequality like the above can be obtained from Lemma 1 by convex combination, resulting in a bound like that of Lemma 1 (see below).

The Alicki-Fannes inequality has several applications in quantum information theory, from the proof of asymptotic continuity of entanglement measures — most notably squashed entanglement E-sq and conditional entanglement of mutual information (CEMI) YHW —, to the continuity of quantum channel capacities LeungSmith , and on to the recent discussion of approximately degradable channels Sutter-et-al .

We present a simple proof of the Alicki-Fannes inequality that yields the stronger form of Lemma 2. One of the themes of the present paper, to which we draw attention here, is the use of entropy inequalities in the proofs. In particular, we make use of the concavity of the conditional entropy (which is equivalent to strong subadditivity of the von Neumann entropy) SSA . In the following proof we will specifically rely on two inequalities expressing the concavity of the entropy and the fact that it is not “too concave” KimRuskai :

By introducing a bipartite state ρ=∑ipiρiA⊗∣i⟩ ⁣⟨i∣I\rho=\sum_{i}p_{i}\rho_{i}^{A}\otimes|i\rangle\!\langle i|^{I}, this is seen to be equivalent to

which consists of two applications of strong subadditivity.

For states ρ\rho and σ\sigma on a Hilbert space A⊗BA\otimes B, if 12∥ρ−σ∥1≤ϵ≤1\frac{1}{2}\|\rho-\sigma\|_{1}\leq\epsilon\leq 1, then

If BB is classical in the sense that both ρ\rho and σ\sigma are so-called qc-states, i.e. with an orthonormal basis {∣x⟩}\{|x\rangle\},

and analogously if both are cq-states, then this can be tightened to

The right hand side is monotonic in ϵ\epsilon, hence we may assume 12∥ρ−σ∥1=ϵ\frac{1}{2}\|\rho-\sigma\|_{1}=\epsilon. Let ϵΔ=(ρ−σ)+\epsilon\Delta=(\rho-\sigma)_{+} be the positive part of ρ−σ\rho-\sigma. Note that because this difference is traceless and its trace norm equals 2ϵ2\epsilon, Δ\Delta is a bona fide state. Furthermore,

By letting ϵΔ′\mathchar58=(1+ϵ)ω−ρ\epsilon\Delta^{\prime}\mathrel{\mathop{\mathchar 58\relax}}=(1+\epsilon)\omega-\rho, we obtain another state Δ′\Delta^{\prime}, such that

This is a slightly optimized version of the trick in the proof of Alicki and Fannes AlickiFannes ; cf. MosonyiHiai .

Now, we use the following well-known variational characterization of the conditional entropy:

where D(ρ∥σ)=Tr⁡ρ(log⁡ρ−log⁡σ)D(\rho\|\sigma)=\operatorname{Tr}\rho(\log\rho-\log\sigma) is the quantum relative entropy Umegaki ; OhyaPetz . Choosing an optimal state ξ\xi for ω\omega (which is ξ=ωB\xi=\omega^{B}), we have, from Eq. (2),

where in the third line we have used the concavity upper bound from Eq. (1). Using the other decomposition in Eq. (2), the concavity of the conditional entropy, i.e. the lower bound in Eq. (1), gives

Putting these two bounds together and multiplying by 1+ϵ1+\epsilon, we arrive at

The proof of the general bound is concluded observing that the conditional entropy of any state is bounded between −log⁡∣A∣-\log|A| and +log⁡∣A∣+\log|A|.

For the case of two qc-states or two cq-states as above, note that the states Δ\Delta and Δ′\Delta^{\prime} are of the same, qc-form (cq-form, resp.), and so their conditional entropies are between and log⁡∣A∣\log|A|. ∎

This asymptotically matches Lemma 2 for large dd and small ϵ\epsilon.

As an application of Lemma 2, we can obtain tighter continuity bounds on various quantum channel capacities, simply substituting our tighter bound rather than the original formulation of Alicki and Fannes in the proofs of Leung and Smith LeungSmith .

As a token, we demonstrate a tight version of the asymptotic continuity of the entanglement of formation BDSW ,

for a state ρAB\rho^{AB} on the bipartite system A⊗BA\otimes B, originally due to Nielsen Nielsen-continuity . We then go on to prove asymptotic continuity for its regularization, the entanglement cost HHT-EC ,

which, albeit following the general “telescoping” strategy of LeungSmith , requires a new idea, and seems not to have been known before antisymm . Note that ECE_{C} is different from EFE_{F} Hastings .

Let ρ\rho and σ\sigma be states on the system A⊗BA\otimes B, denoting the smaller of the two dimensions by dd. Then, 12∥ρ−σ∥1≤ϵ\frac{1}{2}\|\rho-\sigma\|_{1}\leq\epsilon implies, with δ=ϵ(2−ϵ)\delta=\sqrt{\epsilon(2-\epsilon)},

Note that these bounds only depend on the smaller of the two dimensions, in contrast to Nielsen-continuity ; in particular, they apply even in the case that one of the two Hilbert spaces is infinite dimensional.

We may assume w.l.o.g. that EF(ρ)≥EF(σ)E_{F}(\rho)\geq E_{F}(\sigma) and ∣B∣≥∣A∣=d|B|\geq|A|=d. Choose a purifying system R≃ABR\simeq AB, and pure states φABR\varphi^{ABR} and ψABR\psi^{ABR} with φAB=ρ\varphi^{AB}=\rho and ψAB=σ=ψR\psi^{AB}=\sigma=\psi^{R} such that

thus 12∥φ−ψ∥1≤δ=1−(1−ϵ)2\frac{1}{2}\|\varphi-\psi\|_{1}\leq\delta=\sqrt{1-(1-\epsilon)^{2}}. Here, F(ρ,σ)=∥ρσ∥1F(\rho,\sigma)=\|\sqrt{\rho}\sqrt{\sigma}\|_{1} is the fidelity between two quantum states, and we have used that it is related to the trace distance by these well-known inequalities FvdG :

By an observation of Schrödinger (which he called “steering”) in the context of his investigation of quantum entanglement Schroedinger:steering , cf. HughstonJozsaWootters , for any convex decomposition σ=∑xpxσx\sigma=\sum_{x}p_{x}\sigma_{x}, there exists a measurement POVM (Mx)(M_{x}) on RR such that pxσx=Tr⁡Rψ(\openoneAB⊗MxR)p_{x}\sigma_{x}=\operatorname{Tr}_{R}\psi({\openone}^{AB}\otimes M_{x}^{R}). Introducing the qc-channel M(ξ)=∑xTr⁡ξMx∣x⟩ ⁣⟨x∣\mathcal{M}(\xi)=\sum_{x}\operatorname{Tr}\xi M_{x}|x\rangle\!\langle x| from RR to a suitable space XX, we then have

Let us choose an optimal decomposition for the purpose of entanglement of formation, and the corresponding POVM and quantum channel, i.e. EF(σ)=S(A∣X)σ~E_{F}(\sigma)=S(A|X)_{\widetilde{\sigma}}. Applying the same to φABR\varphi^{ABR}, we obtain

with qx=Tr⁡φRMxq_{x}=\operatorname{Tr}\varphi^{R}M_{x}. Hence,

Observe that by the contractivity of the trace norm under cptp maps,

Now we can invoke the classical part of Lemma 2,

For the regularization, consider any integer nn and

with Ωt=ρ⊗t−1⊗σ⊗n−t\Omega_{t}=\rho^{\otimes t-1}\otimes\sigma^{\otimes n-t}. The proof will be concluded by showing that for any ΩA′B′\Omega^{A^{\prime}B^{\prime}},

To see this, assume again w.l.o.g. that EF(ρ⊗Ω)≥EF(σ⊗Ω)E_{F}(\rho\otimes\Omega)\geq E_{F}(\sigma\otimes\Omega), and choose a purification υ\upsilon of Ω\Omega on A′B′R′A^{\prime}B^{\prime}R^{\prime}, with R′≃A′B′R^{\prime}\simeq A^{\prime}B^{\prime}. Besides the purification ψABR\psi^{ABR} of σ\sigma, we now need a state (not generally pure) ΘABR\Theta^{ABR} with ΘAB=ρ\Theta^{AB}=\rho and ΘR=ψR\Theta^{R}=\psi^{R}. Proposition 5 below guarantees the existence of such a state with F(ψ,Θ)≥1−ϵF(\psi,\Theta)\geq 1-\epsilon, hence 12∥ψ−Θ∥1≤δ\frac{1}{2}\|\psi-\Theta\|_{1}\leq\delta, once more invoking Eq. (3). As before we choose an optimal decomposition of σAB⊗ΩA′B′\sigma^{AB}\otimes\Omega^{A^{\prime}B^{\prime}} into states on AA′\mathchar58BB′AA^{\prime}\mathrel{\mathop{\mathchar 58\relax}}BB^{\prime}, which we can represent by a POVM and associated cptp map M\mathchar58RR′⟶X\mathcal{M}\mathrel{\mathop{\mathchar 58\relax}}RR^{\prime}\longrightarrow X:

Applying the same map to ω⊗υ\omega\otimes\upsilon, we get

where we observe that, crucially, the same pxp_{x} appear in the expressions for ρ~\widetilde{\rho} and σ~\widetilde{\sigma}. Using ΘR=σT=ψR\Theta^{R}=\sigma^{T}=\psi^{R}, we even have

where in the second line we have used the chain rule S(AA′∣X)=S(A′∣X)+S(A∣A′X)S(AA^{\prime}|X)=S(A^{\prime}|X)+S(A|A^{\prime}X), as well as S(A′∣X)ρ~=S(A′∣X)σ~S(A^{\prime}|X)_{\widetilde{\rho}}=S(A^{\prime}|X)_{\widetilde{\sigma}}. ∎

Given states ρ\rho and σ\sigma on a Hilbert space AA, with 12∥ρ−σ∥1≤ϵ\frac{1}{2}\|\rho-\sigma\|_{1}\leq\epsilon, there exist purifications ∣φ⟩|\varphi\rangle of ρ\rho and ∣ψ⟩|\psi\rangle of σ\sigma, and a (sub-normalized) vector ∣ϑ⟩|\vartheta\rangle, all three in the tensor square Hilbert space A⊗A=\mathchar58A1A2A\otimes A=\mathrel{\mathop{\mathchar 58\relax}}A_{1}A_{2}, such that

Here, ⋅T{\cdot}^{T} denotes the transpose of a matrix with respect to a chosen basis.

Consequently, there exists a state ΘA1A2\Theta^{A_{1}A_{2}} with the properties ΘA1=ρ\Theta^{A_{1}}=\rho and ΘA2=ψA2=σT\Theta^{A_{2}}=\psi^{A_{2}}=\sigma^{T}, and such that F(ψ,Θ), F(φ,Θ)≥1−ϵF(\psi,\Theta),\,F(\varphi,\Theta)\geq 1-\epsilon.

This proposition can be viewed as a quantum analogue of the coupling of random variables X∼pX\sim p and Y∼qY\sim q such that Pr⁡{X≠Y}=12∥p−q∥1\Pr\{X\neq Y\}=\frac{1}{2}\|p-q\|_{1}, on which the proof of Lemma 1 relied.

Fixing an orthonormal basis {∣i⟩}\{|i\rangle\} of AA, and introducing the unnormalized maximally entangled vector

we have the following two “pretty good purifications” Winterize-or-die of ρ\rho and σ\sigma:

the claimed properties of which can be readily checked.

To obtain ∣ϑ⟩|\vartheta\rangle, we use once more Eq. (2) from the proof of Lemma 2:

with states Δ\Delta and Δ′\Delta^{\prime}. Then define

using (Z⊗\openone)∣Φ⟩=(\openone⊗ZT)∣Φ⟩(Z\otimes{\openone})|\Phi\rangle=({\openone}\otimes Z^{T})|\Phi\rangle, with

We claim that ∥X∥, ∥Y∥≤1\|X\|,\,\|Y\|\leq 1. Indeed, ω≥11+ϵρ\omega\geq\frac{1}{1+\epsilon}\rho, so

and similarly YY†≤\openoneYY^{\dagger}\leq{\openone}. From this it follows that

It remains to bound the inner product ∣⟨ψ∣ϑ⟩∣|\langle\psi|\vartheta\rangle| (the other one, ∣⟨φ∣ϑ⟩∣|\langle\varphi|\vartheta\rangle|, is completely analogous):

where we have first used the definitions of ∣ψ⟩|\psi\rangle, ∣ϑ⟩|\vartheta\rangle and ∣Φ⟩|\Phi\rangle, and then the identity between ω\omega and σ\sigma; the fifth line is by triangle inequality, in the sixth we used (1+ϵ)ω≥ρ(1+\epsilon)\omega\geq\rho once more, the operator monotonicity of the square root, and the Hölder inequality ∣Tr⁡XΔ∣≤∥X∥ ∥Δ∥1|\operatorname{Tr}X\Delta|\leq\|X\|\,\|\Delta\|_{1}; in the last step we use the fact that both ρ\rho and Δ\Delta are states and ∥X∥≤1\|X\|\leq 1.

with bona fide states Δ1\Delta_{1} and Δ2\Delta_{2}. It is straightforward to check that the definition

satisfies all requirements on Θ\Theta. ∎

Although the above proof refers to the unnormalized vector ∣Φ⟩|\Phi\rangle, and thus taken literally only makes sense for finite dimensional Hilbert spaces, the proposition remains true also in the infinite dimensional (separable) case. This can be seen either by finite dimensional approximation, or by considering ∣Φ⟩|\Phi\rangle as a formal device to mediate between normalized entangled vectors (∣φ⟩|\varphi\rangle, ∣ψ⟩|\psi\rangle, ∣ϑ⟩|\vartheta\rangle, etc) and Hilbert-Schmidt class operators (ρ\sqrt{\rho}, σ\sqrt{\sigma}, ρ1/2ω−1/2σ1/2\rho^{1/2}\omega^{-1/2}\sigma^{1/2}, etc).

III Relative entropy distances

The same method employed in Lemma 2 can be used to derive asymptotic continuity bounds for the relative entropy distance with respect to any closed convex set CC of states, or more generally positive semidefinite operators, on a Hilbert space AA, cf. Synak-RadtkeHorodecki ),

Unlike Synak-RadtkeHorodecki , CC has to contain only at least one full-rank state, so that DCD_{C} is guaranteed to be finite; in addition, CC should be bounded, so that DCD_{C} is bounded from below. We recover the conditional entropy S(A∣B)ρS(A|B)_{\rho} for a bipartite state ρ\rho on A⊗BA\otimes B, as DC(ρ)D_{C}(\rho) with

For a closed, convex and bounded set CC of positive semidefinite operators, containing at least one of full rank, let

be the largest variation of DCD_{C}. Then, for any two states ρ\rho and σ\sigma with 12∥ρ−σ∥1≤ϵ\frac{1}{2}\|\rho-\sigma\|_{1}\leq\epsilon,

The only modification with respect to the proof of Lemma 2 is that we replace the invocation of concavity of the conditional entropy with the joint convexity of the relative entropy, which makes DCD_{C} a convex functional.

Namely, with ω\omega as in Eq. (2), we have on the one hand,

On the other hand, with an optimal γ∈C\gamma\in C,

Putting these two inequalities together yields the claim of the lemma. ∎

is the set of separable states, we obtain the relative entropy of entanglement of a state ρ\rho on bipartite system A⊗BA\otimes B, ER(ρ)=DSEP(A\mathchar58B)(ρ)E_{R}(\rho)=D_{\text{SEP}(A\mathrel{\mathop{\mathchar 58\relax}}B)}(\rho) relent . Furthermore, we consider its regularization

which is known to be different from ER(ρ)E_{R}(\rho) in general VollbrechtWerner .

(Cf. Donald/Horodecki DonaldHorodecki & Christandl Christandl:PhD ) For any two states ρ\rho and σ\sigma on the composite system A⊗BA\otimes B, denoting the smaller of the dimensions ∣A∣|A|, ∣B∣|B| by dd, 12∥ρ−σ∥1≤ϵ\frac{1}{2}\|\rho-\sigma\|_{1}\leq\epsilon implies

Note that this bound only depends on the smaller of the two dimensions, in contrast to DonaldHorodecki ; in particular, it applies even in the case that one of the two Hilbert spaces is infinite dimensional.

The first bound, on the single-letter ERE_{R} is a direct application of Lemma 7 to the case where CC is the set of all separable states on A⊗BA\otimes B.

For the regularization, consider any integer nn and

with Ωt=ρ⊗t−1⊗σ⊗n−t\Omega_{t}=\rho^{\otimes t-1}\otimes\sigma^{\otimes n-t}. Now for each tt, Lemma 7 gives

with \kappa_{t}=\sup_{\tau,\tau^{\prime}}\bigl{(}E_{R}(\tau\otimes\Omega_{t})-E_{R}(\tau^{\prime}\otimes\Omega_{t})\bigr{)}. To see this, we have to look into the proof of the lemma, and observe that for states ρ⊗Ωt\rho\otimes\Omega_{t} and σ⊗Ωt\sigma\otimes\Omega_{t}, also the auxiliary operators Δ\Delta and Δ′\Delta^{\prime} are of the form τ⊗Ωt\tau\otimes\Omega_{t} and τ′⊗Ωt\tau^{\prime}\otimes\Omega_{t}. However, by LOCC monotonicity,

so that κt≤log⁡d\kappa_{t}\leq\log d. Although we do not need it, the right hand inequality is in fact an equality, ER(Φd⊗Ωt)=log⁡d+ER(Ωt)E_{R}(\Phi_{d}\otimes\Omega_{t})=\log d+E_{R}(\Omega_{t}) extraweak . Thus, we obtain for all nn,

and taking the limit n→∞n\rightarrow\infty concludes the proof. ∎

Again, in Lemma 7 and Corollary 8, the constant in the linear term (proportional to ϵ\epsilon) is essentially best possible, as we see by taking two states maximizing the difference DC(ρ)−DC(σ)D_{C}(\rho)-D_{C}(\sigma), i.e. attaining κ\kappa, since 12∥ρ−σ∥1≤1=\mathchar58ϵ\frac{1}{2}\|\rho-\sigma\|_{1}\leq 1=\mathrel{\mathop{\mathchar 58\relax}}\epsilon.

Lemma 7 improves upon similar-looking general bounds by Synak-Radtke and Horodecki Synak-RadtkeHorodecki , which were subsequently optimized by Mosonyi and Hiai (MosonyiHiai, , Prop. VI.1). The latter paper also explains lucidly (in Sec. VI) that the coefficient 11+ϵ\frac{1}{1+\epsilon} in the convex decomposition of ω\omega in two ways, into ρ\rho and Δ′\Delta^{\prime} and into σ\sigma and Δ\Delta, is optimal, and gives a nice geometric interpretation of ω\omega as a max⁡\max-relative entropy center of ρ\rho and σ\sigma (cf. Kimura-et-al ). Thus, at least following the same strategy one cannot improve the bound any more.

That the regularized relative entropy measure ER∞E_{R}^{\infty} is asymptotically continuous followed previously from its non-lockability HHHO-lock , which it inherits from ERE_{R}. This has been worked out in (antisymm, , Prop. 13), following (Christandl:PhD, , Prop. 3.23), with a different linear term.

It would be interesting to lift the restriction that CC has to be a convex set: Natural examples are the case that CC is the set of all product states in a bipartite (multipartite) system, in which case DCD_{C} becomes the quantum mutual information (multi-information); or the case that CC is the closure of the set of all Gibbs states for a suitable Hamiltonian operator HH,

Both examples have in common that CC is an exponential family (or the closure of one); it is known that at least in some cases DCD_{C} is continuous, but counterexamples of discontinuous behaviour are known WeisKnauf .

IV Bounded energy

If the Hilbert space in the Fannes inequality (Lemma 1) has infinite dimension, or likewise AA in the Alicki-Fannes inequality (Lemma 2), then the bound becomes trivial: the right hand side is infinite. This is completely natural, since the entropy is not even continuous, and these Fannes-type bounds imply a sort of uniform continuity. Continuity is restored, however, when restricting to states of finite energy, for instance of a quantum harmonic oscillator Wehrl , see also ESP and Shirokov:S-continuity for more recent results and excellent surveys on the status of continuity of the entropy. Shirokov Shirokov:I-continuity has developed an approach do prove (local) continuity of entropic quantities, based on certain finite entropy assumptions, in which he uses Alicki-Fannes inequalities on finite approximations.

Uniform bounds are still out of the question, but what we shall show here is that the Fannes and Alicki-Fannes inequalities discussed above have satisfying analogues, with a dependence on the energy of the states rather than the Hilbert space dimension.

Abstractly, our setting is this: Consider a Hamiltonian HH on a infinite dimensional separable Hilbert space AA. If there is another system BB and we consider bipartite states and conditional entropy, we implicitly assume trivial Hamiltonian on BB, i.e. global Hamiltonian H=HA⊗\openoneBH=H^{A}\otimes{\openone}^{B}. We shall need a number of assumptions on HH, to start with that it has discrete spectrum and that it is bounded from below; for normalization purposes we fix the ground state energy of HH to be . The mathematically precise assumption is the following.

Gibbs Hypothesis. For every β>0\beta>0, let the partition function Z(β)\mathchar58=Tr⁡e−βH<∞Z(\beta)\mathrel{\mathop{\mathchar 58\relax}}=\operatorname{Tr}e^{-\beta H}<\infty be finite, so that 1Z(β)e−βH\frac{1}{Z(\beta)}e^{-\beta H} is a bona fide state with finite entropy. In this case, for every energy EE in the spectrum of HH, the (unique) maximizer of the entropy S(ρ)S(\rho) subject to Tr⁡ρH≤E\operatorname{Tr}\rho H\leq E is of the Gibbs form:

where β=β(E)\beta=\beta(E) is decreasing with EE and is the solution to the equation

This implies that the spectrum is unbounded above, and that the energy levels cannot become “too dense” with growing energy value.

Let us immediately draw some conclusions from these assumptions; the following is a simply consequence of Shirokov’s (Shirokov:Entropy, , Prop. 1), for which we present an elementary proof.

For a Hamiltonian HH satisfying the Gibbs Hypothesis, S\bigl{(}\gamma(E)\bigr{)} is a strictly increasing, strictly concave function of the energy EE.

It is clear from the maximum entropy characterization of γ(E)\gamma(E) that the entropy as a function of EE must be non-decreasing; it is unbounded by looking at the formula for the entropy in terms of log⁡Z\log Z.

Furthermore, for energies E1E_{1} and E2E_{2}, and 0≤p≤10\leq p\leq 1,

From this it follows that S\bigl{(}\gamma(E)\bigr{)} is strictly increasing, because otherwise S\bigl{(}\gamma(E_{1})\bigr{)}=S\bigl{(}\gamma(E_{2})\bigr{)} for some E1<E2E_{1}<E_{2}, but then S\bigl{(}\gamma(E_{2})\bigr{)}<S\bigl{(}\gamma(E_{3})\bigr{)} for some E2<E3E_{2}<E_{3}, since the entropy grows to infinity as E→∞E\rightarrow\infty, contradicting concavity.

But this means that for E1≠E2E_{1}\neq E_{2}, necessarily γ(E1)≠γ(E2)\gamma(E_{1})\neq\gamma(E_{2}), and so by the strict concavity of the von Neumann entropy, we have strict inequality in the second line of Eq. (8) for 0<p<10<p<1. ∎

If HH satisfies the Gibbs Hypothesis, then for any δ>0\delta>0,

The right hand side is clearly attained by letting λ=δ\lambda=\delta. To prove “≤\leq” for any admissible λ\lambda, observe that by concavity (Proposition 11),

Letting t=λδ≤1t=\frac{\lambda}{\delta}\leq 1 and F=EλF=\frac{E}{\lambda} concludes the proof. ∎

Another useful fact proved by Shirokov (Shirokov:Entropy, , Prop. 1(ii)), which we shall invoke later, is that under our assumptions, S\bigl{(}\gamma(E)\bigr{)}=o(E), which can be recast as saying that \delta\,S\bigl{(}\gamma(E/\delta)\bigr{)}\rightarrow 0 for every finite EE and δ→0\delta\rightarrow 0. ∎

We start with an easy-to-prove continuity bound for the entropy, inspired by the proof of Lemma 1, though for the conditional entropy we shall have to resort to a different argument. It uses a quantum coupling as in Proposition 5 (which implies a weaker bound in the following, with the square of the expression on the right hand side).

Let ρ\rho and σ\sigma be states on the same Hilbert space AA, and consider the tensor square A⊗A=\mathchar58A1A2A\otimes A=\mathrel{\mathop{\mathchar 58\relax}}A_{1}A_{2} of the quantum system. Then, there exists a state ω\omega with ωA1=ρ\omega^{A_{1}}=\rho, ωA2=σ\omega^{A_{2}}=\sigma and such that

(This is known as Mirksy’s inequality (HornJohnson, , Cor. 7.4.9.3).)

in A1A2A_{1}A_{2}, we clearly have Tr⁡∣ϕ⟩ ⁣⟨ϕ∣=1−ϵ\operatorname{Tr}|\phi\rangle\!\langle\phi|=1-\epsilon, and ϕA1≤ρ\phi^{A_{1}}\leq\rho, ϕA2≤σ\phi^{A_{2}}\leq\sigma, thus we can write

with bona fide states Δ1\Delta_{1} and Δ2\Delta_{2}.

It is straightforward to check that the definition ω\mathchar58=∣ϕ⟩ ⁣⟨ϕ∣+ϵΔ1⊗Δ2\omega\mathrel{\mathop{\mathchar 58\relax}}=|\phi\rangle\!\langle\phi|+\epsilon\Delta_{1}\otimes\Delta_{2} satisfies all requirements on ω\omega. ∎

Let the Hamiltonian HH on AA satisfying the Gibbs Hypothesis. Then for any two states ρ\rho and σ\sigma on AA with Tr⁡ρH, Tr⁡σH≤E\operatorname{Tr}\rho H,\,\operatorname{Tr}\sigma H\leq E and 12∥ρ−σ∥1≤ϵ≤1\frac{1}{2}\|\rho-\sigma\|_{1}\leq\epsilon\leq 1,

Pick a state ω\omega on A1A2A_{1}A_{2}, according to Proposition 14: ωA1=ρ\omega^{A_{1}}=\rho, ωA2=σ\omega^{A_{2}}=\sigma, and with largest eigenvalue ≥1−ϵ\geq 1-\epsilon, meaning that we can write

with a pure state ∣ψ⟩|\psi\rangle (the normalized vector ∣ϕ⟩|\phi\rangle from the proof of Proposition 14) and some other state ω′\omega^{\prime}. Hence,

Here, we have first used the marginals of ω\omega, then in the second line the Araki-Lieb “triangle” inequality ArakiLieb , in the third line strong subadditivity, and in the last step the maximum entropy principle, noting that with respect to the Hamiltonian HA1⊗\openoneA2+\openoneA1⊗HA2H^{A_{1}}\otimes{\openone}^{A_{2}}+{\openone}^{A_{1}}\otimes H^{A_{2}}, ω\omega has energy ≤2E\leq 2E, and so the energy of ω′\omega^{\prime} is bounded by 2E/ϵ2E/\epsilon. For the last line, observe that the Gibbs state at energy 2E/ϵ2E/\epsilon of the composite system is γ(E/ϵ)⊗2\gamma(E/\epsilon)^{\otimes 2}. ∎

The following two general bounds lack perhaps the simple elegance of Lemma 15, but they turn out to be more flexible, and stronger in certain regimes.

For a Hamiltonian HH on AA satisfying the Gibbs Hypothesis and any two states ρ\rho and σ\sigma with Tr⁡ρH, Tr⁡σH≤E\operatorname{Tr}\rho H,\,\operatorname{Tr}\sigma H\leq E, 12∥ρ−σ∥1≤ϵ<ϵ′≤1\frac{1}{2}\|\rho-\sigma\|_{1}\leq\epsilon<\epsilon^{\prime}\leq 1, and δ=ϵ′−ϵ1+ϵ′\delta=\frac{\epsilon^{\prime}-\epsilon}{1+\epsilon^{\prime}},

For states ρ\rho and σ\sigma on the bipartite system A⊗BA\otimes B and otherwise the same assumption as before,

To interpret these bounds, we remark that in a certain sense they show that the Gibbs entropy at the cutoff energy E/ϵE/\epsilon (E/δE/\delta) takes on the role of the logarithm of the dimension in the finite dimensional case. Before we launch into their proof, let us introduce some notation: Define the energy cutoff projectors

where ∣n⟩|n\rangle is the eigenvector of eigenvalue EnE_{n} of the Hamiltonian HH. We shall also consider the pinching map

which is a unital channel, as well as its action on the original ρ\rho and σ\sigma:

Note that because HH commutes with the action of T\mathcal{T}, we have Tr⁡ξH=Tr⁡T(ξ)H\operatorname{Tr}\xi H=\operatorname{Tr}\mathcal{T}(\xi)H, and so the energy bound EE applies also to T(ρ)\mathcal{T}(\rho) and T(σ)\mathcal{T}(\sigma). Hence,

Our strategy will be to relate S(ρ)S(\rho) to S(ρ≤)S(\rho_{\leq}) (and the same for σ\sigma and σ≤\sigma_{\leq}) via entropy inequalities, including concavity, similar to the first part of the paper, and then apply the usual Fannes (Alicki-Fannes) inequalities to ρ≤\rho_{\leq} and σ≤\sigma_{\leq}.

Proof of Lemma 16. First of all, by concavity of the entropy (monotonicity under unital cptp maps),

Now, by Eq. (9), the maximum entropy principle and Corollary 12,

Thus, from Eq. (10), observing δ≤12\delta\leq\frac{1}{2}, we get

To see this, we think of the action of T\mathcal{T} as a binary measurement on the system AA, which we can implement coherently with two ancilla qubits XX and X′X^{\prime},

Applying this to σ\sigma, we have by unitary invariance and the Araki-Lieb “triangle” inequality,

Thus, using that the energy of σ≤\sigma_{\leq} is at most E/δE/\delta by construction, and so S(\sigma_{\leq})\leq S\bigl{(}\gamma(E/\delta)\bigr{)},

Third, by definitions, contractivity of the trace norm and triangle inequality,

Hence by the Fannes inequality in the form of Lemma 1,

The latter inequality holds because the state 1Tr⁡P≤P≤\frac{1}{\operatorname{Tr}P_{\leq}}P_{\leq} clearly has energy bounded by E/δE/\delta, and so cannot have entropy larger than the Gibbs state.

With these three elements we can conclude the proof: W.l.o.g. S(ρ)≥S(σ)S(\rho)\geq S(\sigma), and so from Eqs. (11), (13) and (15),

Proof of Lemma 17. It is very similar to the previous one, only that we have to be a bit more careful in some details, as the conditional entropy can be negative.

The first step goes through almost unchanged, with the map T⊗id⁡B\mathcal{T}\otimes{\operatorname{id}}_{B}, since the conditional entropy is concave as well (equivalent to strong subadditivity) SSA :

The remainder term λS(A∣B)ρ>\lambda S(A|B)_{\rho_{>}} is upper bounded by λS(ρ>A)\lambda S(\rho_{>}^{A}) (again by strong subadditivity), hence the upper bound \lambda S\bigl{(}\gamma(E/\lambda)\bigr{)} still applies. The only change is due to the fact that the conditional entropy can be negative. However, for any bipartite state ξAB\xi^{AB},

Here, the right hand inequality is strong subadditivity that we have used before; introducing a purification ∣φ⟩ABC|\varphi\rangle^{ABC} of the state, we have −S(A∣B)φ=S(A∣C)φ≤S(ξA)-S(A|B)_{\varphi}=S(A|C)_{\varphi}\leq S(\xi^{A}), which is the left hand inequality. Thus,

Also the second step requires only minor modifications: With the notation of the previous proof, and using the Araki-Lieb “triangle” inequality once again,

Again, since conditional entropies can be negative, we have to be more careful with remainder terms and get

In the third step, the trace norm estimate (14) goes through unchanged, and then we apply the Alicki-Fannes inequality in the form of Lemma 2:

Putting this together with Eqs. (16) and (17), assuming w.l.o.g. that S(A∣B)ρ≥S(A∣B)σS(A|B)_{\rho}\geq S(A|B)_{\sigma}, we obtain

The bounds of Lemmas 15, 16 and 17 are very general, and it may not be immediately apparent how useful they are. However, thanks to (Shirokov:Entropy, , Prop. 1(ii)), restated in Remark 13, \delta S\bigl{(}\gamma(E/\delta)\bigr{)}\rightarrow 0 for every finite EE, as δ→0\delta\rightarrow 0 (cf. (Shirokov:squashed, , Cor. 4)). Thus, choosing ϵ′=ϵ\epsilon^{\prime}=\sqrt{\epsilon}, the lemmas do prove continuity of the entropy and conditional entropy in general, and uniformly for each fixed energy.

where ωi\omega_{i} is the native frequency of the ii-th oscillator and aia_{i} is its annihilation (aka lowering) operator (see e.g. KokLovett or Weedbrook-rev ). Note that we chose the slightly unusual energy convention such that the ground state has energy , rather than ∑i12ℏωi\sum_{i}\frac{1}{2}\hbar\omega_{i}, to be able to apply directly our above results. In the case of a single mode, and choosing units such that ℏω1=1\hbar\omega_{1}=1, the Hamiltonian simply becomes the number operator NN. In that case, it is well-known that

Crucially, and in accordance with Proposition 11, gg is a concave, monotone increasing function of NN.

By using this upper bound in Lemmas 16 and 17, for δ=αϵ(1−ϵ)\delta=\alpha\epsilon(1-\epsilon), with a parameter α\alpha between and 12\frac{1}{2}, and introducing

For each fixed ϵ≤1\epsilon\leq 1, we can make α\alpha arbitrarily small, and then for large energy E≫∑iℏωiE\gg\sum_{i}\hbar\omega_{i}, the bounds of Lemma 18 are asymptotically tight, in the sense that apart from the additive offset terms, the factor multiplying ϵ\epsilon (2ϵ2\epsilon, resp.) cannot be smaller than

V Conclusions

Using entropy inequalities, specifically concavity, we improved the appearance of the Alicki-Fannes inequality for the conditional von Neumann entropy to an almost tight form. It would be nice to know the ultimately best form among all formulas that depend only on the dimension of the Hilbert space and the trace distance, but we have to leave this as an open problem.

In particular, it would be curious to find the optimal form of the fidelity in Proposition 5,

with a fixed purification ψ\psi of σ\sigma, and of Proposition 14,

which may be regarded as quantum state analogues of the coupling random variables,

Furthermore, are there versions of these statements that would allow for alternative proofs or tighter versions of Lemmas 2 and 17 for the conditional entropy?

The same principle lead to the apparently first uniform continuity bounds of the entropy and conditional on infinite dimensional Hilbert spaces under a bound on the expected energy (or, for that matter, bounded expectation of any sufficiently well-behaved Hermitian operator). In the case of a system of harmonic oscillators, we have seen that the bound is, in a certain sense, asymptotically tight, even though here we are much farther away from a universally optimal form.

The Fannes and Alicki-Fannes inequalities already are known to have many applications in quantum information theory. These include the continuity of certain entanglement measures such as entanglement of formation Nielsen-continuity , relative entropy of entanglement DonaldHorodecki , squashed entanglement E-sq and conditional entanglement of mutual information YHW , and of various quantum channel capacities LeungSmith . In fact, we always get explicit continuity bounds in terms of the trace distance of the states or diamond norm distance of the channels, respectively. While in many applications it is of minor interest to have the optimal form of the bound (for example when ϵ\epsilon goes to ), it pays off to have a tighter bound than AlickiFannes in the setting of approximately degradable channels Sutter-et-al . Indeed, this results even in new, tighter upper bounds on the quantum capacity of very quiet depolarizing channels Sutter-et-al , by way of an extension of the methodology of Q_ss .

The infinite dimensional versions of these entropy bounds under an energy constraint are awaiting applications, though it seems clear that explicit bounds on the continuity and asymptotic continuity of entanglement measures ESP , (e.g. for squashed entanglement since the first posting of the present manuscript Shirokov:squashed ) and channel capacities Holevo:constrained ; Holevo-e-constrained ; HolevoShirokov:constrained ; ShirokovHolevo in infinite dimension should be among the first, as well as the extension of approximate degradability Sutter-et-al to Bosonic channels in-prep .

. Acknowledgments. Thanks to David Sutter and Volkher Scholz for stimulating discussions, to Nihat Ay, Milán Mosonyi and Dong Yang for remarks on general relative entropy distances, to Maxim Shirokov for his many insights into entropy and entanglement measures, in particular his keen interest in the asymptotic continuity of entanglement cost and for spotting an error in an earlier version of the proof of Lemma 17, and to Mark Wilde for comments on the history of Lemma 1. The hospitality of the Banff International Research Station (BIRS) during the workshop “Beyond IID in Information Theory” (5-10 July 2016) is gratefully acknowledged, where Volkher Scholz and David Sutter posed the derivation of infinite dimensional Fannes type inequalities as an open problem, and where the present work was initiated.

The author’s work was supported by the EU (STREP “RAQUEL”), the ERC (AdG “IRQUAT”), the Spanish MINECO (grant FIS2013-40627-P) with the support of FEDER funds, as well as by the Generalitat de Catalunya CIRIT, project 2014-SGR-966.

References