Partially smoothed information measures


Abstract

Smooth entropies are a tool for quantifying resource trade-offs in (quantum) information theory and cryptography. In typical bi- and multi-partite problems, however, some of the sub-systems are often left unchanged and this is not reflected by the standard smoothing of information measures over a ball of close states. We propose to smooth instead only over a ball of close states which also have some of the reduced states on the relevant sub-systems fixed. This partial smoothing of information measures naturally allows to give more refined characterizations of various information-theoretic problems in the one-shot setting. In particular, we immediately get asymptotic second-order characterizations for tasks such as privacy amplification against classical side information or classical state splitting. For quantum problems like state merging the general resource trade-off is tightly characterized by partially smoothed information measures as well.

1 Introduction↩︎

One-shot information theory concerns itself with finding tight bounds on the resource trade-offs for various operational problems in information theory and cryptography (see, e.g., [1] for an introduction). Smooth entropies and smooth mutual informations have in many cases proven to be adequate information measures in this context. On the one hand, smooth min-entropy was first introduced in the context of quantum cryptography [2]. More precisely, the smooth conditional min-entropy was introduced to characterize the amount of uniform and independent randomness that can be extracted from a correlated random variable. On the other hand, the smooth max-information has been introduced to quantify the communication requirements in quantum extensions of Slepian-Wolf coding [3]. Since then smooth entropy measures of various kinds have been used to characterize a plethora of other tasks as well.

Our main contributions can be summarized as follows:

  • We introduce a notion of partially smoothed mutual max-information and conditional min-entropy and establish some of their basic mathematical properties.

  • We show that these new definitions are equivalent to their fully smoothed counterparts, up to terms that vanish in the first-order i.i.d.asymptotics. Moreover, for the fully classical case this equivalence even holds for the asymptotic second-order i.i.d.asymptotics.

  • We give several examples of operational problems where the new quantities naturally appear to give tighter bounds for the one-shot problem. In particular, for classical problems this leads to asymptotic second-order i.i.d.expansions.

The remainder of this paper is organized as follows. In 2 we introduce our new definition of smooth mutual max-information and conditional min-entropy with restricted smoothing. In 3 we show that these definitions are, up to small correction terms, equivalent to the standard definitions found in the literature. We are able to show stronger equivalences for the special case of classical distributions. Finally, 4 discusses various operational interpretations of the new measures in detail, both in the classical and quantum context.

2 Definitions and basic properties↩︎

2.1 Basic notation↩︎

For any finite-dimensional inner product space describing a quantum system — let us fix it to \(A\) for the sake of clarity —  we define the set of positive semi-definite (psd) operators acting on \(A\) as \(\mathcal{P}(A)\). We also define two subsets: the set of quantum states (i.e.psd operators with unit trace), denoted \(\mathcal{S}_{\circ}(A)\), and sub-normalized states (i.e.psd operators with trace not exceeding unity), denoted \(\mathcal{S}_{\bullet}(A)\). We describe joint quantum systems using the shorthand \(AB = A \otimes B\).

The Löwner partial order of operators in \(\mathcal{P}(A)\), denoted by \('\geq'\), is given by the relation \(A \geq B\) if and only if \(A - B\) is psd. Moreover, we say that \(A\) dominates \(B\), denoted \(A \gg B\), if and only if the support of \(B\) is contained in the support of \(A\).

2.2 Smooth entropy measures↩︎

Let \(\rho\) and \(\sigma\) be two psd operators. If \(\rho \ll \sigma\) we define the max-divergence [4], [5] as \[\begin{align} D_{\max}(\rho\|\sigma) := \inf \big\{ \lambda : \rho \leq \exp(\lambda) \sigma \big\} \,, \end{align}\] and otherwise it is defined as \(+\infty\). This quantity can be used to define various notions of mutual max-information and conditional min-entropy, respectively. We will concern ourselves with the following two definitions [2], [3]. For any bipartite state \(\rho_{AB} \in \mathcal{S}_{\bullet}(AB)\), we have \[\begin{align} I_{\max}(A ; B)_{\rho} &:= \inf_{\sigma_B \in \mathcal{S}_{\bullet}(B)} D_{\max}(\rho_{AB} \| \rho_A \otimes \sigma_B) \,,\tag{1}\\ H_{\min}(A | B)_{\rho} &:= - D_{\max}(\rho_{AB} \| 1_A \otimes \rho_B) \,.\tag{2} \end{align}\] One immediate observation is that \(\sigma_B \in \mathcal{S}_\circ(B)\) would have worked equally well since the minimizer will always have maximal trace. Moreover, many variations of these definitions can be found in the literature. Most prominently, the min-entropy can be defined using a maximization over \(\sigma_B \in \mathcal{S}_{\bullet}(B)\) similar to the max-information, which yields a quantity with clear operational interpretation [2], [6] (see also [7] about the max-information). The above choices are determined by the applications we discuss in Section 4.

Our goal is to define a smooth max-information and smooth min-entropy based on the above quantities, i.e.quantities for which \(\rho_{AB}\) is replaced with a ball of states close to \(\rho_{AB}\). In particular, we want the states in this ball to have the property that the \(A\) subsystem is (essentially) left intact. To do this we will need to use a metric on (sub-normalized) quantum states, i.e.positive semi-definite operators with trace not exceeding 1. This metric, let us denote it by \(\Delta(\cdot, \cdot)\), needs to have the following properties (we assume \(\rho, \sigma\) and \(\tau\) are sub-normalized quantum states).

  1. Positive definiteness: \(\Delta(\rho, \sigma) \geq 0\) with equality only if \(\rho = \sigma\).

  2. Triangle inequality: \(\Delta(\rho, \sigma) \leq \Delta(\rho, \tau) + \Delta(\tau, \sigma)\).

  3. Strong monotonicity: \(\Delta(\mathcal{E}(\rho), \mathcal{E}(\sigma)) \leq \Delta(\rho, \sigma)\) for any completely positive and trace non-increasing map \(\mathcal{E}\).

We in particular will consider two metrics satisfying the above properties. The purified distance [8] based on the generalized fidelity, \[\begin{align} P(\rho,\tau) := \sqrt{1 - F^2(\rho,\tau)} \quad \textrm{with} \quad F(\rho,\tau) := \mathop{\mathrm{tr}}\Big[\big|\sqrt{\rho}\sqrt{\tau} \big|\Big] + \sqrt{(1-\mathop{\mathrm{tr}}[\rho])(1-\mathop{\mathrm{tr}}[\tau])} \,, \end{align}\] and the generalized trace distance [1], \[\begin{align} T(\rho,\tau) := \frac{1}{2} \mathop{\mathrm{tr}}\Big[\big| \rho - \tau \big|\Big] + \frac{1}{2} \big| \mathop{\mathrm{tr}}[\rho] - \mathop{\mathrm{tr}}[\tau] \big| \,. \end{align}\] We will also use the abbreviations for the standard measures \[\begin{align} \bar{F}(\rho,\tau):=\mathop{\mathrm{tr}}\Big[\big|\sqrt{\rho}\sqrt{\tau} \big|\Big]\quad \textrm{and} \quad \|\rho-\tau\|_1:=\mathop{\mathrm{tr}}\Big[\big| \rho - \tau \big|\Big]\,. \end{align}\] In the following, let \(\Delta\) be a metric satisfying the above Properties 1–3. Moreover, we call a tuple \((\varepsilon, \Delta)\) with \(\varepsilon\geq 0\) valid for a state \(\rho\) if \(\Delta(\rho, 0) > \varepsilon\), where \(0\) denotes the additive identity. The following two definitions are rather standard (see, e.g., [1] for an overview):

Definition 1. Let \(\rho_{AB} \in \mathcal{S}_{\circ}(AB)\) with valid \((\varepsilon, \Delta)\). The \((\varepsilon,\Delta)\)-smooth max-information of \(A\) and \(B\) is defined as1 \[\begin{align} && I_{\max}^{\varepsilon, \Delta}( A ; B)_{\rho}\;:=\;\inf\quad & D_{\max}(\tilde{\rho}_{AB} \|\rho_A \otimes \sigma_B) &&&&\\ && \textrm{s.t.}\quad & \tilde{\rho}_{AB} \in \mathcal{S}_{\circ}(AB), \notag\\ &&& \Delta(\tilde{\rho}_{AB},\rho_{AB}) \leq \varepsilon, \notag\\ &&& \sigma_B \in \mathcal{S}_{\circ}(B) \notag\,. \end{align}\] Moreover, the \((\varepsilon,\Delta)\)-smooth conditional min-entropy of \(A\) given \(B\) is defined as \[\begin{align} && H_{\min}^{\varepsilon, \Delta}( A | B)_{\rho}\;:=\;\sup\quad & - D_{\max}(\tilde{\rho}_{AB} \| 1_A \otimes \rho_B) &&&&\\ && \textrm{s.t.}\quad & \tilde{\rho}_{AB} \in \mathcal{S}_{\bullet}(AB), \notag\\ &&& \Delta(\tilde{\rho}_{AB}, \rho_{AB}) \leq \varepsilon\,. \end{align}\]

As for the non-smooth case there are other definitions in use that we will not discuss here specifically, e.g.in the definition of the smooth max-information one can fix \(\sigma_B\) to be \(\rho_B\) to arrive at a different quantity, and similarly the min-entropy can be further optimized over \(\sigma_B \in \mathcal{S}_{\bullet}(B)\). Note that for the smooth min-entropy it is necessary to smooth over sub-normalized states as otherwise the quantity will not be invariant under the application of local embedding maps (see also [1]).

Given this, we propose the following definition for the smooth max-information.2

Definition 2. Let \(\rho_{AB} \in \mathcal{S}_{\circ}(AB)\) with valid \((\varepsilon, \Delta)\). The \((\varepsilon,\Delta)\)-smooth max-information with fixed \(A\) of \(\rho_{AB}\)* is defined as \[\begin{align} && I_{\max}^{\varepsilon, \Delta}( \dot{A} ; B)_{\rho}\;:=\;\inf\quad & D_{\max}(\tilde{\rho}_{AB} \| \rho_A \otimes \sigma_B) &&&&\\ && \textrm{s.t.}\quad & \tilde{\rho}_{AB} \in \mathcal{S}_{\circ}(AB), \notag\\ &&& \Delta(\tilde{\rho}_{AB},\rho_{AB}) \leq \varepsilon, \notag\\ &&& \tilde{\rho}_A = \rho_A, \notag\\ &&& \sigma_B \in \mathcal{S}_{\circ}(B) \notag\,. \end{align}\]*

For the purpose of this paper we also suggest the following definition of smooth conditional min-entropy.

Definition 3. Let \(\rho_{AB} \in \mathcal{S}_{\circ}(AB)\) with valid \((\varepsilon, \Delta)\). Then, the \((\varepsilon,\Delta)\)-smooth min-entropy with fixed \(B\) of \(\rho_{AB}\)* is defined as \[\begin{align} && H_{\min}^{\varepsilon, \Delta}(A | \dot{B})_{\rho}\;:=\;\sup\quad & -D_{\max}(\tilde{\rho}_{AB} \| 1_A \otimes \rho_B) &&&&\\ && \textrm{s.t.}\quad & \tilde{\rho}_{AB} \in \mathcal{S}_{\bullet}(AB), \notag\\ &&& \Delta(\tilde{\rho}_{AB},\rho_{AB}) \leq \varepsilon, \notag\\ &&& \tilde{\rho}_B \leq \rho_B, \notag\,. \end{align}\]*

Note that smooth versions of all conditional Rényi entropies (see, e.g., [10]) can be defined analogously. However, we will not explore these definitions further here.

If the input states are classical in a fixed basis all the definitions apply for this case as well. It is then immediate to see that the respective optimizations over \(\tilde{\rho}_{AB}\) and \(\sigma_B\) can without loss of generality be restricted to be diagonal in this fixed basis as well.3

2.3 Basic properties↩︎

We will now discuss some basic properties of the quantities introduce above. In the following lemmas we assume that \(\Delta\) satisfies Properties 1–3. Let us explore the above two definitions. The first property is an immediate consequence of the positive definiteness of \(\Delta\).

Lemma 1. Let \(\rho_{AB} \in \mathcal{S}_{\circ}(AB)\). Then, we have \[\begin{align} I_{\max}^{0, \Delta}(\dot{A} ; B)_{\rho} = I_{\max}(A ; B)_{\rho} \qquad \textrm{and} \qquad H_{\min}^{0, \Delta}(A|\dot{B})_{\rho} = H_{\min}(A|B)_{\rho}\,. \end{align}\]

The second lemma, on the other hand, is a direct consequence of the triangle inequality of \(\Delta\).

Lemma 2. Let \(\rho_{AB}, \tilde{\rho}_{AB} \in \mathcal{S}_{\circ}(AB)\) with \(\Delta(\rho_{AB}, \tilde{\rho}_{AB}) \leq \eta\), and \((\varepsilon\!+\!\eta, \Delta)\) valid for \(\rho_{AB}\). Then, we have \[\begin{align} I_{\max}^{\varepsilon+\eta, \Delta}(\dot{A} ; B)_{\rho} \leq I_{\max}^{\varepsilon, \Delta}(\dot{A} ; B)_{\tilde{\rho}} \qquad \textrm{and} \qquad H_{\min}^{\varepsilon+\eta, \Delta}(A ; \dot{B})_{\rho} \geq H_{\min}^{\varepsilon, \Delta}(A ; \dot{B})_{\tilde{\rho}} \,. \end{align}\]

We continue with the following observation that follows immediately from well-known properties of the max-divergence, namely that the smooth max-information is non-increasing under local completely positive trace-preserving (cptp) operations and the smooth min-entropy is non-decreasing under local operations.

Lemma 3. Let \(\rho_{AB} \in \mathcal{S}_{\circ}(AB)\) with valid \((\varepsilon, \Delta)\). For any two completely positive trace preserving maps \(\mathcal{E}: \mathcal{P}(A) \to \mathcal{P}(A')\) and \(\mathcal{F}: \mathcal{P}(B) \to \mathcal{P}(B')\), we have \[\begin{align} I_{\max}^{\varepsilon}( \dot{A} ; B)_{\rho} \geq I_{\max}^{\varepsilon}( \dot{A}' ; B')_{\tau} \end{align}\] where \(\tau_{A'B'} = (\mathcal{E} \otimes \mathcal{F})(\rho_{AB})\). Furthermore, if \(\mathcal{E}\) is also sub-unital (i.e.it satisfies \(\mathcal{E}(1_A) \leq 1_{A'}\)), then \[\begin{align} H_{\min}^{\varepsilon}(A ; \dot{B})_{\rho} \leq H_{\min}^{\varepsilon}(A' ; \dot{B}')_{\tau} \,. \end{align}\]

Clearly every operational definition should be invariant under isometries as embeddings are essentially just a choice of modeling and should not effect operational quantities.

Lemma 4. Let \(\rho_{AB} \in \mathcal{S}_{\circ}(AB)\) with valid \((\varepsilon, \Delta)\). For any two isometries \(U: A \to A'\), \(V: B \to B'\), it holds that \[\begin{align} I_{\max}^{\varepsilon,\Delta}(\dot{A} ; B)_{\rho} = I_{\max}^{\varepsilon,\Delta}(\dot{A}' ; B')_{\rho} \qquad \textrm{and} \qquad H_{\min}^{\varepsilon,\Delta}(A ; \dot{B})_{\rho} = H_{\min}^{\varepsilon,\Delta}(A' ; \dot{B}')_{\rho} \,, \end{align}\] where \(\rho_{A'B'} = (U \otimes V) \rho_{AB} (U \otimes V)^{\dagger}\).

Proof. The result for the smooth max-information follows immediately from Lem. 3. To see this, note that the map \((U \otimes V) \cdot (U \otimes V)^{\dagger}\) is cptp and we can define cptp inverse maps of the form \[\begin{align} \rho_A \mapsto U^{\dagger} \rho_A U + \big(1 - \mathop{\mathrm{tr}}U^{\dagger} \rho_A U \big) \tau_A, \qquad \rho_B \mapsto V^{\dagger} \rho_B V + \big(1 - \mathop{\mathrm{tr}}V^{\dagger} \rho_B V \big) \tau_B, \label{eq:inverseiso} \end{align}\tag{3}\] where \(\tau_{A} \in \mathcal{S}_{\circ}(A)\) and \(\tau_{B} \in \mathcal{S}_{\circ}(B)\) are arbitrary states.

We now present the proof for the smooth min-entropy by proving inequalities in both directions. The direction ‘\(\geq\)’ is guaranteed by Lem. 3 with the same argument as above. However, the map in 3 is not sub-unital so we cannot employ the data-processing inequality to show ‘\(\leq\)’.

Instead, consider the states \(\tilde{\rho}_{A'B'}\) and \(\sigma_{B'}\) that are optimal for \(H_{\min}^{\varepsilon}(A' | \dot{B}')_{\rho}\). We define \[\begin{align} \tilde{\rho}_{AB} := (U \otimes V)^{\dagger} \tilde{\rho}_{A'B'} (U \otimes V) \,. \end{align}\] And note that \(\rho_B := V^{\dagger} \rho_{B'} V\). Note that the maps \(U^{\dagger} (\cdot) U\) and \(V^{\dagger} (\cdot) V\) are in general not trace-preserving as weight outside the range of \(U\) and \(V\) is discarded when we invert the isometries. However, the resulting state \(\tilde{\rho}_{AB}\) and \(\rho_B\) are feasible for the optimization in \(H_{\min}^{\varepsilon,\Delta}(A | \dot{B})_{\rho}\) since the following holds:

  • We have \(\tilde{\rho}_{AB} \in \mathcal{S}_{\bullet}(AB)\) and \(\rho_B \in \mathcal{S}_{\bullet}(B)\);

  • It holds that \(\Delta(\tilde{\rho}_{AB}, \rho_{AB}) \leq \Delta(\tilde{\rho}_{A'B'}, \rho_{A'B'}) \leq \varepsilon\) due to the fact that the metric is monotone under trace non-increasing completely positive maps;

  • We have \(\tilde{\rho}_A = U^{\dagger} \mathop{\mathrm{tr}}_B \left( (1_A \otimes V^{\dagger}) \tilde{\rho}_{A'B'} (1_A \otimes V) \right) U \leq U^{\dagger} \tilde{\rho}_{A'} U \leq U^{\dagger} \rho_{A'} U = \rho_A\), where the first inequality is due to Lem. 10 in App. 7.

Finally, the desired inequality follows from the implication: \[\begin{align} \tilde{\rho}_{A'B'} \leq \exp(\lambda) 1_{A}' \otimes \sigma_{B'} \quad \implies \quad \tilde{\rho}_{AB} \leq \exp(\lambda) U^{\dagger} 1_{A'} U \otimes \sigma_B \,, \end{align}\] that we yield from applying the completely positive map \((U \otimes V)^{\dagger} \cdot (U \otimes V)\) on both sides of the operator inequality. Finally, note that \(U^{\dagger} 1_{A'} U = 1_A\). ◻

One could hope to replace \(\tilde{\rho}_A \leq \rho_A\) in Def. 2 by an equality, thus forcing the state \(\tilde{\rho}_{AB}\) to have the same trace as \(\rho_{AB}\). However, for such a definition one would then need to show a property analogous to the above invariance under isometries, which seems non-trivial. The following argument gives also an indication that sub-normalized states are desirable in this context, although it does not conclusively show that they are necessary for our definition.

For the (unconditional) min-entropy, invariance under isometries can only hold if we allow sub-normalized states. To see this, consider the min-entropy of the state \(\rho = 1/ d\), which is maximal for normalized states of dimension \(d\) and thus cannot be increased by smoothing over this set. However, if embedded into a larger space smoothing will yield a larger min-entropy. Allowing sub-normalized states introduces an alternative to moving weight out of the support of \(\rho\) and it turns out that this is exactly what is needed to ensure the quantity is invariant under isometries.

3 Relation to other Entropy Measures↩︎

3.1 Classical Setting↩︎

Since the (generalized) trace distance is directly connected to error probabilities it is often natural to stick to this distance measure for classical problems. We will do so in this section. We will also continue using the notations \(\mathcal{P}, \mathcal{S}_{\circ}, \mathcal{S}_{\bullet}\), although now we restrict to diagonal matrices in some basis, interpreted as (potentially sub-normalized) probability distributions. In order to establish an asymptotic equipartition property for our locally smoothed information measures we relate them to other well-studied entropic quantities such as information spectrum divergences [11]. Note that standard asymptotic equipartition proofs for mutual information and conditional entropy do not leave any of the marginals unchanged.

Definition 4. For \(P_X,Q_X\in\mathcal{P}(X)\) and \(\varepsilon\in[0,1]\), the max-information spectrum divergence is defined as \[\begin{align} D_s^{\varepsilon}(P_X\|Q_X): = \inf\left\{a: \Pr_{x \leftarrow p_X}\left\{\frac{P_X(x)}{Q_{X}(x)}> 2^a\right\}<\varepsilon\right\}\,. \end{align}\]

Importantly, the max-information spectrum divergence has the following asymptotic second-order expansion in for independent and identically distributed (i.i.d.) case [12] \[\begin{align} \label{eq:spectrum-divergence95second-order} \frac{1}{n}D_s^{\varepsilon}(P_X^{\times n}\|Q_X^{\times n})=D(P_X\|Q_X)+\sqrt{\frac{V(P_X\|Q_X)}{n}}\cdot\Phi^{-1}(\varepsilon)+O\left(\frac{\log n}{n}\right) \end{align}\tag{4}\] with the Kullback-Leibler divergence \(D(P_X\|Q_X):=\sum_xP_X(x)\log\left(\frac{P_X(x)}{Q_X(x)}\right)\), the information variance \(V(P_X\|Q_X):=\mathbb{E}\left[(\log P_X-\log Q_X-D(P_X\|Q_X))^2\right]\), and the cumulative standard Gaussian distribution \[\begin{align} \Phi(x):=\int_{-\infty}^x\frac{1}{2\pi}\exp(x^2/2)\,\mathrm{d}x\,. \end{align}\] We then define the information spectrum max-information and conditional min-entropy as \[\begin{align} I_s^{\varepsilon}(X;Y)_P:= D_s^{\varepsilon}(P_{XY} \| P_X\times P_Y) \qquad \textrm{and} \qquad H_s^{\varepsilon}(X|Y)_P:= -D_s^{\varepsilon}(P_{XY} \| 1_X\times P_Y)\,, \end{align}\] respectively. This leads to the following equivalence result.

Theorem 1. Let \(P_{XY}\in\mathcal{S}_{\circ}(XY)\) and \(0<\varepsilon+\delta\leq1\). Then, we have \[\begin{align} \label{eq:IsImaxHsHmin} &I_s^{\frac{\varepsilon}{1-\delta}+\delta}(X;Y)_P - 2\log\frac{1}{\delta}\leq I_{\max}^{\varepsilon, T}( \dot{X} ; Y)_P\leq I_s^{\varepsilon}(X;Y)_P+1\\ &H_s^{\frac{\varepsilon}{1-\delta}}(X|Y)_P + \log\frac{1}{\delta}\geq H_{\min}^{\varepsilon, T}(X| \dot{Y})_P\geq H_s^{\varepsilon}(X|Y)_P - 1\,. \end{align}\qquad{(1)}\] In particular, this implies the asymptotic second-order expansions \[\begin{align} \frac{1}{n}I_{\max}^{\varepsilon,T}(\dot{X};Y)_P&=I(X;Y)_P+\sqrt{\frac{V(X;Y)_P}{n}}\cdot\Phi^{-1}(\varepsilon)+O\left(\frac{\log n}{n}\right)\\ \frac{1}{n}H_{\min}^{\varepsilon, T}(X| \dot{Y})_P&=H(X|Y)_P+\sqrt{\frac{V(X|Y)_P}{n}}\cdot\Phi^{-1}(\varepsilon)+O\left(\frac{\log n}{n}\right) \end{align}\] with the mutual information variance \(V(X;Y)_P:=V(P_{XY}\|P_X\times P_Y)\) and the conditional information variance \(V(X|Y)_P:=V(P_{XY}\|1_x\times P_Y)\).

For the proof of Thm. 1 we first need to introduce some additional quantities and lemmas. Recall that for the classical special case we have \[\begin{align} I_{\max}^{\varepsilon, T}(\dot{X}; Y)_{P} = \inf_{P'_{XY}\in \mathcal{S}_{\circ}(XY), Q_Y\in \mathcal{S}_{\circ}(Y): P'_X = P_X, T(P'_{XY}, P_{XY})\leq \varepsilon} D_{\max}(P'_{XY}\| P_X\times Q_Y)\,. \end{align}\] Now, with \(Q_Y\in\mathcal{S}_{\circ}(Y)\) we define the following intermediate quantities \[\begin{align} I_s^{\varepsilon}(X;Y)_{P|Q}&:= D_s^{\varepsilon}(P_{XY} \| P_X\times Q_Y)\\ I_{\max}^{\varepsilon, T}(\dot{X}; Y)_{P|Q}&:= \inf_{P'_{XY} \in \mathcal{S}_{\circ}(XY): P'_X = P_X, T(P'_{XY}, P_{XY})\leq \varepsilon} D_{\max}(P'_{XY}\| P_X\times Q_Y)\,, \end{align}\] where their utility is that they roughly capture both the smooth max-information and the smooth min-entropy (as we will see). To continue, note that \[\begin{align} \label{eq:reducetoImax} &I_s^{\varepsilon}(X;Y)_P = I_s^{\varepsilon}(X;Y)_{P|P}\\ &H_s^{\varepsilon}(X|Y)_P = \log|X| - I_s^{\varepsilon}(Y; X)_{P|U}\quad\text{with U_X the uniform distribution}\\ &I_{\max}^{\varepsilon, T}(\dot{X}; Y)_{P} = \inf_{Q_Y}I_{\max}^{\varepsilon, T}(\dot{X}; Y)_{P|Q}\\ &H_{\min}^{\varepsilon, T}(X| \dot{Y})_P \geq \log|X| - I_{\max}^{\varepsilon, T}(\dot{Y}; X)_{P|U}\quad\text{with U_X the uniform distribution.} \end{align}\tag{5}\] Here, we observe that \(H_{\min}^{\varepsilon,T}(X| \dot{Y})_P\) is only lower bounded since it involves a supremum over sub-normalized distributions. The last ingredient is the following lemma, which says that \(I_s^{\varepsilon}(X;Y)_{P|Q}\) cannot be too small in comparison to \(I_s^{\varepsilon}(X;Y)_{P}\).

Lemma 5. Let \(P_{XY}\in \mathcal{S}_{\circ}(XY)\), \(Q_Y\in \mathcal{S}_{\circ}(Y)\), and \(0 < \varepsilon+\delta <1\). Then, we have \[\begin{align} \label{eq:I95sQsmall} I_s^{\varepsilon}(X;Y)_{P|Q} \geq I_s^{\varepsilon+ \delta}(X;Y)_{P} - \log\frac{1}{\delta}\,. \end{align}\qquad{(2)}\]

Proof. Let \(c:= I_s^{\varepsilon}(X;Y)_{P|Q}\), \(\mathrm{Bad}_1\) be the set of all \((x,y)\) for which \(P_{XY}(x,y) \geq 2^cP_X(x)Q_Y(y)\), and \(\mathrm{Bad}_2\) be the set of all \((x,y)\) for which \(Q_Y(y) > \frac{1}{\delta}P_Y(y)\). Now, observe that \[\begin{align} \Pr_{(x,y) \leftarrow P_{XY}}\left\{\mathrm{Bad}_2\right\} =\Pr_{y \leftarrow P_Y}\left\{Q_Y(y) > \frac{1}{\delta}P_Y(y)\right\} = \sum_{y:Q_Y(y) > \frac{1}{\delta}P_Y(y)}P_Y(y) \leq \delta\sum_yQ_Y(y) \leq \delta\,. \end{align}\] For all \((x,y) \notin \mathrm{Bad}_1\cup \mathrm{Bad}_2\), we have \[\begin{align} P_{XY}(x,y) \leq \frac{2^c}{\delta}P_X(x)P_Y(y)\,. \end{align}\] Furthermore, we have \[\begin{align} \Pr_{(x,y)\leftarrow P_{XY}}\left\{\mathrm{Bad}_1\cup \mathrm{Bad}_1\right\} \leq \Pr_{(x,y)\leftarrow P_{XY}}\left\{\mathrm{Bad}_1 \right\} + \Pr_{(x,y) \leftarrow P_{XY}}\left\{\mathrm{Bad}_2\right\} \leq \varepsilon+\delta\,, \end{align}\] which proves the claim. ◻

The proof of the equivalence result in Thm. 1 is then as follows.

Proof of Thm. 1. We first show that for every \(Q_Y\in \mathcal{S}_{\circ}(Y)\), \[\begin{align} \label{eq:arbitQimax} I_s^{\frac{\varepsilon}{1-\delta}}(X;Y)_{P|Q} - \log\frac{1}{\delta}\leq I_{\max}^{\varepsilon, T}( \dot{X} ; Y)_{P|Q}\leq I_s^{\varepsilon}(X;Y)_{P|Q}+1\,. \end{align}\tag{6}\]

  • For the lhs the argument is similar to that given in [13]. Let \(d:= I_{\max}^{\varepsilon, T}(\dot{X};Y)_{P|Q}\) and \(P'_{XY}\) be the probability distribution achieving the infimum in the definition. Define the set \[\begin{align} \mathcal{A}:=\left\{(x,y): \frac{P_{XY}(x,y)}{P'_{XY}(x,y)} \geq \frac{1}{\delta}\right\}\,. \end{align}\] Using Lem. 11, we find \[\begin{align} \varepsilon\geq T(P_{XY}, P'_{XY}) \geq P_{XY}(\mathcal{A}) - P'_{XY}(\mathcal{A}) \geq P_{XY}(\mathcal{A}) - \delta P_{XY}(\mathcal{A})= (1-\delta)P_{XY}(\mathcal{A})\,. \end{align}\] Thus, we get that \(P_{XY}(\mathcal{A}) \leq \frac{\varepsilon}{1-\delta}\). Moreover, for every \((x,y)\in \mathcal{A}^{\mathrm{c}}\), we have \[\begin{align} P_{XY}(x,y)\leq \frac{1}{\delta}P'_{XY}(x,y) \leq \frac{2^d}{\delta} P_X(x)Q_Y(y)\,. \end{align}\] Hence, we conclude that \(I_s^{\frac{\varepsilon}{1-\delta}}(X:Y)_{P|Q} \leq d + \log\frac{1}{\delta}\).

  • For the rhs let \(c:= I_s^{\varepsilon}(X;Y)_{P|Q}\). It holds that \[\begin{align} \Pr_{(x,y) \leftarrow P_{XY}}\left\{\frac{P_{Y\mid X=x}(y)}{Q_Y(y)}> 2^c\right\} \leq \varepsilon\,. \end{align}\] Now, let \(\varepsilon_x := \Pr_{y \leftarrow P_{Y\mid X=x}}\left\{\frac{P_{Y\mid X=x}(y)}{Q_Y(y)}> 2^c\right\}\) and \(\mathrm{Good}_x\) be the set of all \(y\) that satisfy \(\frac{P_{Y\mid X=x}(y)}{Q_Y(y)}\leq 2^c\). Then, we have \(\varepsilon= \sum_x P_X(x) \varepsilon_x\). Define random variable \(P'_Y\) jointly correlated with \(P_X\) as \[\begin{align} P'_{Y\mid X=x}(y) = P_{Y\mid X=x}(y)\cdot 1(y\in \mathrm{Good}_x) + \varepsilon_x Q_Y(y), \quad P'_{XY}(x,y)= P_X(x) P'_{Y\mid X=x}(y)\,. \end{align}\] We get \(P'_{Y\mid X=x}(y) \leq (2^c + \varepsilon_x) Q_Y(y) \leq (2^c+1) Q_Y(y)\), which implies \(D_{\max}(P'_{XY}\| P_X\times Q_Y) \leq \log(2^c+1) \leq c+1\). Moreover, we calculate \[\begin{align} T(P_{XY}, P'_{XY}) &=& \frac{1}{2}\|P_{XY}-P'_{XY}\|_1 + \frac{1}{2}|P_{XY}(1_{XY}) - P'_{XY}(1_{XY})|\\ &=& \sum_x P_X(x)\frac{1}{2}\|P'_{Y\mid X=x} - P_{Y\mid X=x}\|_1 + 0 \\ &\leq& \sum_x P_X(x)\frac{1}{2}\left(\varepsilon_x\|Q_Y(y)\|_1 + \|P_{Y\mid X=x}-P_{Y\mid X=x}\cdot1(\mathrm{Good}_x)\|_1 \right) \\ &=& \sum_x P_X(x)\varepsilon_x = \varepsilon\,. \end{align}\] Since \(P'_X = P_X\), we conclude that \(I_{\max}^{\varepsilon, T}(\dot{X};Y)_{P|Q}\leq I_s^{\varepsilon}(X;Y)_{P|Q}+1\).

Eq. ?? is now proved as follows:

  • For the lhs let \(Q_Y\) be the probability distribution achieving the infimum in \(I_{\max}^{\varepsilon, T}( \dot{X} ; Y)_P\). From Eq. 6 and ?? , we obtain \[\begin{align} I_s^{\frac{\varepsilon}{1-\delta}+\delta}(X;Y)_P - 2\log\frac{1}{\delta}\leq I_s^{\frac{\varepsilon}{1-\delta}}(X;Y)_{P|Q} - \log\frac{1}{\delta} \leq I_{\max}^{\varepsilon, T}( \dot{X} ; Y)_{P|Q} = I_{\max}^{\varepsilon, T}( \dot{X} ; Y)_P\,. \end{align}\]

  • For the rhs we use Eq. 5 and 6 to conclude \[\begin{align} I_{\max}^{\varepsilon, T}( \dot{X} ; Y)_P = \inf_{Q_Y}I_{\max}^{\varepsilon, T}( \dot{X} ; Y)_{P|Q}\leq \inf_{Q_Y} I_s^{\varepsilon}(X;Y)_{P|Q}+1 \leq I_s^{\varepsilon}(X;Y)_{P|P}+1 = I_s^{\varepsilon}(X;Y)_{P}+1\,. \end{align}\]

Now, we proceed to the min-entropy constraints.

  • The first inequality is equivalent to first part of Eq. 6 . Let \(d:= H_{\min}^{\varepsilon, T}(X|\dot{Y})_P\) and \(P'_{XY}\) be the (sub normalized) random variables achieving the infimum in the definition. Let \(\mathcal{A}:=\{(x,y): \frac{P_{XY}(x,y)}{P'_{XY}(x,y)} \geq \frac{1}{\delta}\}\). Then, using Lem. 11 we have \[\begin{align} \varepsilon\geq T(P_{XY},P'_{XY}) \geq P_{XY}(\mathcal{A}) - P'_{XY}(\mathcal{A}) \geq P_{XY}(\mathcal{A}) - \delta P_{XY}(\mathcal{A})= (1-\delta)P_{XY}(\mathcal{A})\,. \end{align}\] Thus, we get \(P_{XY}(\mathcal{A}) \leq \frac{\varepsilon}{1-\delta}\). Moreover, for every \((x,y)\in \mathcal{A}^{\mathrm{c}}\), we have \[\begin{align} P_{XY}(x,y)\leq \frac{1}{\delta}P'_{XY}(x,y) \leq \frac{2^{-d}}{\delta} 1_X(x)P_Y(y)\,. \end{align}\] Thus, we find \(H_s^{\frac{\varepsilon}{1-\delta}}(X|Y)_P \geq d - \log\frac{1}{\delta}\).

  • For the second inequality let \(c:= H^{\varepsilon}_s(X|Y)_P\). Then, we have \[\begin{align} \Pr_{(x,y) \leftarrow P_{XY}}\left\{\frac{P_{XY}(x,y)}{P_Y(y)}> 2^{-c}\right\} \leq \varepsilon\,. \end{align}\] Let \(U_X\) be uniform distribution over \(X\) and let \(c':= \log|X| - c\). Then, we find \[\begin{align} \Pr_{(x,y) \leftarrow P_{XY}}\left\{\frac{P_{X\mid Y=y}(x)}{U_X(x)}> 2^{c'}\right\} \leq \varepsilon\,. \end{align}\] which implies \(c' \geq I_{s}^{\varepsilon}(Y; X)_{P|U}\). From Eq. 6 and 5 , we find that \[\begin{align} c'\geq I_{s}^{\varepsilon}(Y; X)_{P|U} \geq I_{\max}^{\varepsilon, T}(\dot{Y}; X)_{P|U} -1 \geq \log|X| - H_{\min}^{\varepsilon, T}(X| \dot{Y})_P -1\,, \end{align}\] and substituting the value of \(c'\), the proof concludes.

 ◻

Instead of smoothing over nearby states that have one of the reduced states (essentially) left intact we could alternatively even smooth over nearby states that have both reduced states (essentially) left intact. This would follow the intuition to smooth the correlations between the systems while leaving the individual systems unchanged and naturally extends to multi-partite scenarios. By iteratively applying the methods from the proof of Thm. 1 we then find similar equivalence statements with the same max-information spectrum divergence based measures. This again leads to an asymptotic equipartition property, however, the expansion only becomes asymptotically tight in first-order (and not in second-order). We note that for many operational problems it seems more adapted to only fix one of the reduced states (see 4).

3.2 Quantum setting↩︎

For quantum problems Uhlmann’s theorem [14] indicates that it is natural to work with fidelity based distance measures such as the purified distance — which is what we will use in this section. Now, the equivalence proof from 3.1 crucially uses the idea of conditioning on the classical side information and hence we cannot give a direct quantum analogue. Instead we find the following equivalence result with the standard smooth max-mutual information (based on a different proof technique).

Theorem 2. Let \(\rho_{AB}\in\mathcal{S}_{\circ}(AB)\) and \(0\leq2\varepsilon+\delta\leq1\) with \(\delta>0\). Then, we have \[\begin{align} \label{eq:max95trace} I^{2\varepsilon+\delta, P}_{\max}( \dot{A} ; B)_\rho\leq I_{\max}^{\varepsilon,P}(A;B)_\rho+\log\frac{8+\delta^2}{\delta^2}\,, \end{align}\qquad{(3)}\] and by definition we also have the opposite inequality \(I_{\max}^{\varepsilon,P }(\dot{A} ; B)_\rho\geq I_{\max}^{\varepsilon,P }(A;B)_\rho\).

Proof. Let \(\tilde{\rho}_{AB}\) and \(\sigma_B\) be the optimizers on the right-hand side of Eq. ?? . Moreover, for some \(\gamma>0\) let \[\begin{align} P_A^\gamma:=\left\{\frac{1}{\gamma}\tilde{\rho}_A-\rho_A\right\}_+\quad\text{and}\quad\bar{\rho}_{AB}:=P_A^\gamma\tilde{\rho}_{AB}P_A^\gamma\,, \end{align}\] where \(\{X\}_+\) denotes the projector onto the positive part of any Hermitian operator \(X\). Let \(V_A\) be the unitary from the polar decomposition of \(\rho_A^{\frac{1}{2}}\bar{\rho}_A^{\frac{1}{2}}\) such that \[\begin{align} F(\rho_A,\bar{\rho}_A)=\mathrm{Tr}\Bigg[\left|\rho_A^{\frac{1}{2}}\bar{\rho}_A^{\frac{1}{2}}\right|\Bigg]=\mathrm{Tr}\left[\rho_A^{\frac{1}{2}}\bar{\rho}_A^{\frac{1}{2}}V_A\right]\,. \end{align}\] For \(\gamma=\frac{\delta^2}{8}\) define the bipartite quantum state \[\begin{align} \hat{\rho}_{AB}:=\underbrace{\rho^{\frac{1}{2}}_AV_A\bar{\rho}_A^{-\frac{1}{2}}\bar{\rho}_{AB}\bar{\rho}_A^{-\frac{1}{2}}V_A^\dagger\rho^{\frac{1}{2}}_A}_{=:\tau_{AB}}+\underbrace{\left(\rho^{\frac{1}{2}}_A(1_A-V_AP^\gamma_AV_A^\dagger)\rho^{\frac{1}{2}}_A\right)\otimes\sigma_B}_{=:\sigma_{AB}}\,, \end{align}\] which by inspection has \(\hat{\rho}_A=\rho_A\). We calculate \[\begin{align} \hat{\rho}_{AB}\leq\;&\left\|(\rho_A\otimes\sigma_B)^{-\frac{1}{2}}\tilde{\rho}_{AB}(\rho_A\otimes\sigma_B)^{-\frac{1}{2}}\right\|_{\infty}\cdot\left(\rho^{\frac{1}{2}}_AV_A\bar{\rho}_A^{-\frac{1}{2}}P_A^\gamma\rho_AP_A^\gamma\bar{\rho}_A^{-\frac{1}{2}}V_A^\dagger\rho^{\frac{1}{2}}_A\right)\otimes\sigma_B\notag\\ &+\left(\rho_A^{\frac{1}{2}}(1_A-V_AP^\gamma_AV_A^\dagger)\rho_A^{\frac{1}{2}}\right)\otimes\sigma_B\,, \end{align}\] and by the definition of \(P_A^\gamma\) we have \(P_A^\gamma\rho_AP_A^\gamma\leq\frac{8}{\delta^2}\cdot\bar{\rho}_A\) as well as \(1_A-P^\gamma_A\leq1_A\) leading to \[\begin{align} \label{eq:proof1} \hat{\rho}_{AB}\leq\left(\frac{8}{\delta^2}\cdot\left\|(\rho_A\otimes\sigma_B)^{-\frac{1}{2}}\tilde{\rho}_{AB}(\rho_A\otimes\sigma_B)^{-\frac{1}{2}}\right\|_{\infty}+1\right)\cdot \rho_A\otimes\sigma_B\,. \end{align}\tag{7}\] Using that \(D_{\max}(\tilde{\rho}_{AB}\|\rho_A\otimes\sigma_B)\geq0\) [5] we get \[\begin{align} \label{eq:proof2} \frac{8}{\delta^2}\cdot\left\|(\rho_A\otimes\sigma_B)^{-\frac{1}{2}}\tilde{\rho}_{AB}(\rho_A\otimes\sigma_B)^{-\frac{1}{2}}\right\|_{\infty}+1\leq\frac{8+\delta^2}{\delta^2}\left\|(\rho_A\otimes\sigma_B)^{-\frac{1}{2}}\tilde{\rho}_{AB}(\rho_A\otimes\sigma_B)^{-\frac{1}{2}}\right\|_{\infty}\,. \end{align}\tag{8}\] Hence, the claim follows as soon as we establish that \(\hat{\rho}_{AB}\) is close enough to \(\rho_{AB}\) in purified distance. Now, notice that \(\mathrm{Tr}[\sigma_{AB}]=1-\mathrm{Tr}[\tau_{AB}]\) and hence \(\bar{\tau}_{AB}:=\frac{\tau_{AB}}{\mathrm{Tr}[\tau_{AB}]}\) and \(\bar{\sigma}_{AB}:=\frac{\sigma_{AB}}{\mathrm{Tr}[1-\tau_{AB}]}\) are normalized. We can then write \[\begin{align} \hat{\rho}_{AB}=\mathrm{Tr}[\tau_{AB}]\cdot\bar{\tau}_{AB}+\big(1-\mathrm{Tr}[\tau_{AB}]\big)\cdot\bar{\sigma}_{AB}\,. \end{align}\] Since the fidelity \(\bar{F}^2(\rho,\sigma)\) is concave in each argument (this follows from the operator concavity of the logarithm) we can estimate \[\begin{align} \bar{F}^2\left(\hat{\rho}_{AB},\bar{\rho}_{AB}\right)&\geq\mathrm{Tr}[\tau_{AB}]\cdot F^2\left(\bar{\tau}_{AB},\bar{\rho}_{AB}\right)+\big(1-\mathrm{Tr}[\tau_{AB}]\big)\cdot\bar{F}^2\left(\bar{\sigma}_{AB},\bar{\rho}_{AB}\right)\\ &\geq\mathrm{Tr}[\tau_{AB}]\cdot\bar{F}^2\left(\bar{\tau}_{AB},\bar{\rho}_{AB}\right)\\ &=\bar{F}^2\left(\tau_{AB},\bar{\rho}_{AB}\right)\,.\label{eq:previous} \end{align}\tag{9}\] By the triangle inequality for the purified distance we get for the quantity of interest \[\begin{align} \label{eq:interest} P\left(\hat{\rho}_{AB},\rho_{AB}\right)\leq P\left(\hat{\rho}_{AB},\bar{\rho}_{AB}\right)+P\left(\rho_{AB},\bar{\rho}_{AB}\right)\,, \end{align}\tag{10}\] and since \(\hat{\rho}_{AB}\) is normalized we get for the first term on the rhs that \[\begin{align} P\left(\hat{\rho}_{AB},\bar{\rho}_{AB}\right)=\sqrt{1-\bar{F}^2\left(\hat{\rho}_{AB},\bar{\rho}_{AB}\right)}\,. \end{align}\] We continue with \[\begin{align} \bar{F}^2\left(\hat{\rho}_{AB},\bar{\rho}_{AB}\right)\geq\bar{F}^2\left(\tau_{AB},\bar{\rho}_{AB}\right)\geq\bar{F}^2\left(\tau_{ABC},\bar{\rho}_{ABC}\right)\,, \end{align}\] where the first step is Eq. 9 and the second step follows since the fidelity is monotone under partial trace (this holds for general non-negative operators) together with choosing \(\tau_{ABC}\) as an extension of \(\tau_{AB}\) and \(\bar{\rho}_{ABC}\) as an extension of \(\bar{\rho}_{AB}\). We choose the purification of \(\bar{\rho}_{AB}\) on \(ABC\) defined through the pure state vector \[\begin{align} |\bar{\rho}_{ABC}\rangle:=\bar{\rho}_A^{\frac{1}{2}}|\Phi_{A:BC}\rangle\,, \end{align}\] where \(|\Phi\rangle_{A:BC}\) denotes the non-normalized maximally entangled pure state vector in the cut \(A:BC\) (on the subspace on \(A\) spanned by the projector \(P_A^\gamma\)). Furthermore, we take the purification of \(\tau_{AB}\) on \(ABC\) given by \[\begin{align} |\tau_{ABC}\rangle:=\rho_A^{\frac{1}{2}}V_A\bar{\rho}_A^{-\frac{1}{2}}|\bar{\rho}_{ABC}\rangle \end{align}\] which is fine since \[\begin{align} \tau_{AB}=\mathrm{Tr}_C\Big[|\tau_{ABC}\rangle\langle{\omega}_{ABC}|\Big]=\rho^{\frac{1}{2}}_AV_A\bar{\rho}_A^{-\frac{1}{2}}\bar{\rho}_{AB}\bar{\rho}_A^{-\frac{1}{2}}V_A^\dagger\rho^{\frac{1}{2}}_A\,. \end{align}\] We calculate \[\begin{align} &\bar{F}^2\left(\tau_{ABC},\bar{\rho}_{ABC}\right)=\left|\langle\bar{\rho}_{ABC}\middle|\tau_{ABC}\rangle\right|^2=\left|\langle\Phi_{A:BC}|\bar{\rho}_A^{\frac{1}{2}}|\tau_{ABC}\rangle\right|^2=\left|\langle\Phi_{A:BC}|\bar{\rho}_A^{\frac{1}{2}}\rho_A^{\frac{1}{2}}V_AP_A^\delta\Phi_{A:BC}\rangle\right|^2\\ &=\left|\mathrm{Tr}\left[\bar{\rho}_A^{\frac{1}{2}}\rho_A^{\frac{1}{2}}V_AP_A^\delta\right]\right|^2=\left|\mathrm{Tr}\left[P_A^\delta\bar{\rho}_A^{\frac{1}{2}}\rho_A^{\frac{1}{2}}V_A\right]\right|^2=\left|\mathrm{Tr}\left[\bar{\rho}_A^{\frac{1}{2}}\rho_A^{\frac{1}{2}}V_A\right]\right|^2=\bar{F}^2\left(\bar{\rho}_A,\rho_A\right)=F^2\left(\bar{\rho}_A,\rho_A\right)\,. \end{align}\] Hence, together with Eq. 10 we arrive at \[\begin{align} \label{eq:arrive-at} P\left(\hat{\rho}_{AB},\rho_{AB}\right)\leq P\left(\bar{\rho}_A,\rho_A\right)+P\left(\bar{\rho}_{AB},\rho_{AB}\right)\leq2\cdot P\left(\bar{\rho}_{AB},\rho_{AB}\right)\,, \end{align}\tag{11}\] where the last step follows from the monotonicity of the purified distance under partial trace. Using again the triangle inequality for the purified distance we then bound \[\begin{align} P\left(\bar{\rho}_{AB},\rho_{AB}\right)&\leq P\left(\bar{\rho}_{AB},\tilde{\rho}_{AB}\right)+P\left(\tilde{\rho}_{AB},\rho_{AB}\right)\\ &\leq P\left(P_A^\gamma\tilde{\rho}_{AB}P_A^\gamma,\tilde{\rho}_{AB}\right)+\varepsilon\\ &\leq\sqrt{2\cdot\mathrm{Tr}\big[(1_A-P_A^\gamma)\tilde{\rho}_A\big]}+\varepsilon\\ &\leq\sqrt{2\cdot\frac{\delta^2}{8}}+\varepsilon=\frac{\delta}{2}+\varepsilon\,. \end{align}\] Together with Eq. 11 we conclude that \(P\left(\hat{\rho}_{AB},\rho_{AB}\right)\leq2\varepsilon+\delta.\). ◻

The standard asymptotic equipartition property for the max-divergence from [1] gives the asymptotic first-order expansion \[\begin{align} \label{eq:Imax-asymptotic} \lim_{n\to\infty}\frac{1}{n}I_{\max}^{\varepsilon,P}(\dot{A}^n; B^n)_{\rho^{\otimes n}} = I(A \!:\! B)_\rho,\quad\text{with the quantum mutual information I(A \!:\! B)_\rho.} \end{align}\tag{12}\] We also find the following equivalence result for the smooth conditional min-entropy.

Theorem 3. Let \(\rho_{AB}\in\mathcal{S}_{\circ}(AB)\) and \(0\leq2\varepsilon+\delta\leq1\) with \(\delta>0\). Then, we have4 \[\begin{align} H^{2\varepsilon+\delta,P}_{\min}(A|\dot{B})_\rho\geq H_{\min}^{\varepsilon,P}(A|B)_\rho-\log\frac{8+\delta^2}{\delta^2}\,, \end{align}\] and by definition we also have the opposite inequality \(H_{\min}^{\varepsilon,P}(A|\dot{B})_\rho\leq H_{\min}^{\varepsilon,P}(A|B)_\rho\).

Proof. The first part of the proof is very similar to the proof of Thm. 2, just with the roles of the systems \(A\) and \(B\) interchanged. In the following we only sketch the steps which are different. For \(\tilde{\rho}_{AB}\in \mathcal{S}_{\bullet}(AB)\) the optimizer in \(H_{\min}^{\varepsilon,P}(A|B)_\rho\) we define the bipartite quantum state \[\begin{align} \hat{\rho}_{AB}:=\rho^{\frac{1}{2}}_BV_B\bar{\rho}_B^{-\frac{1}{2}}\bar{\rho}_{AB}\bar{\rho}_B^{-\frac{1}{2}}V_B^\dagger\rho^{\frac{1}{2}}_B+\frac{1_A}{|A|}\otimes\left(\rho^{\frac{1}{2}}_B(1_B-V_BP^\gamma_BV_B^\dagger)\rho^{\frac{1}{2}}_B\right)\,, \end{align}\] with \(P_B^\gamma\), \(\bar{\rho}_{AB}\), and \(V_B\) as in the proof of Thm. 2 (where \(A\leftrightarrow B)\). We then find similarly as in the proof of Thm. 2 that \[\begin{align} \hat{\rho}_{AB}\leq\left(\frac{8}{\delta^2}\cdot\left\|\rho_B^{-\frac{1}{2}}\tilde{\rho}_{AB}\rho_B^{-\frac{1}{2}}\right\|_{\infty}+\frac{1}{|A|}\right)\cdot1_A\otimes\rho_B\,, \end{align}\] and using that \(D_{\max}(\tilde{\rho}_{AB}\|1_A\otimes\rho_B)\geq-\log|A|\) [8] we get \[\begin{align} \frac{8}{\delta^2}\left\|\rho_B^{-\frac{1}{2}}\tilde{\rho}_{AB}\rho_B^{-\frac{1}{2}}\right\|_{\infty}+\frac{1}{|A|}\leq\frac{8+\delta^2}{\delta^2}\left\|\rho_B^{-\frac{1}{2}}\tilde{\rho}_{AB}\rho_B^{-\frac{1}{2}}\right\|_{\infty}\,. \end{align}\] As in the proof of Thm. 2 this leads to the statement \[\begin{align} H^{2\varepsilon+\delta,P}_{\min}(A|\dot{B})\geq -D_{\max}(\tilde{\rho}_{AB}\|1_A\otimes\rho_B)-\log\frac{8+\delta^2}{\delta^2}\,, \end{align}\] concluding the proof. ◻

Employing the standard asymptotic equipartition property from [16] this implies the asymptotic first-order expansion \[\begin{align} \label{eq:Hmin-asymptotic} \lim_{n\to\infty}\frac{1}{n}H_{\min}^{\varepsilon,P}(A^n|\dot{B}^n)_{\rho^{\otimes n}} = H(A|B)_\rho,\quad\text{with the conditional entropy H(A|B)_\rho.} \end{align}\tag{13}\]

4 Operational Examples↩︎

4.1 Overview↩︎

It is generally neat that existing proofs and protocols readily apply and give tight bounds when combined with our novel restricted smoothing. In the following we discuss various basic classical and quantum examples in bipartite settings.

4.2 Classical state splitting↩︎

Let \(\varepsilon\in (0,1]\) be the error parameter. There are two parties Alice and Bob. Alice possesses random variable \(X\), taking values over a finite set \(\mathcal{X}\) and a random variable \(Y\), taking values over a finite set \(\mathcal{Y}\). Alice sends a message to Bob and at the end Bob outputs random variable \(\hat{Y}\) such that \(T(P_{XY},P_{X\hat{Y}})\leq \varepsilon\). They are allowed to use shared randomness between them which is independent of \(XY\) at the beginning of the protocol.

We note that a generalization of this task (with additional side information) was studied in [13]. These results together with [17] imply that the minimal number \(R(P_{XY},\varepsilon)\) of bits communicated from Alice to Bob to achieve classical state splitting with error \(\varepsilon\in(0,1]\) in generalized trace distance is bounded as \[\begin{align} \label{eq:state-splitting95previous} I_s^{\varepsilon/(1-\delta)}(P_{XY}\| P_X \times P_Y) - \log\frac{1}{\delta}\leq R(P_{XY},\varepsilon)\leq I_s^{\varepsilon- 3\delta}(P_{XY}\| P_X \times P_Y) + 2\log\frac{1}{\delta}\,, \end{align}\tag{14}\] for \(\delta \in (0,1)\) small enough. We show an even tighter characterization in terms of the smooth max-information.

Theorem 4. Let \(P_{XY}\in\mathcal{S}_{\circ}(XY)\). Then, the minimal number \(R(P_{XY},\varepsilon)\) of bits communicated from Alice to Bob to achieve classical state splitting with error \(\varepsilon\in(0,1]\) in generalized trace distance is bounded as \[\begin{align} \label{eq:state-splitting} I_{\max}^{\varepsilon,T}(\dot{X};Y)_P\leq R(P_{XY},\varepsilon)\leq I_{\max}^{\varepsilon-\delta,T}(\dot{X};Y)_P+\log\log\frac{1}{\delta}+1,\quad\text{for any \delta\in(0,\varepsilon].} \end{align}\qquad{(4)}\] In particular, this implies the asymptotic second-order expansion5 \[\begin{align} \label{eq:state-splitting2} \frac{1}{n}R\left(P_{XY}^{\times n},\varepsilon\right)=I(X;Y)_P+\sqrt{\frac{V(X;Y)_P}{n}}\cdot\Phi^{-1}(\varepsilon)+O\left(\frac{\log n}{n}\right)\,. \end{align}\qquad{(5)}\]

Proof. Eq. ?? immediately follow from Eq. ?? and Thm. 1. The converse in Eq. ?? is as follows, which uses the converse argument from  [13]. Let \(T\) be Alice’s message and \(S\) be shared randomness. Observe that \(P_{XS} = P_{X}\times P_S\), let \(\mathcal{D}: ST \rightarrow Y\) be Bob’s decoding operation, \(U_T\) be the uniformly distributed over \(T\), and let the output random variable after Alice’s message and Bob’s decoding be \(P'_{XY}:= (1_X\times\mathcal{D})(P_{XST})\). Now, we consider that \[\begin{align} R &\geq D_{\max}(P_{XST}\| P_{XS}\times U_T) = D_{\max}(P_{XST}\| P_{X}\times P_S\times U_T)\nonumber\\ &\geq D_{\max}((1_X\times\mathcal{D})(P_{XST})\| P_{X}\times \mathcal{D}(P_S\times U_T))\nonumber\\ &=D_{\max}(P'_{XY}\| P_X \times \mathcal{D}(P_S\times U_T)) \geq \min_{Q_Y}D_{\max}(P'_{XY}\| P_X \times Q_Y)\,. \end{align}\] We then use the fact that \(P'_X=P_X\) and \(T(P'_{XY}, P_{XY}) \leq \varepsilon\) to further lower bound \(R\) by \(I_{\max}^{\varepsilon,T}(\dot{X};Y)_P\).
The achievability in Eq. ?? uses the rejection sampling argument [4], [18]. Let \(P'_{XY}, Q_Y\) be random variables achieving the infimum in the definition of \(I_{\max}^{\varepsilon- \delta}(X;Y)_P\). Let \(K:= I_{\max}^{\varepsilon- \delta}(X;Y)_P\) and \(R:= K+ \log\log\frac{1}{\delta}\). By definition, it holds that \(P'_{XY} \leq 2^K P_X \times Q_Y\) and \(P'_X=P_X\). Thus we conclude that for all \(x\) satisfying \(P_X(x)>0\), \(P'_{Y\mid X=x} \leq 2^K Q_Y\).

The protocol \(\mathcal{P}\): Alice and Bob share the random variable \(P_X \times Q_{Y_1}\times \ldots Q_{Y_{2^R}}\), with \(X\) belonging to Alice and \(Y_1, \ldots Y_{2^R}\) acting as shared randomness between Alice and Bob. They proceed in the following step, with Alice obtaining a sample \(x\) from \(P_X(x)\).

  1. Alice sets \(i=1\).

  2. (While \(i\leq 2^R\)):

  3. Alice takes a sample \(y\) from \(Q_{Y_i}\).

  4. With probability \(\frac{P'_{Y\mid X=x}(y)}{2^K Q_Y(y)}\) she accepts this sample, sends \(i\) to Bob and exits the while loop.

  5. With probability \(1-\frac{P'_{Y\mid X=x}(y)}{2^K Q_Y(y)}\) she updates \(i\rightarrow i+1\) and goes to Step \(2\) (End While).

  6. If \(i> 2^R\), Alice sends \(2^R+1\) to Bob.

  7. Bob receives Alice’s message, which we call \(j\). If \(j> 2^R\), Bob outputs a sample distributed as \(Q_Y\). Else he outputs the sample from \(Q_{Y_j}\).

Let the output of Bob be \(P''_{Y\mid X=x}\).

Analysis of the protocol: The probability of Alice’s acceptance on Step \(4\) is \[\sum_y Q_Y(y)\frac{P'_{Y\mid X=x}(y)}{2^K Q_Y(y)} = 2^{-K}\sum_y P'_{Y\mid X=x}(y) = 2^{-K}.\] Conditioned on Alice’s acceptance, the distribution of \(Y_i\) is equal to \(P'_{Y\mid X=x}\). To argue this, observe that the probability of any \(y\), conditioned on acceptance, is equal to \[\frac{1}{2^{-K}}\cdot Q_Y(y)\cdot\frac{P'_{Y\mid X=x}(y)}{2^K Q_Y(y)} = P'_{Y\mid X=x}(y).\] Thus, conditioned on the event that Alice accepts an \(i\), the sample output by Bob (which is the same as that observed by Alice) is distributed as \(P'_{Y\mid X=x}\). Let \(\gamma\) be the probability that Alice does not find any sample, that is, \(i> 2^R\). Then Bob’s output \(P''_{Y\mid X=x}\) is equal to \((1-\gamma)P'_{Y\mid X=x} + \gamma Q_{Y}\). Let \(P''_{XY}:= P_XP''_{Y\mid X}\) be the overall output distribution.

Since probability of acceptance at any step is equal to \(2^{-K}\), we have \[\gamma = (1- 2^{-K})^{2^R} \leq \left(2^{-2^{-K}}\right)^{2^R} = 2^{-2^{R-K}} = 2^{- 2^{\log\log\frac{1}{\delta}}}= 2^{- \log\frac{1}{\delta}} = \delta.\] Consider, \[\begin{align} T(P''_{XY}, P_{XY}) &=& \sum_xP_X(x)T(P''_{Y\mid X=x}, P_{Y\mid X=x})\\ &=& \sum_xP_X(x)T((1-\gamma)P'_{Y\mid X=x} + \gamma Q_{Y}, P_{Y\mid X=x}))\\ &\leq& (1-\gamma)\sum_xP_X(x)T(P'_{Y\mid X=x}, P_{Y\mid X=x})) + \gamma\sum_xP_X(x)T(Q_{Y}, P_{Y\mid X=x})\\ &\leq& (1-\gamma) T(P'_{XY}, P_{XY}) + \gamma \leq \varepsilon-\delta + \gamma \leq \varepsilon. \end{align}\] Furthermore, the number of bits communicated is \(\log(2^R+1)\leq R+1\), which completes the proof. ◻

4.3 Strong privacy amplification against side information↩︎

For a set of two-universal hash functions \(\{f_{X\to Z}^s\}_{s\in S}\) and classical-quantum states \[\begin{align} \rho_{XB}=\sum_{X\in\mathcal{X}}|x\rangle\langle x|\otimes\rho_B^x\in\mathcal{S}_\circ(XB) \end{align}\] we use the same composable security criterion for \(\varepsilon\)-random and secret bits as, e.g, in [1], \[\begin{align} \label{eq:security-criterion} T\Bigg(\underbrace{\frac{1}{|S|}\sum_{\substack{s\in S\\z\in\mathcal{Z}}}|s\rangle\langle s|_S\otimes|z\rangle\langle z|\otimes\Bigg(\sum_{x:f^s(x)=z}\rho_B^x\Bigg)}_{=:\;\omega_{SZB}},\frac{1_S}{|S|}\otimes\frac{1_Z}{|Z|}\otimes\rho_B\Bigg)\leq\varepsilon\,. \end{align}\tag{15}\] Note that in contrast to the setting studied in [19] or [15] we have a composable security definition by putting the reduced state on \(B\) on the lhs of Eq. 15 . We refer to [20] for a more detailed discussion.

Theorem 5. Let \(\rho_{XB}\in\mathcal{S}_{\circ}(XB)\) be classical-quantum on \(XB\) and \(\varepsilon\in(0,1]\). Then, the maximal number of \(\varepsilon\)-random and secret bits \(\ell(\rho_{XB},\varepsilon)\) that can be extracted from \(\rho_{XB}\) is bounded as \[\begin{align} \label{eq:pa-qsi} H_{\min}^{\varepsilon-\delta,P}(X|\dot{B})_\rho-\log\frac{1}{\delta^4}\leq\ell(\rho_{XB},\varepsilon)\leq H_{\min}^{\varepsilon,P}(X|\dot{B})_\rho,\quad\text{for any \delta\in(0,\varepsilon].} \end{align}\qquad{(6)}\] Moreover, when \(B = Y\) is classical then we also have \[\begin{align} \label{eq:pa-cl} H_{\min}^{\varepsilon-\delta,T}(X|\dot{Y})_\rho-\log\frac{1}{4\delta^2}\leq\ell(\rho_{XY},\varepsilon)\leq H_{\min}^{\varepsilon,T}(X|\dot{Y})_\rho\,. \end{align}\qquad{(7)}\] In particular, this implies the asymptotic second-order expansion \[\begin{align} \frac{1}{n}\ell\left(P_{XY}^{\times n},\varepsilon\right)=H(X|Y)_P+\sqrt{\frac{V(X|Y)_P}{n}}\cdot\Phi^{-1}(\varepsilon)+O\left(\frac{\log n}{n}\right) \end{align}\] as first given in [21] (see also [22]).

Proof. We first prove the lower bound in Eq. ?? . Let \(\tilde{\rho}_{XB}\in\mathcal{S}_{\bullet}(XB)\) be the optimizer in the definition of \(H_{\min}^{\varepsilon-\delta,P}(X|\dot{B})_\rho\) and let \[\begin{align} \label{eq:pa95tripartite} \tilde{\omega}_{SZB}:=\frac{1}{|S|}\sum_{\substack{s\in S\\z\in\mathcal{Z}}}|s\rangle\langle s|_S\otimes|z\rangle\langle z|\otimes\Bigg(\sum_{x:f^s(x)=z}\tilde{\rho}_B^x\Bigg)\,. \end{align}\tag{16}\] Since by definition \(\rho_B\geq\tilde{\rho}_B\) and by data-processing \(P\left(\omega_{SZB},\tilde{\omega}_{SZB}\right)\leq\varepsilon-\delta\) we get that \[\begin{align} P\left(\omega_{SZB},\frac{1_S}{|S|}\otimes\frac{1_Z}{|Z|}\otimes\rho_B\right)&\leq P\left(\omega_{SZB},\frac{1_S}{|S|}\otimes\frac{1_Z}{|Z|}\otimes\tilde{\rho}_B\right)\\ &\leq P\left(\tilde{\omega}_{SZB},\frac{1_S}{|S|}\otimes\frac{1_Z}{|Z|}\otimes\tilde{\rho}_B\right)+P\left(\omega_{SZB},\tilde{\omega}_{SZB}\right)\\ &\leq P\left(\tilde{\omega}_{SZB},\frac{1_S}{|S|}\otimes\frac{1_Z}{|Z|}\otimes\tilde{\rho}_B\right)+\varepsilon-\delta\\ &\leq \sqrt{2T\left(\tilde{\omega}_{SZB},\frac{1_S}{|S|}\otimes\frac{1_Z}{|Z|}\otimes\tilde{\rho}_B\right)}+\varepsilon-\delta\,, \end{align}\] where in the last step we employed the equivalence of generalized trace distance and purified distance [1]. Now, standard achievability proofs such as [15] applied to \(\tilde{\rho}_{XB}\in\mathcal{S}_{\bullet}(XB)\) give \[\begin{align} T\left(\tilde{\omega}_{SZB},\frac{1_S}{|S|}\otimes\frac{1_Z}{|Z|}\otimes\tilde{\rho}_B\right)\leq\frac{1}{2}\sqrt{|Z|\cdot2^{-H_{\min}^{\varepsilon-\delta,P}(X|\dot{B})_\rho}}\,, \end{align}\] and choosing \(\log|Z|=H_{\min}^{\varepsilon-\delta,P}(X|\dot{B})_\rho-\log\frac{1}{\delta^4}\) leads to the claim. For the upper bound in Eq. ?? we follow [1] but adapted to our partially smoothed conditional min-entropy. Namely, assume by contradiction that there exists a protocol which extracts \(\ell>H_{\min}^{\varepsilon,P}(X|\dot{B})_\rho\) bits of \(\varepsilon\)-random and secret bits. Then, since applying a function on \(X\) cannot increase the smooth conditional min-entropy (Lem. 6) we have for all \(s\in\mathcal{S}\) that \[\begin{align} \ell>H_{\min}^{\varepsilon,P}(X|\dot{B})_\rho\geq H_{\min}^{\varepsilon,P}(Z|\dot{B})_{\rho^s},\quad\text{with \rho_{ZB}^s:=\sum_{z\in\mathcal{Z}}|z\rangle\langle z|_Z\otimes\left(\sum_{x:f^s(x)=z}\rho_B^x\right).} \end{align}\] Hence, for all \(\tilde{\rho}_{ZB}\in\mathcal{S}_{\bullet}(ZB)\) with \(P\left(\tilde{\rho}_{ZB},\rho^s_{ZB}\right)\leq\varepsilon\) we have \(H_{\min}(Z|B)_{\tilde{\rho}}<\ell\). This in turn implies \[\begin{align} P\left(\rho_{ZB}^s,\frac{1_Z}{|Z|}\otimes\rho_B\right)>\varepsilon\quad\implies\quad P\left(\omega_{SZB},\frac{1_S}{|S|}\otimes\frac{1_Z}{|Z|}\otimes\rho_B\right)>\varepsilon\,, \end{align}\] which is in contradiction to Eq. 15 and thus concludes the proof.

The upper bound in Eq. ?? follows in the same way as the upper bound in Eq. ?? , just by noting that in the classical case Lem. 6 also holds for the generalized trace distance. For the lower bound in Eq. ?? , denote in the security criterion Eq. 15 the state \(\omega_{SZB}\) for \(B=Y\) classical by \(Q_{SZY}\), let \(\tilde{P}_{XY}\in\mathcal{S}_{\bullet}(XY)\) be the optimizer in the definition of \(H_{\min}^{\varepsilon-\delta,T}(X|\dot{Y})_P\), and let \(\tilde{Q}_{SZY}\) be defined as in Eq. 16 . Since by definition \(P_Y\geq\tilde{P}_Y\) and by data-processing \(T\left(Q_{SZY},\tilde{Q}_{SZY}\right)\leq\varepsilon-\delta\) we get that \[\begin{align} T\left(Q_{SZY},\frac{1_S}{|S|}\times\frac{1_Z}{|Z|}\times P_Y\right)&\leq T\left(\tilde{Q}_{SZY},\frac{1_S}{|S|}\times\frac{1_Z}{|Z|}\times\tilde{P}_Y\right)+T\left(Q_{SZY},\tilde{Q}_{SZY}\right)\\ &\leq T\left(\tilde{Q}_{SZY},\frac{1_S}{|S|}\times\frac{1_Z}{|Z|}\times\tilde{P}_Y\right)+\varepsilon-\delta\,. \end{align}\] Now, standard achievability proofs such as [15] applied to \(\tilde{P}_{XY}\in\mathcal{S}_{\bullet}(XY)\) lead to the claim for \(\log|Z|=H_{\min}^{\varepsilon-\delta,T}(X|\dot{Y})_P-\log\frac{1}{4\delta^2}\). ◻

For the case of quantum side information we do not have the asymptotic second-order expansion of \(H_{\min}^{\varepsilon,P}(X|\dot{B})_\rho\) and thus we cannot give the asymptotic second-order expansion of \(\ell(\rho_{XB},\varepsilon)\). This seems to be an open problem for the composable security definition used here (see [23] for a discussion of the Markovian case). However, note that by Thm. 5 finding the asymptotic second-order expansion of \(H_{\min}^{\varepsilon,P}(X|\dot{B})_\rho\) is now equivalent to finding the asymptotic second-order expansion of \(\ell(\rho_{XB},\varepsilon)\). Hence, we believe that our ideas provide a promising approach to study the question.

4.4 Quantum state merging↩︎

A pure tripartite state \(\rho_{ABR}\) is shared between parties Alice (A), Bob (B), and the reference \(R\). The goal is to send the \(A\)-marginal from Alice to Bob using classical communication and entanglement assistance while not changing the overall state [24][27]. More precisely, for \(\rho_{ABR}\in\mathcal{S}_\circ(ABR)\) of rank-one and \(A_{0}B_{0}\) additional quantum systems, a quantum channel \[\begin{align} \mathcal{E}:AA_{0}\otimes BB_{0}\rightarrow A_{1}\otimes B_{1}\bar{B}B \end{align}\] is a quantum state merging of \(\rho_{ABR}\) with error \(\varepsilon\in[0,1]\), if it is a local operation and classical forward communication process for the bipartition \(AA_{0}\rightarrow A_{1}\) versus \(BB_{0}\rightarrow B_{1}\bar{B}B\), and \[\begin{align} P\Big((\mathcal{E}_{AA_{0}BB_{0}\to A_{1}B_{1}\bar{B}B}\otimes\mathcal{I}_R)(\Phi_{A_{0}B_{0}}\otimes\rho_{ABR}),\Phi_{A_{1}B_{1}}\otimes\rho_{B\bar{B}R}\Big)\leq\varepsilon\,, \end{align}\] where \(\rho_{B\bar{B}R}=(\mathcal{I}_{A\to\bar{B}}\otimes\mathcal{I}_{BR})(\rho_{ABR})\), and \(\Phi_{A_{0}B_{0}}\), \(\Phi_{A_{1}B_{1}}\) are maximally entangled states on \(A_{0}B_{0}\), \(A_{1}B_{1}\), respectively. The difference \(\log|A_0|-\log|A_1|\) quantifies the entanglement cost.

Theorem 6. Let \(\rho_{ABR}\in\mathcal{S}_\circ(ABR)\) be of rank-one. For free classical communication assistance the minimal entanglement cost \(E(\rho_{ABR},\varepsilon)\) for quantum state merging of \(\rho_{ABR}\) with error \(\varepsilon\in(0,1]\) in purified distance is bounded as \[\begin{align} \label{eq:entanglement-cost} -H_{\min}^{\varepsilon,P}(A|\dot{R})_\rho\leq E(\rho_{ABR},\varepsilon)\leq -H_{\min}^{\varepsilon-\delta,P}(A|\dot{R})_\rho+\log\frac{1}{\delta^4},\quad\text{for any \delta\in(0,\varepsilon].} \end{align}\qquad{(8)}\] Alternatively, for unlimited entanglement assistance — not necessarily constraint to the form of maximally entangled states — the minimal classical communication cost \(C(\rho_{ABR},\varepsilon)\) for quantum state merging of \(\rho_{ABR}\) with error \(\varepsilon\in(0,1]\) in purified distance is bounded as \[\begin{align} \label{eq:classical-cost} I_{\max}^{\varepsilon,P}(\dot{R};A)_\rho\leq C(\rho_{ABR},\varepsilon)\leq I_{\max}^{\varepsilon-\delta,P}(\dot{R};A)_\rho+\log\frac{1}{\delta^4},\quad\text{for any \delta\in(0,\varepsilon].} \end{align}\qquad{(9)}\]

Proof. For the lower bounds in Eq. ?? and Eq. ?? we first model the form of a general quantum state merging protocol.

On Alice’s side, we consider an arbitrary local operation from \(A A_0\) to a classical register \(X_{A}\) (to be sent to Bob) and \(A_1\). We consider an isometric purification of this operation in two steps. First, Alice performs an isometry \(U\) from \(AA_{0}\) to \(A_{1}X_{A}A'\), where \(X_{A}\) is a register that is to be measured and \(A'\) an arbitrary garbage register to be discarded. Then, the measurement of \(X_{A}\) and the communication to Bob is modelled by the isometry \(V=\sum_{x}|xxx\rangle_{X_{A}X_{B}X_{R}}\langle x|_{X_{A}}\) from \(X_{A}\) to \(X_{A}X_{B}X_{R}\), where \(X_B\) is a classical copy of \(X_A\) to be sent to Bob and \(X_R\) is a coherent copy of \(X_A\) and \(X_B\) that is used to purify this operation.

On Bob’s side, we consider an arbitrary local operation from \(B_{0}X_{B}B\) to \(B_{1}\bar{B}BX_{B}\), where \(\bar{B}B\) is the merged state. This can again be purified to an isometry \(W\) from \(B_{0}BX_{B}\) to \(B_{1}\bar{B}BX_{B}B'\), with \(B'\) an arbitrary register to be discarded. Overall, the pure state vector \(|\Phi\rangle_{A_{0}B_{0}}\otimes|\rho\rangle_{ABR}\) is taken to some pure state vector \[\begin{align} \text{|\omega\rangle_{A_{1}B_{1}\bar{B}BRX_{A}X_{B}X_{R}A'B'} \varepsilon-close in purified distance to |\Phi\rangle_{A_{1}B_{1}}\otimes|\rho\rangle_{\bar{B}BR}\otimes|\xi\rangle_{X_{A}X_{B}X_{R}A'B'},} \end{align}\] where \(|\xi\rangle_{X_{A}X_{B}X_{R}A'B'}\) denotes another pure state vector. By basic properties of the purified distance [28], this latter vector has without loss of generality the form \[\begin{align} |\xi\rangle_{X_{A}X_{B}X_{R}A'B'}=\sum_{x}\sqrt{p_{x}}|xxx\rangle_{X_{A}X_{B}X_{R}}\otimes|\xi^{x}\rangle_{A'B'}\,, \end{align}\] where \(\{p_{x}\}\) denotes some probability distribution and the \(|\xi^{x}\rangle_{A'B'}\) are pure state vectors on \(A'B'\).

For the lower bound in Eq. ?? we follow [26] and analyse the correlations between Alice and the reference system, measured in terms of the smooth conditional min-entropy. For the modelling as above we estimate \[\begin{align} \log|A_0|+H_{\min}^{\varepsilon,P}(A|\dot{R})_{\rho}&\geq H_{\min}^{\varepsilon,P}(A_{0}A|\dot{R})_{\Phi\otimes\rho}\tag{17}\\ &=H_{\min}^{\varepsilon,P}(A_{1}A'X_{A}|\dot{R})_{U(\Phi\otimes\rho)U^{\dagger}}\tag{18}\\ &\geq H_{\min}(A_{1}A'X_{A}|\dot{R})_{V^\dagger(\Phi\otimes\rho\otimes\xi)V}\\ &=-D_{\max}\left(V^\dagger\left(\frac{1_{A_1}}{|A_1|}\otimes\rho_{R}\otimes\xi_{X_{A}X_{B}X_{R}A'}\right)V\middle\|1_{A_1A'X_A}\otimes\rho_{R}\right)\\ &\geq -D_{\max}\left(\frac{1_{A_1}}{|A_1|}\otimes\rho_{R}\otimes\xi_{X_{A}X_{R}A'}\middle\|1_{A_1A'X_A}\otimes\xi_{X_{R}}\otimes\rho_{R}\right)\tag{19}\\ &=\log|A_1|-D_{\max}(\xi_{X_{A}X_{R}A'}\|1_{A'X_A}\otimes\xi_{X_{R}})\\ &\geq\log|A_1|\tag{20}\,, \end{align}\] where in Eq. 17 we used the dimension upper bound for the smooth conditional min-entropy from Lem. 7, in Eq. 18 the isometric invariance of the smooth conditional min-entropy (Lem. 4), in Eq. 19 the monotonicity property under projective measurements when conditioned on the measurement outcomes from Lem. 9, and in Eq. 20 that \(D_{\max}(\xi_{X_{A}X_{R}A'}\|1_{A'X_A}\otimes\xi_{X_{R}})\geq0\) following from the classical-quantum structure of \(\xi_{X_{A}X_{R}A'}\) [2].

For the lower bound in Eq. ?? we analyse the correlations between Alice and the reference system, measured in terms of the smooth max-information. For the modelling as above we find \[\begin{align} I_{\max}^{\varepsilon,P}(\dot{R};A)_{\rho}&\leq I_{\max}^{\varepsilon,P}(\dot{R};A_{0}A)_{\Phi\otimes\rho}\\ &=I_{\max}^{\varepsilon,P}(\dot{R};A_{1}A'X_{A}X_{B}X_{R})_{\omega}\\ &\leq I_{\max}^{\varepsilon,P}(\dot{R};A_{1}A'X_{A}X_{R})_{\omega}+\log|X|\\ &\leq\log|X|\,, \end{align}\] where the first step is due to the monotonicity of the smooth max-information under local quantum operations (Lem. 3), the second step due to the invariance of the smooth max-information under isometries (Lem. 4), the third step due to the dimension upper bound on the smooth max-information of coherent classical states from Lem. 8, and the forth step follows because the output state has to be \(\varepsilon\)-close in purified distance to the perfect state (which has no correlations to \(R\)). Note that this chain of arguments does not depend on the structure of the entanglement assistance and thus also applies to assistance that is not necessarily constraint to the form of maximally entangled states.

The upper bound in Eq. ?? follows from the analysis in [26] adapted to our partially smoothed conditional min-entropy. The protocol is such that Alice applies a Haar random rank-\(|A_1|\) projective measurement to decouple her systems from the reference, sends the resulting classical measurement outcomes to Bob, who then recovers the full state by Uhlmann’s theorem. In particular, fix \(N\) orthogonal subspaces of dimension \(|A_1|\) on \(AA_0\), denote the projectors on these subspaces followed by a fixed unitary mapping it to \(A_1\) by \(P^x_{A_0A\to A_1}\), and define the isometry \[\begin{align} W_{A_{0}A\to A_{1}X_{A}X_{B}}=\sum_{x}P^{x}_{A_{0}A\to A_{1}}\otimes|x\rangle_{X_{A}}\otimes|x\rangle_{X_{B}}\,. \end{align}\] Now, by standard one-shot decoupling results as for example outlined in [27], there exists a unitary operator \(U_{A_{0}A}\) such that for \[\begin{align} \omega_{A_{1}X_{A}X_{B}B_{0}BR}=W_{A_{0}A\to A_{1}X_{A}X_{B}}U_{A_{0}A} \left(\Phi_{A_0B_0} \otimes \rho_{ABR}\right) U_{A_{0}A}^{\dagger}W^\dagger_{A_{0}A\to A_{1}X_{A}X_{B}}\,, \end{align}\] we have \[\begin{align} T\left(\omega_{A_1X_AR},\tau_{A_1X_A}\otimes\rho_R\right)\leq\frac{1}{2}\sqrt{|A_1|\cdot2^{-H_{\min}(A_0A|R)_{\Phi\otimes\rho}}}=\frac{1}{2}\sqrt{\frac{|A_1|}{|A_0|}\cdot2^{-H_{\min}(A|R)_\rho}}\,, \end{align}\] where \(\tau_{A_1X_AX_B}=W_{A_{0}A\to A_{1}X_{A}X_{B}}\left(\frac{1_{A_0}}{|A_0|}\otimes\frac{1_{X_A}}{|X_A|}\right)W^\dagger_{A_{0}A\to A_{1}X_{A}X_{B}}\) and we have used the additivity of the conditional min-entropy [6]. By the equivalence of generalized trace distance and purified distance [1] this implies \[\begin{align} P\left(\omega_{A_1X_AR},\tau_{A_1X_A}\otimes\rho_R\right)\leq\left(\frac{|A_1|}{|A_0|}\cdot2^{-H_{\min}(A|R)_\rho}\right)^{\frac{1}{4}}\,. \end{align}\] Moreover, by Uhlmann’s theorem there exists an isometry \(V_{BB_{0}X_{B}\to BB'B_{1}X_{B}}\) with \[\begin{align} \label{eq:uhlmann-merging} &P\left(\omega_{A_1X_AR},\tau_{A_1X_A}\otimes\rho_R\right)\notag\\ &=P\left(V_{BB_{0}X_{B}\rightarrow BB'B_{1}X_{B}}\big(\omega_{A_{1}X_{A}X_{B}B_{0}BR}\big)V^\dagger_{BB_{0}X_{B}\to BB'B_{1}X_{B}},\tau_{X_AX_B}\otimes\Phi_{A_1B_1}\otimes\rho_{BB'R}\right)\,. \end{align}\tag{21}\] We conclude that applying the isometry \(W_{A_{0}A\to A_{1}X_{A}X_{B}}\), sending \(X_B\) to Bob, and then applying the isometry \(V_{BB_{0}X_{B}\rightarrow BB'B_{1}X_{B}}\), realises quantum state merging for an \[\begin{align} \text{entanglement cost}\quad\log|A_0|-\log|A_1|\quad\text{and error}\quad\left(\frac{|A_1|}{|A_0|}\cdot2^{-H_{\min}(A|R)_\rho}\right)^{\frac{1}{4}}\,. \end{align}\] Now, let \(\tilde{\rho}_{AR}\in\mathcal{S}_{\bullet}(AR)\) be the optimizer in \(H_{\min}^{\varepsilon-\delta,P}(A|\dot{R})_\rho\) and choose \(A_0B_0\) and \(A_1B_1\) such that \[\begin{align} \label{eq:choice-merging} \log|A_0|-\log|A_1|=-H_{\min}^{\varepsilon-\delta,P}(A|\dot{R})_\rho+\log\frac{1}{\delta^4}\,. \end{align}\tag{22}\] For the quantum state \[\begin{align} \tilde{\omega}_{A_{1}X_{A}X_{B}B_{0}BR}:=W_{A_{0}A\to A_{1}X_{A}X_{B}}U_{A_{0}A} \left(\Phi_{A_0B_0} \otimes \tilde{\rho}_{ABR}\right) U_{A_{0}A}^{\dagger}W^\dagger_{A_{0}A\to A_{1}X_{A}X_{B}} \end{align}\] we estimate \[\begin{align} P\left(\omega_{A_1X_AR},\tau_{A_1X_A}\otimes\rho_R\right)&\leq P\left(\omega_{A_1X_AR},\tau_{A_1X_A}\otimes\tilde{\rho}_R\right)\tag{23}\\ &\leq P\left(\tilde{\omega}_{A_1X_AR},\tau_{A_1X_A}\otimes\tilde{\rho}_R\right)+P\left(\tilde{\omega}_{A_1X_AR},\omega_{A_1X_AR}\right)\tag{24}\\ &\leq P\left(\tilde{\omega}_{A_1X_AR},\tau_{A_1X_A}\otimes\tilde{\rho}_R\right)+\varepsilon-\delta\tag{25}\\ &\leq\left(\frac{|A_1|}{|A_0|}\cdot2^{-H_{\min}^{\varepsilon-\delta,P}(A|\dot{R})_\rho}\right)^{\frac{1}{4}}+\varepsilon-\delta\\ &\leq\varepsilon\,,\tag{26} \end{align}\] where in Eq. 23 we used that by definition \(\rho_R\geq\tilde{\rho}_R\), in Eq. 24 the triangle inequality, in Eq. 25 that by data-processing \(P\left(\omega_{A_{1}X_{A}R},\tilde{\omega}_{A_{1}X_{A}R}\right)\leq\varepsilon-\delta\), and in Eq. 26 we applied the choice from Eq. 22 . By Uhlmann’s theorem as in Eq. 21 this concludes the proof.

The upper bound in Eq. ?? follows from the protocol given in [29] based on the convex split lemma (cf. Lem. 12).6 We sketch the argument in our context. Given \(\tilde{\rho}_{AR}\in\mathcal{S}_{\circ}(AR)\) and \(\sigma_A\in\mathcal{S}_{\circ}(A)\) the optimizers in \(I_{\max}^{\varepsilon-\delta,P}(\dot{R};A)_\rho\), the protocol from [29] together with quantum teleportation realises a \(\delta\)-error quantum state merging of a purification \(\tilde{\rho}_{ABR}\) with \(P(\tilde{\rho}_{ABR},\rho_{ABR})\leq\varepsilon-\delta\), for a classical communication cost of \(D_{\max}\left(\tilde{\rho}_{AR}\middle\|\sigma_A\otimes\rho_R\right)+\log\frac{1}{\delta^4}\). Following the same line of argument as in Eq. 23 - 26 based on \(\rho_R=\tilde{\rho}_R\) and \(P(\tilde{\rho}_{ABR},\rho_{ABR})\leq\varepsilon-\delta\) this leads to the desired statement. ◻

The asymptotic first-order expansions then follow from Eq. 12 and Eq. 13 and we recover the original results on quantum state merging [24], [25]. As shown in these references in first-order asymptotically the entanglement cost and classical communication cost can actually be simultaneously minimised — whereas this becomes unclear in the one-shot setting. The asymptotic second-order expansions are an open problem but are now again reduced to giving the asymptotic second-order expansions of \(H_{\min}^{\varepsilon,P}(A|\dot{R})_\rho\) and \(I_{\max}^{\varepsilon,P}(\dot{R};A)_\rho\), respectively.

5 Outlook↩︎

As we have seen our locally smoothed information measures naturally appear in a plethora of operational tasks in quantum information theory. It might be insightful to study mathematical properties of these measures that go beyond what we presented in 2. The main open problem raised by our work is to give asymptotic second-order expansions of the partially smoothed information measures \[\begin{align} H_{\min}^{\varepsilon,\Delta}(A|\dot{B})_\rho\quad\text{and}\quad I_{\max}^{\varepsilon,\Delta}(\dot{A};B)_\rho \end{align}\] for the quantum case. This, however, seems to require new ideas as the classical proof technique from Thm. 1 does not directly translate to the quantum setting and the quantum proof technique from Thm. 2 is not tight enough.

Locally smoothed information measures also appear naturally when defining smooth entropies for quantum channels, as realised in the recent works [31], [32]. We especially point to [31] where our equivalence results from 3 already found applications in the context of quantum channel simulations. Finally, another question our approach might shine some light on is the quantum joint typicality conjecture [33].
Acknowledgements. We thank Kun Fang and Xi Wang for discussions and Joseph M. Renes for pointing out a gap in the proof of a previous version of Theorem 5. MT acknowledges an Australian Research Council Discovery Early Career Researcher Awards, project DE160100821. Part of the work done when R.J. was visiting Tata Institute of Fundamental Research (TIFR), Mumbai, India as a “VAJRA Adjunt Faculty” under the “Visiting Advanced Joint Research (VAJRA) Faculty Scheme” of the Science and Engineering Research Board (SERB), Department of Science and Technology (DST), Government of India. This work is supported by the Singapore Ministry of Education and the National Research Foundation, also through the Tier 3 Grant “Random numbers from quantum processes” MOE2012-T3-1-009.

6 More properties of partially smoothed information measures↩︎

Lemma 6. Let \(\rho_{XB}=\sum_x|x\rangle\langle x|_X\otimes\rho_B^x\in\mathcal{S}_\circ(XB)\), \(\varepsilon\in[0,1]\), and \(f:X\to Z\) be a function. Then, we have for \(\omega_{ZB}:=\sum_x|f(x)\rangle\langle f(x)|_Z\otimes\rho_B^x\in\mathcal{S}_\circ(ZB)\) that \[\begin{align} H_{\min}^{\varepsilon,P}(Z|\dot{B})_\omega\leq H_{\min}^{\varepsilon,P}(X|\dot{B})_\rho\,. \end{align}\] Moreover, when \(B = Y\) is classical then we also have \(H_{\min}^{\varepsilon,T}(Z|\dot{Y})_\omega\leq H_{\min}^{\varepsilon,T}(X|\dot{Y})_\rho\).

Proof. For purified distance the proof follows along similar lines as [1]. We first use of the invariance of the smooth conditional min-entropy under isometries (Lem. 4) to assert that \[\begin{align} H_{\min}^{\varepsilon,P}(X|\dot{B})_\rho = H_{\min}^{\varepsilon,P}(XZ|\dot{B})_\omega \,, \label{eq:smoothed-iso} \end{align}\tag{27}\] where \(\omega_{XZB}:=\sum_x |x\rangle\!\langle x|_X \otimes |f(x)\rangle\!\langle f(x)|_Z\otimes\rho_B^x\). Moreover, by [1], we have \[\begin{align} H_{\min}(XZ|B)_\omega \geq H_{\min}(Z|B)_\omega \,. \label{eq:unsmoothed} \end{align}\tag{28}\] To lift this to smooth entropies, let us assume that \(\tilde{\omega}_{ZB}\) achieves the maximum in the definition of \(H_{\min}^{\varepsilon,P}(Z|\dot{B})_\omega\). Then, by Uhlmann’s theorem (specifically by [1]), there exists a state \(\tilde{\omega}_{XZB}\) that extends \(\tilde{\omega}_{ZB}\) and is \(\varepsilon\)-close to \(\omega_{XZB}\). Moreover, by the monotonicity of the purified distance under measurements we can ensure that \(\omega_{XZB}\) is classical on \(Z\). Finally, Eq. 28 applied to the state \(\omega_{XZB}\) then yields \[\begin{align} H_{\min}^{\varepsilon, P}(XZ|\dot{B})_{\omega} \geq H_{\min}(XZ|B)_{\tilde{\omega}} \geq H_{\min}(Z|B)_{\tilde{\omega}} = H_{\min}^{\varepsilon,P}(Z|\dot{B})_{\omega} \,, \label{eq:concludehere} \end{align}\tag{29}\] concluding the proof of the first statement.

The proof for the classical case and generalized trace distance is adapted from [22]. First, note that 27 holds for generalized trace distance. Given 28 , it thus remains to extend the optimizer \(\tilde{\omega}_{ZY}\) to a suitable \(\tilde{\omega}_{XZY}\) without invoking Uhlmann’s theorem. As in [22], we define \[\begin{align} \tilde{\omega}_{XZY}(x,z,y) := \frac{\omega_{XZY}(x,z,y)}{\omega_{ZY}(z,y)} \, \tilde{\omega}_{ZY}(z,y) = \omega_{X|ZY}(x|z,y) \, \tilde{\omega}_{ZY}(z,y) \,. \end{align}\] Using the definition of the generalized trace distance we then find that \[\begin{align} T(\tilde{\omega}_{XZY}, \omega_{XZY}) &= \sum_{x,y,z:\atop \omega_{XZY}(x,z,y) \geq \tilde{\omega}_{XZY}(x,z,y)} \omega_{XZY}(x,z,y) - \tilde{\omega}_{XZY}(x,z,y) \\ &= \sum_{x,y,z: \atop \omega_{ZY}(z,y) \geq \tilde{\omega}_{ZY}(z,y)} \omega_{X|ZY}(x|z,y) \Big( \omega_{ZY}(z,y) - \tilde{\omega}_{ZY}(z,y) \Big) \\ &= \sum_{y,z: \atop \omega_{ZY}(z,y) \geq \tilde{\omega}_{ZY}(z,y)} \omega_{ZY}(z,y) - \tilde{\omega}_{ZY}(z,y) \\ &= T(\tilde{\omega}_{ZY}, \omega_{ZY}) \,. \end{align}\] The proof concludes with the same argument given in 29 . ◻

We would also like to note here that for generalized trace distance the above lemma does in fact not hold in the quantum case (when asking for identical smoothing parameters for both min-entropies) and a counter-example can be constructed using the states given in [34].

Lemma 7. Let \(\rho_{ABR}\in\mathcal{S}_{\circ}(ABR)\) and \(\varepsilon\in[0,1]\). Then, we have \[\begin{align} H_{\min}^{\varepsilon,P}(AB|\dot{R})_\rho\leq H_{\min}^{\varepsilon,P}(A|\dot{R})_\rho+\log|B|\,. \end{align}\]

Proof. This is implied by [26]. ◻

Lemma 8. Let \(\rho_{ABXX'}\in\mathcal{S}_{\circ}(ABXX')\) be coherently classical on \(XX'\) and \(\varepsilon\in[0,1]\). Then, we have \[\begin{align} I_{\max}^{\varepsilon,P}(BXX';\dot{A})_{\rho}\leq I_{\max}^{\varepsilon,P}(BX';\dot{A})_\rho+\log|X|\,. \end{align}\]

Proof. Let \(\tilde{\rho}_{ABX'}\in\mathcal{S}_{\bullet}(ABX')\) and \(\sigma_{BX'}\in\mathcal{S}_{\circ}(BX')\) be the optimizers in \(I_{\max}^{\varepsilon,P}(BX';\dot{A})_\rho\), where both can assumed to be classical on \(X'\). Now taking an extension \(\tilde{\rho}_{ABXX'}\) of \(\tilde{\rho}_{ABX'}\) with \(P(\tilde{\rho}_{ABXX'},\rho_{ABXX'})\leq\varepsilon\) we get \[\begin{align} \label{eq:apply} \tilde{\rho}_{ABXX'}\leq|X|\cdot1_{X}\otimes\tilde{\rho}_{ABX'}\leq|X|\cdot2^{D_{\max}(\tilde{\rho}_{ABX'}\|\rho_{A}\otimes\sigma_{BX'})}\cdot1_{X}\otimes\rho_{A}\otimes\sigma_{BX'}\,. \end{align}\tag{30}\] Applying \(\Pi_{XX'}:=\sum_{X}|x\rangle\langle x|_{X}\otimes|x\rangle\langle x|_{X'}\) to both sides of Eq. 30 leads to \[\begin{align} \label{eq:apply2} \Pi_{XX'}\tilde{\rho}_{ABXX'}\Pi_{XX'}\leq|X|\cdot2^{D_{\max}(\tilde{\rho}_{ABX'}\|\rho_{A}\otimes\sigma_{BX'})}\cdot\rho_{A}\otimes\Big(\Pi_{XX'}(1_{X}\otimes\sigma_{BX'})\Pi_{XX'}\Big)\,. \end{align}\tag{31}\] For \(\hat{\rho}_{ABXX'}:=\Pi_{XX'}\tilde{\rho}_{ABXX'}\Pi_{XX'}\) and \(\hat{\sigma}_{BXX'}:=\Pi_{XX'}(1_{X}\otimes\sigma_{BX'})\Pi_{XX'}\) this is \[\begin{align} \hat{\rho}_{ABXX'}\leq|X|\cdot2^{D_{\max}(\tilde{\rho}_{ABX'}\|\rho_A\otimes\sigma_{BX'})}\cdot\rho_{A}\otimes\hat{\sigma}_{BXX'}\,, \end{align}\] which in turn implies \[\begin{align} 2^{D_{\max}(\hat{\rho}_{ABXX'}\|\rho_A\otimes\hat{\sigma}_{BXX'})}\leq|X|\cdot2^{D_{\max}(\tilde{\rho}_{ABX'}\|\rho_A\otimes\sigma_{BX'})}\,. \end{align}\] Since \(\Pi_{XX'}\) is trace non-increasing we have \(\hat{\rho}_{ABXX'}\in\mathcal{S}_{\bullet}(ABXX')\) with \(P(\hat{\rho}_{ABXX'},\rho_{ABXX'})\leq\varepsilon\) and together with \(\hat{\sigma}_{BXX'}\in\mathcal{S}_{\bullet}(BXX')\) this finishes the proof. ◻

Lemma 9. Let \(\rho_{AB}\in\mathcal{S}_\circ(AB)\), \(\sigma_B\in\mathcal{S}_{\circ}(B)\), \(\{ P_A^x \}\) be a projective measurement on \(A\), and \(\omega_{ABX}:=\sum_x |x\rangle\langle x|_X\otimes P_A^x\rho_{AB}P_A^x\). Then, we find that \[\begin{align} \label{eq:lemma-first} D_{\max}(\rho_{AB}\|1_A\otimes\sigma_B)\leq D_{\max}(\omega_{ABX}\|1_A\otimes\sigma_B\otimes\omega_X)\,. \end{align}\qquad{(10)}\] Moreover, for \(\varepsilon\in[0,1]\) we have \[\begin{align} \label{eq:lemma-second} H_{\min}^{\varepsilon,P}(A|\dot{B})_\rho\geq\sup_{\tilde{\omega}_{ABX}}-D_{\max}(\tilde{\omega}_{ABX}\|1_A\otimes\rho_B\otimes\tilde{\omega}_X)\,, \end{align}\qquad{(11)}\] where the supremum is over all classical-quantum \(\tilde{\omega}_{ABX}\in\mathcal{S}_{\circ}(ABX)\) with \(P(\tilde{\omega}_{ABX},\omega_{ABX})\leq\varepsilon\).

Proof. Eq. ?? is [26]. For Eq. ?? note that the isometric purification of \(\{P_{A}^{x}\}\), \(V=\sum_{x}|x\rangle_{X}\otimes P_{A}^{x}\), can be inverted on the image of \(V\) \[\begin{align} P_{V}=\sum_{x}P_{A}^{x}\otimes|x\rangle\langle x|_X\,. \end{align}\] Now, for every classical-quantum \(\tilde{\omega}_{ABX}\in\mathcal{S}_{\bullet}(ABX)\) with \(P(\tilde{\omega}_{ABX},\omega_{ABX})\leq\varepsilon\) we have for \(\tilde{\omega}^V_{ARX}:=P_{V}\tilde{\omega}_{ABX}P_V\) that \[\begin{align} D_{\max}(\tilde{\omega}^V_{ABX}\|1_A\otimes\rho_B\otimes\tilde{\omega}_{X}^V)\leq D_{\max}(\tilde{\omega}_{ARX}\|1_A\otimes\rho_B\otimes\tilde{\omega}_{X})\,. \end{align}\] Hence, we can restrict the supremum in Eq. ?? to states in the image of \(V\) and the claim now follows by Eq. ?? together with the invariance under isometries. ◻

7 Assorted additional lemmas↩︎

Lemma 10 (Special case of Lem. A.1 in [28]). Let \(\rho_{AB} \in \mathcal{S}_{\bullet}(AB)\) and \(L: B \to B'\) be a contraction. Then, we have \[\begin{align} \mathop{\mathrm{tr}}_{B'} \left[ ( 1_A \otimes L ) \rho_{AB} ( 1_A \otimes L )^{\dagger} \right] \leq \rho_A\,. \end{align}\]

Lemma 11 (Def. 3.3 & 3.4 in [1]). Let \(P_X, Q_X \in \mathcal{S}_{\bullet}(X)\). Then, for \(\mathcal{X}\) the set associated to \(X\) we have \[\begin{align} T(P_X, Q_X) = \max_{S\in \mathcal{X}}|P_X(S) - Q_X(S)|\,. \end{align}\]

Lemma 12 (Variation of convex-split lemma from [29]). Let \(\varepsilon, \delta\in (0,1)\) and \(\rho_{AB}, \rho'_{AB}\in \mathcal{S}_{\circ}(AB), \sigma_B\in \mathcal{S}_{\circ}(B)\) such that \(\Delta(\rho_{AB},\rho'_{AB})\leq \varepsilon\). Then, for the quantum state \[\begin{align} \tau_{AB_1\ldots B_{2^R}} = \frac{1}{2^R}\sum_{j=1}^{2^R}\rho_{AB_j}\otimes \sigma_{B_1}\otimes \ldots \sigma_{B_{j-1}}\otimes\sigma_{B_{j+1}}\otimes\ldots \sigma_{B_{2^R}}&\\ \text{with R\geq\left\lceil D_{\max}(\rho'_{AB}\|\rho'_A\otimes \sigma_B)+2\log\frac{2}{\delta}\right\rceil}& \end{align}\] we have \(\Delta(\tau_{AB_1\ldots B_{2^R}},\rho'_A\otimes\sigma_{B_1}\otimes\ldots\sigma_{B_{2^R}})\leq\varepsilon+\delta\).

Proof. We only sketch the minor additional steps compared to the proof in [29]. For \[\begin{align} \tau'_{AB_1\ldots B_{2^R}} = \frac{1}{2^R}\sum_{j=1}^{2^R}\rho'_{AB_j}\otimes \sigma_{B_1}\otimes \ldots \sigma_{B_{j-1}}\otimes\sigma_{B_{j+1}}\otimes\ldots \sigma_{B_{2^R}} \end{align}\] we have from [29] that \[P(\tau'_{AB_1\ldots B_{2^R}}, \rho'_A\otimes \sigma_{B_1} \otimes\ldots \sigma_{B_{2^R}})\leq \delta\,,\] which implies by the Fuchs-Van de Graaf inequality that \[T(\tau'_{AB_1\ldots B_{2^R}}, \rho'_A\otimes \sigma_{B_1} \otimes\ldots \sigma_{B_{2^R}})\leq \delta\,.\] Now, by the concavity of fidelity \[F(\tau_{AB_1\ldots B_{2^R}}, \tau'_{AB_1\ldots B_{2^R}})\geq F(\rho_{AB}, \rho'_{AB})\quad\implies\quad P(\tau_{AB_1\ldots B_{2^R}}, \tau'_{AB_1\ldots B_{2^R}})\leq P(\rho_{AB}, \rho'_{AB})\,,\] and by the triangle inequality \(T(\tau_{AB_1\ldots B_{2^R}}, \tau'_{AB_1\ldots B_{2^R}})\leq T(\rho_{AB}, \rho'_{AB})\). The proof now follows by the triangle inequality for either \(P\) or \(T\). ◻

References↩︎

[1]
M. Tomamichel. Quantum Information Processing with Finite Resources — Mathematical Foundations. volume 5 of SpringerBriefs in Mathematical Physics, http://dx.doi.org/10.1007/978-3-319-21891-5(2016).
[2]
R. Renner. Security of Quantum Key Distribution. PhD thesis, ETH Zurich, (2005). Available at http://arxiv.org/abs/quant-ph/0512258.
[3]
M. Berta, M. Christandl, and R. Renner. The Quantum Reverse Shannon Theorem Based on One-Shot Information Theory. http://dx.doi.org/10.1007/s00220-011-1309-7(2011).
[4]
R. Jain, J. Radhakrishnan, and P. Sen. Privacy and Interaction in Quantum Communication Complexity and a Theorem About the Relative Entropy of Quantum States. In Proc. IEEE FOCS 2002, pages 429–438, Vancouver(2002).
[5]
N. Datta. Min- and Max- Relative Entropies and a New Entanglement Monotone. http://dx.doi.org/10.1109/TIT.2009.2018325(2009).
[6]
R. König, R. Renner, and C. Schaffner. The Operational Meaning of Min- and Max-Entropy. http://dx.doi.org/10.1109/TIT.2009.2025545(2009).
[7]
N. Ciganovic, N. J. Beaudry, and R. Renner. Smooth Max-Information as One-Shot Generalization for Mutual Information. http://dx.doi.org/10.1109/TIT.2013.2295314(2014).
[8]
M. Tomamichel, R. Colbeck, and R. Renner. Duality Between Smooth Min- and Max-Entropies. http://dx.doi.org/10.1109/TIT.2010.2054130(2010).
[9]
A. Anshu, R. Jain, and N. A. Warsi. Building blocks for communication over noisy quantum networks. Preprint, http://arxiv.org/abs/1702.01940 .
[10]
M. Tomamichel, M. Berta, and M. Hayashi. Relating Different Quantum Generalizations of the Conditional Rényi Entropy. http://dx.doi.org/10.1063/1.4892761(2014).
[11]
T. S. Han. Information-Spectrum Methods in Information Theory. Springer (2002).
[12]
V. Strassen. Asymptotische Abschätzungen in Shannons Informationstheorie. In Trans. Third Prague Conference on Information Theory, pages 689–723, Prague(1962).
[13]
A. Anshu, R. Jain, and N. A. Warsi. A unified approach to source and message compression. Preprint, http://arxiv.org/abs/1707.03619 .
[14]
A. Uhlmann. The Transition Probability for States of Star-Algebras. Annals of Physics 497(4): 524–532, (1985).
[15]
M. Tomamichel, C. Schaffner, A. Smith, and R. Renner. Leftover Hashing Against Quantum Side Information. http://dx.doi.org/10.1109/TIT.2011.2158473(2011).
[16]
M. Tomamichel, R. Colbeck, and R. Renner. A Fully Quantum Asymptotic Equipartition Property. http://dx.doi.org/10.1109/TIT.2009.2032797(2009).
[17]
M. Braverman and A. Rao. “Information Equals Amortized Communication”. In 2011 IEEE 52nd Annual Symposium on Foundations of Computer Science, pages 748–757, (2011).
[18]
P. Harsha, R. Jain, D. McAllester, and J. Radhakrishnan. “The Communication Complexity of Correlation”. http://dx.doi.org/10.1109/TIT.2009.2034824(2010).
[19]
M. Tomamichel and M. Hayashi. A Hierarchy of Information Quantities for Finite Block Length Analysis of Quantum Tasks. http://dx.doi.org/10.1109/TIT.2013.2276628(2013).
[20]
C. Portmann and R. Renner. Cryptographic Security of Quantum Key Distribution, (2014). http://arxiv.org/abs/1409.3525.
[21]
M. Hayashi. Security Analysis of \(\varepsilon\) -Almost Dual Universal2 Hash Functions: Smoothing of Min Entropy Versus Smoothing of Renyi Entropy of Order 2. http://dx.doi.org/10.1109/TIT.2016.2535174(2016).
[22]
S. Watanabe and M. Hayashi. Non-Asymptotic Analysis of Privacy Amplification via Rényi Entropy and Inf-Spectral Entropy. In Proc. IEEE ISIT 2013, pages 2715–2719, Istanbul, Turkey(2013).
[23]
M. Hayashi and S. Watanabe. Uniform Random Number Generation From Markov Chains: Non-Asymptotic and Asymptotic Analyses. http://dx.doi.org/10.1109/TIT.2016.2530084(2016).
[24]
M. Horodecki, J. Oppenheim, and A. Winter. Partial Quantum Information.. http://dx.doi.org/10.1038/nature03909(2005).
[25]
M. Horodecki, J. Oppenheim, and A. Winter. Quantum State Merging and Negative Information. http://dx.doi.org/10.1007/s00220-006-0118-x(2006).
[26]
M. Berta. Single-Shot Quantum State Merging. Diploma thesis, ETH Zurich, Available at http://arxiv.org/abs/0912.4495.
[27]
F. Dupuis, M. Berta, J. Wullschleger, and R. Renner. One-Shot Decoupling. http://dx.doi.org/10.1007/s00220-014-1990-4(2014).
[28]
M. Tomamichel. A Framework for Non-Asymptotic Quantum Information Theory. PhD thesis, ETH Zurich, (2012). Available at http://arxiv.org/abs/1203.2142.
[29]
A. Anshu, V. K. Devabathini, and R. Jain. “Quantum Communication Using Coherent Rejection Sampling”. http://dx.doi.org/10.1103/PhysRevLett.119.120506(2017).
[30]
C. Majenz, M. Berta, F. Dupuis, R. Renner, and M. Christandl. “Catalytic Decoupling of Quantum Information”. http://dx.doi.org/10.1103/PhysRevLett.118.080503(2017).
[31]
K. Fang, X. Wang, M. Tomamichel, and M. Berta. “Quantum Channel Simulation and the Channel’s Smooth Max-Information”. Preprint, http://arxiv.org/abs/1807.05354(2018).
[32]
P. Faist, M. Berta, and F. Brandao. “Thermodynamic capacity of quantum processes”. Preprint, http://arxiv.org/abs/1807.05610(2018).
[33]
P. Sen. A one-shot quantum joint typicality lemma. Preprint, http://arxiv.org/abs/1806.07278 .
[34]
J. M. Renes. “On privacy amplification, lossy compression, and their duality to channel coding”. Preprint, http://arxiv.org/abs/1708.05685(2017).

  1. The original definition of the smooth max-information in [3] was slightly different and based on \(D_{\max}(\tilde{\rho}_{AB} \| \tilde{\rho}_A \otimes \sigma_B)\).↩︎

  2. A similar locally smoothed quantity has previously made an appearance in [9] as a proof tool.↩︎

  3. To see this, for example for the smooth max-information, note that if this were not so then the full dephasing map (in the classical basis) could be applied to both sides of the operator inequality \[\begin{align} \tilde{\rho}_{AB} \leq \rho_A \otimes \sigma_B \,, \end{align}\] yielding a new feasible solution since the distance between \(\rho_{AB}\) and \(\tilde{\rho}_{AB}\) is also reduced when the dephasing map is applied due to Lem. 3.↩︎

  4. Similar equivalence results for alternative min-entropy definitions based on a maximization over \(\sigma_B \in \mathcal{S}_{\bullet}(B)\) can be derived by additionally employing [15].↩︎

  5. Alternatively this expansion can also directly be deduced from Eq. 14 .↩︎

  6. Analogously to the protocol minimizing the entanglement cost, this protocol can also be thought of in terms of decoupling [30].↩︎