July 16, 2026
The entanglement-assisted capacity of a quantum channel admits an additive single-letter characterization, implying that joint encodings across channel uses cannot increase the ultimate communication rate. Here, we show that this additive picture does not extend to communication reliability. Specifically, we prove that the Petz–Rényi channel information can be strictly superadditive for every \(\alpha\in[\frac{1}{2},1)\), yielding a genuine multi-copy enhancement of the entanglement-assisted random-coding error exponent, even though the entanglement-assisted capacity remains additive. We establish this phenomenon analytically already for measurement channels, which are entanglement-breaking and have additive unassisted capacity. Remarkably, this strict superadditivity is witnessed by a separable, classically correlated two-copy channel-input marginal, demonstrating that no entanglement between the transmitted systems is required. Our results show that, although correlations across channel uses cannot increase the ultimate rate of entanglement-assisted communication, they can enhance its reliability.
gap_data.dat alpha Ione ItwoPer IthreePer Isand gapTwo gapThree gapSand 0.5 0.990734014905 0.990779744273 0.990825179241 1.02243519042 4.57293687331e-05 9.11643362278e-05 0.0316100111742 0.51 0.993208295661 0.993249587424 0.993290642125 1.0233974593 4.12917636826e-05 8.23464641928e-05 0.0301068171766 0.52 0.995639302675 0.995676483741 0.995713475883 1.02436035076 3.71810655441e-05 7.41732076472e-05 0.0286468748771 0.53 0.998027054855 0.998060442382 0.998093681019 1.02532383584 3.33875278539e-05 6.66261640907e-05 0.0272301548208 0.54 1.00037164828 1.0004015476 1.00043133099 1.02628788549 2.98993191969e-05 5.96827137103e-05 0.0258565545032 0.55 1.00267325141 1.0026999544 1.00272656835 1.02725247059 2.67029915859e-05 5.33169374348e-05 0.0245259022344 0.56 1.0049321002 1.00495588411 1.00497960066 1.02821756189 2.37839077439e-05 4.7500458787e-05 0.0232379612261 0.57 1.00714849306 1.00716961968 1.00719069626 1.0291831301 2.11266231782e-05 4.22032008316e-05 0.0219924338427 0.58 1.00932278582 1.00934150104 1.00936017987 1.03014914583 1.8715221553e-05 3.73940542888e-05 0.0207889659587 0.59 1.0114553868 1.0114719204 1.01148842826 1.03111557963 1.65336034683e-05 3.30414559395e-05 0.0196271513728 0.6 1.01354675186 1.01356131759 1.01357586574 1.03208240197 1.45657304504e-05 2.91138799071e-05 0.0185065362289 0.61 1.01559737962 1.01561017544 1.01562295986 1.03304958326 1.27958260427e-05 2.55802458899e-05 0.0174266234007 0.62 1.01760780683 1.01761901537 1.01763021708 1.03401709387 1.12085384516e-05 2.24102504389e-05 0.0163868767897 0.63 1.01957860393 1.019588393 1.01959817856 1.03498490408 9.78906733473e-06 1.95746281406e-05 0.0153867255238 0.64 1.0215103708 1.02151889406 1.02152741615 1.03595298416 8.52325959344e-06 1.70453504638e-05 0.0144255680074 0.65 1.02340373273 1.02341113041 1.0234185285 1.03692130431 7.39767717528e-06 1.47957691352e-05 0.0135027758049 0.66 1.02525933664 1.02526573628 1.02527213735 1.03788983469 6.39964251548e-06 1.28007132536e-05 0.0126176973436 0.67 1.02707784749 1.02708336476 1.02708888404 1.03885854546 5.51726279085e-06 1.10365449828e-05 0.0117696614179 0.68 1.02885994503 1.02886468447 1.02886942621 1.0398274067 4.73943873058e-06 9.48118263744e-06 0.0109579804888 0.69 1.03060632065 1.03061037651 1.03061443475 1.04079638851 4.05585967567e-06 8.11409615342e-06 0.010181953762 0.7 1.0323176746 1.03232113159 1.03232459088 1.04176546094 3.45698812021e-06 6.91628068616e-06 0.00944087005979 0.71 1.03399471337 1.0339976474 1.03400058358 1.04273459405 2.93403599749e-06 5.87021315246e-06 0.00873401047092 0.72 1.03563814728 1.03564062622 1.03564310708 1.04370375787 2.47893482852e-06 4.95979597059e-06 0.00806065078948 0.73 1.03724868839 1.03725077269 1.03725285868 1.04467292243 2.0843014763e-06 4.17029121258e-06 0.0074200637499 0.74 1.03882704847 1.03882879187 1.03883053672 1.04564205777 1.74340097159e-06 3.48824883467e-06 0.00681152105581 0.75 1.04037393728 1.04037538739 1.04037683871 1.04661113393 1.45010763752e-06 2.90143080117e-06 0.00623429522495 0.76 1.04189006099 1.04189125985 1.04189245972 1.04758012096 1.19886541761e-06 2.39873341079e-06 0.00568766123913 0.77 1.04337612077 1.04337710542 1.04337809088 1.04854898891 9.84648710345e-07 1.97010971093e-06 0.00517089803012 0.78 1.04483281159 1.04483361451 1.04483441808 1.04951770788 8.02922826448e-07 1.60649254344e-06 0.00468328980053 0.79 1.04626082107 1.04626147067 1.04626212079 1.05048624797 6.4960700441e-07 1.29972026652e-06 0.00422412718574 0.8 1.04766082858 1.04766134962 1.04766187104 1.05145457932 5.21037848866e-07 1.04246494481e-06 0.00379270827874 0.81 1.04903350442 1.04903391835 1.04903433258 1.05242267211 4.13935029586e-07 8.28163829292e-07 0.00338833952305 0.82 1.0503795091 1.05037983447 1.05038016006 1.05339049653 3.25368735066e-07 6.50954836035e-07 0.00301033647454 0.83 1.05169949278 1.05169974551 1.0516999984 1.05435802286 2.52729146943e-07 5.05615749047e-07 0.00265802446127 0.84 1.05299409475 1.05299428845 1.05299448226 1.05532522139 1.93697982542e-07 3.87507347321e-07 0.00233073913335 0.85 1.05426394305 1.05426408927 1.05426423557 1.05629206249 1.46221668107e-07 2.92520413225e-07 0.00202782692069 0.86 1.05550965413 1.05550976262 1.05550987116 1.05725851656 1.08488151573e-07 2.17028176763e-07 0.00174864540181 0.87 1.05673183267 1.05673191157 1.05673199051 1.05822455411 7.89028939963e-08 1.57839728621e-07 0.00149256359761 0.88 1.05793107131 1.05793112737 1.05793118347 1.05919014566 5.60693582674e-08 1.12159791454e-07 0.00125896219307 0.89 1.0591079506 1.05910798936 1.05910802815 1.06015526184 3.87691043713e-08 7.75508419704e-08 0.00104723369697 0.9 1.06026303893 1.06026306487 1.06026309082 1.06111987336 2.59456254259e-08 5.18988336751e-08 0.00085678253755 0.91 1.0613968925 1.06139690919 1.06139692588 1.062083951 1.66889913e-08 3.3381800213e-08 0.000687025116191 0.92 1.06251005536 1.06251006558 1.0625100758 1.06304746561 1.02202251107e-08 2.04423145078e-08 0.00053738980999 0.93 1.06360305948 1.06360306536 1.06360307124 1.06401038817 5.87994808399e-09 1.17607728001e-08 0.000407316930352 0.94 1.06467642483 1.06467642795 1.06467643107 1.06497268972 3.11677528231e-09 6.2340697049e-09 0.000296258653533 0.95 1.06573065956 1.06573066103 1.06573066251 1.06593434143 1.47710155218e-09 2.95421798135e-09 0.000203678914141 0.96 1.06676626009 1.06676626068 1.06676626128 1.06689531454 5.95251181679e-10 1.18980936214e-09 0.000129053266611 0.97 1.06778371134 1.06778371153 1.06778371171 1.06785558044 1.86533677393e-10 3.71660258125e-10 7.18687296166e-05 0.98 1.06878348693 1.06878348697 1.06878348701 1.06881511061 3.59927643245e-11 7.19984072362e-11 3.16236054625e-05 0.99 1.069766049362182517854280510466977718847 1.069766049364398699547648213350562703672 1.069766049366614889148361910425024304336 1.06977387665 2.21618169336770288358498482533e-12 4.43237129408139995804658548925e-12 7.8272784183e-06 0.995 1.070251017196577334763029653711077500931 1.070251017196714831114543795957311373188 1.070251017196852327590227583605154771527 1.07025296427 1.37496351514142246233872257021e-13 2.74992827197929894077270595739e-13 1.94707628043e-06 0.999 1.070636011796563317019652820570118250675 1.070636011796563535747043041905893957932 1.070636011796563754474441242174124996043 1.07063608937 2.18727390221335775707256886822e-16 4.37454788421604006745367665702e-16 7.75742490244e-08
The fundamental limits of reliable information transmission over a noisy point-to-point channel are quantified by its channel capacity. For a classical channel \(\mathscr{W}\), the capacity is characterized by the mutual information of channel \(I_1(\mathscr{W})\) [1]. Moreover, Shannon proved that the capacity is additive under independent uses of channels, i.e. \(I_1(\mathscr{W}_1 \otimes \mathscr{W}_2) = I_1(\mathscr{W}_1) + I_1(\mathscr{W}_2)\), reducing the evaluation of this fundamental quantity to a computable single-letter formula, cast as a fixed-dimensional convex optimization. Operationally, the additivity of \(I_1(\mathscr{W})\) means that encoding classical data jointly across both channels yields no advantage for the maximum achievable rate.
Quantum mechanics, however, fundamentally departs from this paradigm. For classical communication over a quantum channel \(\mathscr{N}\), the Holevo information of a channel \(\chi(\mathscr N)\) [2], [3] need not be additive: there exist quantum channels \(\mathscr{N}_1\) and \(\mathscr{N}_2\) such that \(\chi(\mathscr{N}_1\otimes\mathscr{N}_2) > \chi(\mathscr{N}_1) + \chi(\mathscr{N}_2)\) [4], [5]. Consequently, ensembles containing entangled states across independent channel uses can outperform product-state encodings, and hence the classical capacity generally requires regularization over arbitrarily many channel uses. An even more striking nonadditivity arises in quantum communication [6]–[8]: two channels with individually vanishing quantum capacities can have positive quantum capacity when used jointly, a phenomenon known as superactivation [9]. More generally, no fixed finite block length suffices to determine the quantum capacity of all channels [10]. These results establish correlations across channel uses as an operational resource, while explaining why quantum channel capacities often resist single-letter characterization. For unassisted communication, however, this nonadditivity disappears for entanglement-breaking channels: their Holevo information is strongly additive, and hence their classical capacity is single-lettered [11].
To restore the elegant phenomenon of additivity for general quantum channels, preshared entanglement between the sender and receiver emerges as an operational resolution [12]–[14]. Bennett et al. showed that allowing the sender and receiver unlimited preshared entanglement reduces the classical capacity of a quantum channel, \(C_{\mathrm{EA}}(\mathscr{N})\), to an additive, single-letter optimization of the quantum mutual information. The corresponding entanglement-assisted quantum capacity satisfies \(Q_{\mathrm{EA}}(\mathscr{N}) = \frac{1}{2} C_{\mathrm{EA}}(\mathscr{N})\) by teleportation [15] and superdense coding [16]. Hence, the regularization and superadditivity that obstruct unassisted capacity formulas disappear in the entanglement-assisted setting, yielding a natural quantum analogue of Shannon’s coding theorem. More broadly, this result reveals that the nonadditive complexity of quantum communication is not an immutable property of the channel alone, but depends fundamentally on which correlations are available as operational resources.
While channel capacity establishes the ultimate quantity of reliably transmissible information, the operational performance of practical physical systems is equally governed by the quality of that communication. This quality is characterized by the error exponent for rates below capacity and the strong converse exponent for rates above capacity. Together, these exponents dictate the exponential decay of the decoding error probability and success probability, respectively, at any fixed transmission rate. For transmission rates above capacity, the strong converse exponent [17], [18] is determined by a simple formula in terms of the channel’s sandwiched Rényi information \(\widetilde{I}_{\alpha}(\mathscr{N})\) of order \(\alpha>1\). Because this sandwiched quantity is additive in this regime [17], the entanglement-assisted strong converse exponent circumvents intractable asymptotic limits, yielding a computable single-letter formula. Furthermore, the recent proof extending this result to \(\alpha \in [\frac{1}{2}, 1)\) [19], together with the mutual-information point \(\alpha=1\), demonstrates additivity of the sandwiched–Rényi channel information throughout its full data-processing range \(\alpha\geq\frac{1}{2}\). Taken together, these results show that correlations across multiple channel uses provide no advantage for the entanglement-assisted capacity or the above-capacity strong-converse exponent. Recent progress establishes an exponential decay rate for the decoding error probability, expressed in terms of the channel’s Petz–Rényi information \(I_{\alpha}(\mathscr{N})\) of order \(\alpha\in[\frac{1}{2},1)\) [20]. Given that entanglement assistance resolves the nonadditivity of channel capacity, prior results on the additivity of \(I_{1}(\mathscr{N})\) and \(\widetilde{I}_{\alpha}(\mathscr{N})\) for \(\alpha\geq\frac{1}{2}\) naturally suggest that \(I_{\alpha}(\mathscr{N})\), and thereby the quality of entanglement-assisted communication, might exhibit the same well-behaved additive structure when coding with rates at capacity and above.
However, when turning to the practically more relevant regime of transmission rates below capacity, we prove that the Petz–Rényi channel information exhibits strict superadditivity for every \(\alpha\in[1/2,1)\), yielding a genuine multi-copy improvement in communication reliability. Interestingly, this failure of additivity already occurs for an entanglement-breaking measurement channel, a class for which the Holevo information—and hence the unassisted classical capacity—is additive [11]. Hence, even a channel that outputs only classical data and whose output is necessarily separable from any retained reference can exhibit a collective entanglement-assisted advantage in its Petz–Rényi information; see Figure 1.
A striking feature of our result is that the strict superadditivity established can even be witnessed without entanglement between the transmitted systems. In sharp contrast to many celebrated literature whereas quantum nonadditivity is often regarded as an intrinsically entanglement-driven phenomenon [4], [9], [10], [21], [22] 1, the witness can be chosen to have a separable, indeed classically correlated, two-copy channel marginal. This demonstrates that entanglement across the transmitted inputs is not an essential ingredient of the superadditive advantage.
Let \(\mathscr{N}_{\mathsf{A}\to\mathsf{B}}\) be a quantum channel from Alice’s system \(\mathsf{A}\) to Bob’s system \(\mathsf{B}\). In entanglement-assisted (EA) communication, an entangled state \(\theta_{\underline{\mathsf{A}}\underline{\mathsf{R}}}\) is shared between Alice (holding \(\underline{\mathsf{A}}\)) and Bob (holding \(\underline{\mathsf{R}}\)). To send a message \(m \in \{1,2,\ldots, \lfloor 2^{nR} \rfloor\}\) of rate \(R\), Alice applies an encoding operation \(\mathscr{E}_{\underline{\mathsf{A}}\to\mathsf{A}^n}^m\) on her part of shared entanglement to prepare a length-\(n\) quantum codeword on system \(\mathsf{A}^n \equiv \mathsf{A}_1 \ldots \mathsf{A}_n\). The quantum codeword then undergoes the \(n\)-fold product channel \(\mathscr{N}_{\mathsf{A}\to\mathsf{B}}^{\otimes n}\). At receiver, Bob applies a quantum measurement \(\big\{M_{\underline{\mathsf{R}}\mathsf{B}^n}^{m}\big\}_m\) on the noisy quantum system \(\mathsf{B}^n \equiv \mathsf{B}_1 \ldots \mathsf{B}_n\) and his part of shared entanglement. The decoding error probability is \[\begin{align} \!\operatorname{Err}\! =\!1- \frac{1}{\lfloor 2^{nR} \rfloor} \!\sum_{m} \operatorname{Tr}\!\left[ \mathscr{N}_{\mathsf{A}\to\mathsf{B}}^{\otimes n} \circ \mathscr{E}_{\underline{\mathsf{A}}\to \mathsf{A}^n}^m\!\left(\theta_{\underline{\mathsf{A}}\underline{\mathsf{R}}}\right) {M}_{\underline{\mathsf{R}}\mathsf{B}^n}^m \right]\!. \end{align}\] We call such an encoder \(\{\mathscr{E}_{\underline{\mathsf{A}}\to \mathsf{A}^n}^m\}_m\), decoder \(\big\{M_{\underline{\mathsf{R}}\mathsf{B}^n}^{m}\big\}_m\), and shared entanglement \(\theta_{\underline{\mathsf{A}}\underline{\mathsf{R}}}\) an \((n,R,\varepsilon)\) code if \(\operatorname{Err}\leq \varepsilon\). The maximum achievable rate over all EA-codes is \[\begin{align} R^{\star}(n,\varepsilon) \equiv \sup\left\{R: \exists\, (n,R,\varepsilon)\text{ code}\right\}, \end{align}\] which is the ultimate quantity of information bits Alice can send to Bob with an error tolerance \(\varepsilon\). The well-known BSST theorem [12]–[14] characterizes the channel capacity \[\begin{align} \label{eq:first-order} C_{\mathrm{EA}}(\mathscr{N}) \equiv \lim_{\varepsilon\to0}\lim_{n\to\infty} R^{\star}(n,\varepsilon) = I_1(\mathscr{N}), \end{align}\tag{1}\] in terms of the single-letter mutual information of channel \(\mathscr{N}\): \[\begin{align} \label{eq:I951} I_1(\mathscr{N}) = \max_{\rho_{\mathsf{A}}} I_1(\mathsf{R}:\mathsf{B})_{\omega}, \quad \omega_{\mathsf{R}\mathsf{B}} = (\operatorname{id}_{\mathsf{R}}\otimes\mathscr N)(\psi_{\mathsf{R}\mathsf{A}}^{\rho}), \end{align}\tag{2}\] where \(\psi_{\mathsf{R}\mathsf{A}}^{\rho}\) is a purification of the input state \(\rho_{\mathsf{A}}\) and \(\mathsf{R}\cong\mathsf{A}\). The quantity \(I_1(\mathscr{N})\) is strong additive under tensor product of channels [13]: \[\begin{align} \label{eq:I951-aditive} {I}_{1}(\mathscr{N}_1 \otimes \mathscr{N}_2) = {I}_{1}(\mathscr{N}_1) + {I}_{1}(\mathscr{N}_2). \end{align}\tag{3}\] The channel capacity \(C_{\mathrm{EA}}(\mathscr{N})\) is strongly additive as well.
On the other hand, for a given rate \(R\), the minimum error probability is defined as \[\begin{align} \varepsilon^{\star}(n,R) \equiv \inf\left\{\varepsilon: \exists\,(n,R,\varepsilon)\text{ code}\right\}, \end{align}\] determining the ultimate quality of EA-communication. For transmission rates exceeding capacity, \(R > C_{\mathrm{EA}}(\mathscr{N})\), Gupta and Wilde [17] and Li and Yao [18] demonstrated that the success probability decays exponentially as \[\begin{align} 1- \varepsilon^{\star}(n,R) &\simeq 2^{-n E_{\mathrm{sc}}(R;\mathscr{N})}, \tag{4} \\ E_{\mathrm{sc}}(R;\mathscr{N}) &\equiv\sup_{\alpha>1} \frac{\alpha-1}{\alpha} \left[ R - \widetilde{I}_{\alpha}(\mathscr{N}) \right]. \tag{5} \end{align}\] Here, the strong converse exponent is determined by the difference between the rate \(R\) and the sandwiched Rényi information of \(\mathscr{N}\) for \(\alpha >1\), \[\begin{align} \label{eq:defn:sandwiched95info} \widetilde{I}_{\alpha}(\mathscr{N}) \equiv \max_{\rho_{\mathsf{A}}} \min_{\sigma_{\mathsf{B}}} \widetilde{D}_{\alpha} \left( \mathscr{N}(\psi_{\mathsf{R}\mathsf{A}}^{\rho}) \Vert \rho_{\mathsf{R}} \otimes \sigma_{\mathsf{B}} \right), \end{align}\tag{6}\] where \(\widetilde{D}_{\alpha}\) is the sandwiched Rényi relative entropy [23], [24]. The sandwiched Rényi information generalizes \(I_1(\mathscr{N})\) to a parametric family and coincides \(I_1(\mathscr{N})\) when \(\alpha \to 1\). Moreover, Gupta and Wilde [17] proved its additivity: \[\begin{align} \label{eq:sadwiched-aditive95621} \widetilde{I}_{\alpha}(\mathscr{N}_1 \otimes \mathscr{N}_2) = \widetilde{I}_{\alpha}(\mathscr{N}_1) + \widetilde{I}_{\alpha}(\mathscr{N}_2), \quad \forall\,\alpha>1, \end{align}\tag{7}\] an underlying property that guarantees a single-letter formula for the strong converse bound and establishes its weak additivity for rates above capacity, i.e., \[\begin{align} \label{eq:E95sc-additive} E_{\mathrm{sc}}(2R;\mathscr{N}^{\otimes 2}) = 2 E_{\mathrm{sc}}(R;\mathscr{N}). \end{align}\tag{8}\]
For transmission rates below capacity, the convergence rate of the decoding error probability dictates the quality of reliable entanglement-assisted communication, making it the central operational quantity of interest. Recently, it was established [20] (see prior findings in [25], [26]) that \[\begin{align} \varepsilon^{\star}(n,R) &\leq 1.5 \cdot 2^{-E_{\mathrm{r}}(nR;\mathscr{N}^{\otimes n}) }, \tag{9} \\ E_{\mathrm{r}}(R;\mathscr{N}) &\equiv \max_{\frac{1}{2}\leq \alpha < 1} \frac{1-\alpha}{\alpha} \left[ I_{\alpha}(\mathscr{N}) - R \right], \tag{10} \end{align}\] Fundamentally, it is generally anticipated that the exponent \(\frac{1}{n}E_{\mathrm{r}}(nR;\mathscr{N}^{\otimes n})\) to be not merely an achievable bound, but exactly tight, at any transmission rate \(R\) above the so-called critical rate \(R_{\mathrm{crit}}\), completely characterizing the asymptotic error decay. This expectation is firmly grounded in its rigorously proven asymptotic tightness for classical channels [27], classical-quantum channels [28], [29], and covariant quantum channels [30].
Unlike 5 , the random coding exponent \(E_{\mathrm{r}}(R;\mathscr{N})\) is instead characterized by the Petz–Rényi information of order \(\alpha \in [\frac{1}{2},1]\): \[\begin{align} {I}_{\alpha}(\mathscr{N}) &\equiv \max_{\rho_{\mathsf{A}}} \min_{\sigma_{\mathsf{B}}} {D}_{\alpha} \left( \mathscr{N}(\psi_{\mathsf{R}\mathsf{A}}^{\rho}) \Vert \rho_{\mathsf{R}} \otimes \sigma_{\mathsf{B}} \right) \\ &\!\!\overset{\text{\cite{HT14}}}{=} \max_{\rho_{\mathsf{A}}} \frac{\alpha}{\alpha-1} \log \operatorname{Tr}\left[ \left( \operatorname{Tr}_{\mathsf{R}}[\rho_{\mathsf{R}}^{1-\alpha} \omega_{\mathsf{R}\mathsf{B}}^{\alpha} ] \right)^{1/\alpha} \right], \label{eq:objective} \end{align}\tag{11}\] where \(D_{\alpha}\) is the Petz–Rényi relative entropy [31]. Both \(\widetilde{I}_{\alpha}(\mathscr{N})\) and \(I_{\alpha}(\mathscr{N})\) are non-decreasing in \(\alpha\) and converge to \(I_1(\mathscr{N})\) as \(\alpha\to1\) [32]. The associated exponent functions are depicted in Figure 2.
Given that joint entanglement across multiple channel uses do not increase the channel capacity, i.e., the additivity of \(I_1(\mathscr{N})\) in 3 , and recalling the additivity of \(\widetilde{I}_{\alpha}(\mathscr{N})\) in 7 for \(\alpha>1\) and for \(\alpha\in[\frac{1}{2},1)\) by Li–Xu [19], it is natural to expect the following question: \[\begin{align}\label{eq:Q:supperadditivity} \textit{Is Petz--{}R\'{e}nyi{ }I_{\alpha}(\mathscr{N}), \alpha \in [\frac{1}{2},1), additive too?} \end{align}\tag{12}\] Operationally, this corresponds to the question: \[\begin{align}\tag{13} \begin{figure}\includegraphics[width=0.8\textwidth]{_pdflatex/rtvopknu.png}\tag{14}\end{figure} \end{align}\]
Indeed, we show that \(I_{\alpha}(\mathscr{N})\) is equivalent to a convex optimization, as opposed to the non-convex optimizations of the (unassisted) Holevo capacity \(\chi(\mathscr{N})\) and quantum channel capacity \(Q(\mathscr{N})\).
Proposition 1 (Convex optimization reduction). \(I_{\alpha}(\mathscr{N})\) given in 11 is equivalent to a convex optimization.
We defer the detailed proof to Appendix 9. Moreover, additivity of \(I_{\alpha}(\mathscr{N})\) does indeed hold for some channels.
Proposition 2 (Additivity for special channels). \(I_{\alpha}(\mathscr{N})\), \(\alpha \in (0,1)\), is strongly additive for
isometric channels,
projective measurement channels,
covariant quantum channels,
and classical-quantum channels.
We defer the detailed proof to Appendix 9.
In quantum information, a favorable convex optimization landscape typically ensures that fundamental quantities remain computable via single-letter formulas. Unexpectedly, the Petz–Rényi information \(I_{\alpha}(\mathscr{N})\) possesses fundamentally distinct behaviors from the sandwiched Rényi information \(\widetilde{I}_{\alpha}(\mathscr{N})\). Indeed, we prove that \(I_{\alpha}(\mathscr{N})\) can exhibit strict superadditivity for any \(\alpha \in [\frac{1}{2},1)\), which falsifies 12 in general. This finding answers 13 in the affirmative—although joint preshared entanglement does not increase the ultimate quantity of EA-communication, it in general enhance the communication quality for certain channels. Namely, as opposed to the additivity of the strong converse exponent in 8 , \[\begin{align} \label{eq:E95r-supperadditive} \exists\, \mathscr{N}: E_{\mathrm{r}}(2R;\mathscr{N}^{\otimes 2}) > 2 E_{\mathrm{r}}(R;\mathscr{N}), \quad \forall\, R<C_{\mathrm{EA}}(\mathscr{N}). \end{align}\tag{15}\]
Entanglement-breaking. Our first example is the single-heavy Fourier measurement \(\mathscr{M}_{\mathsf{A}\to \mathsf{Y}}\) with rank-one effects \[\begin{align} \label{eq:Fourier95POVM0} M_{\mathsf{A}}^y&=\frac{1}{\lambda d_{\mathsf{A}}} |v_y\rangle\!\langle v_y|, \qquad y=0,\ldots,d_{\mathsf{A}}-1, \end{align}\tag{16}\] where \(\lambda\in(1/d_{\mathsf{A}},1)\), \(a_0=\lambda\), \(a_1=\cdots=a_{d_{\mathsf{A}}-1}=\frac{1-\lambda}{d_{\mathsf{A}}-1}\), \(|v_y\rangle = \sum_{j=0}^{d_{\mathsf{A}}-1}\sqrt{a_j}\,\mathrm{e}^{2\pi \mathrm{i}jy/d_{\mathsf{A}}}|j\rangle\), and the residual effect \(M_{\mathsf{A}}^{\infty} = \mathbf{1}_{\mathsf{A}}-\sum_{j=0}^{d_{\mathsf{A}}-1} M_{\mathsf{A}}^j\geq 0_{\mathsf{A}}\).
Theorem 3 (Strict superadditivity for measurement channels). For every \(d_{\mathsf{A}}\ge3\), every \(\lambda\in(1/d_{\mathsf{A}},1)\), and every \(0<\alpha<1\), the Fourier measurement defined in 16 satisfies \[\begin{align} I_{\alpha}(\mathscr{M}^{\otimes 2}) > 2 I_{\alpha}(\mathscr{M}) \qquad \forall\,0<\alpha<1. \end{align}\]
We defer the detailed proof to Appendix 10.
A numerical certificated superadditivity is in Figure 3.
Our second example is the amplitude damping channels with Choi matrix: \[\begin{align} \label{eq:GAD} \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}}=\begin{pmatrix} 1&0&0&\sqrt{1-\gamma}\\ 0&0&0&0\\ 0&0&\gamma&0\\ \sqrt{1-\gamma}&0&0&1-\gamma \end{pmatrix}. \end{align}\tag{17}\]
Theorem 4 (Strict superadditivity for amplitude damping channels). For the amplitude damping channel \(\mathscr{N}\) given in 17 with \(0<\gamma<1\), one has \[\begin{align} I_\alpha(\mathscr{N}^{\otimes 2}) > 2 I_{\alpha}(\mathscr{N}) \qquad \forall\,0<\alpha<1. \end{align}\]
We defer the detailed proof to Appendix 11.
To analytically prove the strict superadditivity in the above two examples, we first employ Proposition 1 to show that the one-copy Petz–Rényi reduces to a single-parameter convex optimization and the one-copy optimizer \(\rho_{\mathsf{A}}^{\star}\) can be chosen as diagonal. Denoting the trace functional in 11 by \(Q_{\alpha}(\rho) \equiv \operatorname{Tr}\big[ \big( \operatorname{Tr}_{\mathsf{R}}[\rho_{\mathsf{R}}^{1-\alpha} \omega_{\mathsf{R}\mathsf{B}}^{\alpha} ] \big)^{1/\alpha} \big]\), we choose a correlated diagonal two-copy ansatz as \[\begin{align} \label{eq:pi-kappa0} \rho_{\mathsf{A}_1\mathsf{A}_2}(\kappa) = \rho_{\mathsf{A}_1}^{\star} \otimes \rho_{\mathsf{A}_2}^{\star} + \kappa \Delta, \end{align}\tag{18}\] where \(\Delta\) is a diagonal traceless operator and \(\kappa\) is sufficiently small. For any \(\alpha \in (0,1)\), we prove that \[\begin{align} c_{1}(\alpha) \equiv \left.\frac{\mathrm{d}}{\mathrm{d}\kappa} Q_{\alpha}\left( \rho_{\mathsf{A}_1\mathsf{A}_2}(\kappa) \right) \right|_{\kappa=0} < 0. \end{align}\] By Taylor’s expansion: \[\begin{align} \label{eq:pi-kappa0-minimizing} Q_{\alpha}\left(\rho_{\mathsf{A}_1\mathsf{A}_2}(\kappa)\right) = Q_{\alpha}\left(\rho_{\mathsf{A}}^\star\right)^2 + c_{1}(\alpha) \kappa + \mathcal{O}(\kappa^2), \end{align}\tag{19}\] the strict superadditivity is then witnessed by the infinitesimal diagonal classically correlated path \(\rho_\alpha(\kappa)\).
Since \(I_{\alpha}\left(\mathscr{N}\right)\) can be strictly superadditive, the largest Petz channel information in the asymptotic limit is expressed by its regularization: \[\begin{align} I_{\alpha}^{\infty}(\mathscr{N}) &\coloneq \sup_{n\in\mathbb{N}} \frac{1}{n} I_{\alpha}\left(\mathscr{N}^{\otimes n}\right) = \lim_{n\to\infty} \frac{1}{n} I_{\alpha}\left(\mathscr{N}^{\otimes n}\right). \end{align}\] While this regularization captures the optimal communication quality, it unfortunately demands an intractable, infinite-dimensional optimization. To determine the ultimate limits of this multi-copy advantage, we bypass this incomputability by deriving a computable single-letter upper bound. See Figure 4 for the numerical results.
Proposition 5 (Single-letter bounds). For any quantum channel \(\mathscr{N}_{\mathsf{A}\to\mathsf{B}}\) and \(\alpha \in (0,1)\), \[\begin{align} I_{\alpha}(\mathscr{N}) \leq I_{\alpha}^{\infty}(\mathscr{N}) \leq \widetilde{I}_{\frac{1}{2-\alpha}}(\mathscr{N}) \leq {I}_{\frac{1}{2-\alpha}}(\mathscr{N}), \end{align}\] where the upper bounds both converge to \(I_{1}(\mathscr{N})\) as \(\alpha \to 1\).
We defer the detailed proof to Appendix 12.
The single-letter upper bound in terms of \(\widetilde{I}_{\frac{1}{2-\alpha}}(\mathscr{N})\) is a double-state optimization. One may further relax it to the Petz version. We show in Appendix 12 that the sandwiched Rényi information still admits a one-state optimization expression for quantum-classical channels.
In this paper, we analytically prove the strict superadditivity of the Petz–Rényi information \(I_{\alpha}(\mathscr{N})\) of order \(\alpha \in [\frac{1}{2},1)\) for a family of quantum-classical channels and the amplitude damping channels. This implies that joint preshared entanglement can increase the random coding exponent for transmission rates below capacity. Our proof extends to \(\alpha \in(0,1)\cup(1,2)\) as well, and the strict superadditivity vanishes at \(\alpha = 1\) [13] and \(\alpha = 2\). See Table ¿tbl:tab:renyi95additivity? for the summary.
One may wonder if a stronger resource for assisting communication would ease the strict superadditivity of \(I_{\alpha}(\mathscr{N})\). In fact, Ref. [33] considers non-signaling assistance with one-bit forward activation, and the resulting error exponent is given by the same \(I_{\alpha}(\mathscr{N})\) of order \(\alpha \in (0,1)\). Our results then imply that joint assisting resource across multiple channel uses can still enhance the error exponent. While Girardi et al. demonstrated that the zero-rate error exponent admits a regularized expression [34], our investigation targets the constant-rate regime near capacity—the operational domain most crucial for evaluating practical, high-throughput communication.
Finally, numerical evidence from our single-heavy Fourier measurement and amplitude-damping channel also demonstrate strict superadditivity for the Petz–Rényi channel entropy, an additivity question posed by Gour and Wilde [35]. Our examples then show that product strategies for channel discrimination in [32] are suboptimal in general.
max width=
We consider finite-dimensional Hilbert space. We denote by \(\rho_{\mathsf{A}}\) and \(\sigma_{\mathsf{B}}\) quantum states (i.e. density matrix) on quantum systems \(\mathsf{A}\) and \(\mathsf{B}\), respectively. The optimizations \(\max_{\rho_{\mathsf{A}}}\) or \(\min_{\sigma_{\mathsf{B}}}\) are over the state space on systems \(\mathsf{A}\) or \(\mathsf{B}\). We drop the subscript \(\mathsf{A}\) and \(\mathsf{B}\) if the name of the quantum systems are irrelevant. We use \(\mathsf{R}\) to stand for the reference system of \(\mathsf{A}\); hence, \(\mathsf{R}\cong \mathsf{A}\). For \(p\geq 1\), we define the Schatten norm \(\|X\|_p \equiv (\operatorname{Tr}[|X|^p])^{1/p}\).
A quantum channel \(\mathscr{N}_{\mathsf{A}\to\mathsf{B}}\) is a completely positive and trace-preserving map from system \(\mathsf{A}\) to system \(\mathsf{B}\). We denote its unnormalized Choi operator by \[\begin{align} \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}} &\equiv \mathrm{id} \otimes \mathscr{N}_{\bar{\mathsf{A}}\to\mathsf{B}}(|\tilde{\Phi}\rangle\langle\tilde{\Phi}|_{\mathsf{A}\bar{\mathsf{A}}}), \\ |\tilde{\Phi}\rangle_{\mathsf{A}\bar{\mathsf{A}}} &\equiv \sum_i |i\rangle_{\mathsf{A}} |i\rangle_{\bar{\mathsf{A}}}, \quad \bar{\mathsf{A}} \cong \mathsf{A}. \label{eq:unnormalized95MES} \end{align}\tag{20}\] The joint input-output state given an input \(\rho_{\mathsf{A}}\) is denoted by \[\begin{align} \omega_{\mathsf{R}\mathsf{B}}^{\rho} \equiv \sqrt{\rho_{\mathsf{R}}} \Gamma_{\mathsf{R}\mathsf{B}}^{\mathscr{N}} \sqrt{\rho_{\mathsf{R}}}. \end{align}\] The superscript \(\rho\) will be dropped if the context is clear. We use \[\begin{align} \mathscr{M}_{\mathsf{A}\to\mathsf{Y}}(\rho_{\mathsf{A}}) = \sum_{y\in\mathsf{Y}} \operatorname{Tr}\left[\rho_{\mathsf{A}} M_{\mathsf{A}}^y\right] |y\rangle\langle y|_{\mathsf{Y}} \end{align}\] for a quantum-classical (measurement) channel, described by the associated positive operator-valued measure (POVM) \(\{M_{\mathsf{A}}^y\}_{y\in\mathsf{Y}}\).
For quantum states \(\rho\) and \(\sigma\) and \(\alpha \in (0,1)\cup(1,+\infty)\), we define the Petz [31] and sandwiched [23], [24] Rényi relative entropies, respectively, as \[\begin{align} D_{\alpha}(\rho\Vert\sigma) &\equiv \frac{1}{\alpha-1} \log \operatorname{Tr}\left[ \rho^{\alpha} \sigma^{1-\alpha} \right], \\ \widetilde{D}_{\alpha}(\rho\Vert\sigma) &\equiv \frac{1}{\alpha-1} \log \operatorname{Tr}\left[ \left( \sigma^{\frac{1-\alpha}{2\alpha}} \rho \sigma^{\frac{1-\alpha}{2\alpha}} \right)^{\alpha} \right]. \end{align}\] For \(\alpha < 1\), both quantities are defined to be infinite for orthogonal states. For \(\alpha > 1\), both quantities are defined for \(\text{supp} (\rho) \subseteq \text{supp}(\sigma)\), and infinite otherwise. The end points \(\alpha \in \{0,1,+\infty\}\) are defined by continuous extension. It is known that \(\widetilde{D}_{\alpha}(\rho\Vert\sigma) \leq D_{\alpha}(\rho\Vert\sigma)\) [23], [36], [37]; both quantities are non-decreasing in \(\alpha\) and \[\begin{align} \lim_{\alpha\to 1} \widetilde{D}_{\alpha}(\rho\Vert\sigma) = \lim_{\alpha\to 1} {D}_{\alpha}(\rho\Vert\sigma) = D_1(\rho\Vert\sigma) \equiv \operatorname{Tr}\left[\rho\left(\log \rho - \log \sigma\right)\right]. \end{align}\] The Petz–Rényi relative entropy is contractive under any quantum channel for \(\alpha \in [0,2]\) [31], while sandwiched Rényi relative entropy satisfies this property for \(\alpha \in [\frac{1}{2},+\infty]\) [23], [24], [37].
Define the Petz–Rényi information and sandwiched Rényi information of \(\mathscr{N}\) as \[\begin{align} {I}_{\alpha}(\mathscr{N}) &\equiv \sup_{\rho_{\mathsf{A}}} \min_{\sigma_{\mathsf{B}}} {D}_{\alpha} \left( \mathscr{N}(\psi_{\mathsf{R}\mathsf{A}}^{\rho}) \Vert \rho_{\mathsf{R}} \otimes \sigma_{\mathsf{B}} \right) \tag{21} \\ &= \sup_{\rho_{\mathsf{A}}} \frac{\alpha}{\alpha-1} \log \operatorname{Tr}\left[ \left( \operatorname{Tr}_{\mathsf{A}}[\rho_{\mathsf{A}}^{1-\alpha} (\sqrt{\rho_{\mathsf{A}}} \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}} \sqrt{\rho_{\mathsf{A}}})^{\alpha} ] \right)^{1/\alpha} \right],\tag{22} \\ \widetilde{I}_{\alpha}(\mathscr{N}) &\equiv \sup_{\rho_{\mathsf{A}}} \min_{\sigma_{\mathsf{B}}} \widetilde{D}_{\alpha} \left( \mathscr{N}(\psi_{\mathsf{R}\mathsf{A}}^{\rho}) \Vert \rho_{\mathsf{R}} \otimes \sigma_{\mathsf{B}} \right), \tag{23} \end{align}\] where \(\psi_{\mathsf{R}\mathsf{A}}^{\rho}\) is a purification of the input state \(\rho_{\mathsf{A}}\). Equality 22 follows from [38]. Note that for \(\alpha \in (0,1)\), both quantities are finite; we may change \(\sup_{\rho_{\mathsf{A}}}\) to \(\max_{\rho_{\mathsf{A}}}\). By the relations between the Petz and sandwiched Rényi relative entropies, we have [32] \[\begin{align} \widetilde{I}_{\alpha}(\mathscr{N}) \leq I_{\alpha}(\mathscr{N}), \quad \forall\, \alpha \geq 0; \quad \lim_{\alpha\to1} \widetilde{I}_{\alpha}(\mathscr{N}) = \lim_{\alpha\to1} {I}_{\alpha}(\mathscr{N}) = {I}_{1}(\mathscr{N}). \end{align}\]
The additivity notions of an extended-real-valued function \(f\) on the set of quantum channels are the following: \[\begin{align} &\text{(strongly additive)} &&f(\mathscr{N}_1\otimes\mathscr{N}_2) = f(\mathscr{N}_1) + f(\mathscr{N}_2); \\ &\text{(weakly additive)} &&f(\mathscr{N}\otimes\mathscr{N}) = 2 f(\mathscr{N}); \\ &\text{(superadditive)} &&f(\mathscr{N}_1\otimes\mathscr{N}_2) \geq f(\mathscr{N}_1) + f(\mathscr{N}_2); \\ &\text{(subadditive)} &&f(\mathscr{N}_1\otimes\mathscr{N}_2) \leq f(\mathscr{N}_1) + f(\mathscr{N}_2). \end{align}\]
In this section, we derive basic properties of the Petz–Rényi information \(I_{\alpha}(\mathscr{N})\). We first show that \(I_{\alpha}(\mathscr{N})\) is equivalent to a convex optimization (Proposition 1). Second, \(I_{\alpha}(\mathscr{N})\) is strongly additive for some special channels (Proposition 3). Third, \(I_{\alpha}(\mathscr{N})\) is strongly additive for any channel at \(\alpha = 2\) (Proposition 5).
Though the objective function on the right-most side of 22 is not concave in \(\rho_{\mathsf{A}}\) for \(\alpha \in (0,1)\), we can consider the minimization of the map \(\rho_{\mathsf{A}} \mapsto \operatorname{Tr}\left[ \left( \operatorname{Tr}_{\mathsf{A}} \left[ \rho_{\mathsf{A}}^{1-\alpha} \left( \sqrt{\rho_{\mathsf{A}}} \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}} \sqrt{\rho_{\mathsf{A}}} \right)^{\alpha} \right] \right)^{1/\alpha} \right]\) for \(\alpha \in (0,1)\), since the logarithm is monotone. The following Proposition 1 shows the convexity, which in turn, implies that the objective function 22 is quasi-concave in \(\rho_{\mathsf{A}}\) for any \(\alpha \in (0,1)\).
Proposition 1 (Convex optimization reduction). Let \(\rho_{\mathsf{A}}\) be a state and let \(\Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}}\) be the Choi operator of a channel \(\mathscr{N}_{\mathsf{A}\to \mathsf{B}}\). Then, the map \[\begin{align} \label{eq:app-objective} \rho_{\mathsf{A}} \mapsto \operatorname{Tr}\left[ \left( \operatorname{Tr}_{\mathsf{A}} \left[ \rho_{\mathsf{A}}^{1-\alpha} \left( \sqrt{\rho_{\mathsf{A}}} \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}} \sqrt{\rho_{\mathsf{A}}} \right)^{\alpha} \right] \right)^{1/\alpha} \right] \end{align}\qquad{(1)}\] on density operators is convex for any \(\alpha \in (0,1)\) and concave for any \(\alpha \in (1,2\,]\).
Proof. Write \(\Gamma_{\mathsf{A}\mathsf{B}}=\Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}}\). Without loss of generality, we only prove the case of \(\rho_{\mathsf{A}} > 0\) and \(\Gamma_{\mathsf{A}\mathsf{B}}>0\). The case of \(\rho_{\mathsf{A}} \geq 0\) and \(\Gamma_{\mathsf{A}\mathsf{B}} \geq 0\) follows from substituting \(\rho_{\mathsf{A}} \leftarrow (1-\epsilon) \rho_{\mathsf{A}} + \epsilon \mathbf{1}_{\mathsf{A}}/d_{\mathsf{A}}\), \(\Gamma_{\mathsf{A}\mathsf{B}}\leftarrow (1-\epsilon)\Gamma_{\mathsf{A}\mathsf{B}} + \epsilon \mathbf{1}_{\mathsf{A}}\otimes\mathbf{1}_{\mathsf{B}}/d_{\mathsf{B}}\), continuity, and letting \(\epsilon \searrow 0\).
First, consider \(\alpha \in (0,1)\). It is sufficient to show the convexity of the map \[\begin{align} \rho_{\mathsf{A}} \mapsto \operatorname{Tr}_{\mathsf{A}} \left[ \rho_{\mathsf{A}}^{1-\alpha} \left( \sqrt{\rho_{\mathsf{A}}} \Gamma_{\mathsf{A}\mathsf{B}} \sqrt{\rho_{\mathsf{A}}} \right)^{\alpha} \right] \end{align}\] on positive semi-definite operators, since \(\operatorname{Tr}[ \left(\cdot\right)^{1/\alpha} ]\) is convex and non-decreasing for \(\alpha \in (0,1)\).
Via polar decomposition, we have \(L f(L^\dagger L) = f( L L^\dagger) L\) for any unitary-invariant functional calculus \(f\). Then, \[\begin{align} \left( \sqrt{\rho_{\mathsf{A}}} \Gamma_{\mathsf{A}\mathsf{B}} \sqrt{\rho_{\mathsf{A}}} \right)^{\alpha} = \sqrt{\rho_{\mathsf{A}}} \sqrt{\Gamma_{\mathsf{A}\mathsf{B}}} \big( \sqrt{\Gamma_{\mathsf{A}\mathsf{B}}} \rho_{\mathsf{A}} \sqrt{\Gamma_{\mathsf{A}\mathsf{B}}} \big)^{\alpha-1} \sqrt{\Gamma_{\mathsf{A}\mathsf{B}}} \sqrt{\rho_{\mathsf{A}}}, \end{align}\] which means that it is equivalent to consider the map \[\begin{align} \rho_{\mathsf{A}} \mapsto &\operatorname{Tr}_{\mathsf{A}}\left[ \sqrt{\Gamma_{\mathsf{A}\mathsf{B}}} \big( \sqrt{\Gamma_{\mathsf{A}\mathsf{B}}} \rho_{\mathsf{A}} \sqrt{\Gamma_{\mathsf{A}\mathsf{B}}} \big)^{\alpha-1} \sqrt{\Gamma_{\mathsf{A}\mathsf{B}}} \rho_{\mathsf{A}}^{2-\alpha} \right] \\ &= \langle \tilde{\Phi} |_{\mathsf{A}\bar{\mathsf{A}}} \sqrt{\Gamma_{\mathsf{A}\mathsf{B}}} \big( \sqrt{\Gamma_{\mathsf{A}\mathsf{B}}} \rho_{\mathsf{A}} \sqrt{\Gamma_{\mathsf{A}\mathsf{B}}} \big)^{\alpha-1} \sqrt{\Gamma_{\mathsf{A}\mathsf{B}}} \otimes (\rho_{\bar{\mathsf{A}}}^{\top})^{2-\alpha} |\tilde{\Phi} \rangle_{\mathsf{A}\bar{\mathsf{A}}}, \label{eq:concavity95end} \end{align}\tag{24}\] where \(|\tilde{\Phi}\rangle_{\mathsf{A}\bar{\mathsf{A}}}\) was defined in 20 .
Now, invoke Lemma 2 below with \(X \leftarrow \sqrt{\Gamma_{\mathsf{A}\mathsf{B}}} \rho_{\mathsf{A}} \sqrt{\Gamma_{\mathsf{A}\mathsf{B}}}\), \(Y \leftarrow \rho_{\bar{\mathsf{A}}}^{\top}\), and \(p \leftarrow \alpha - 1 \in (-1,0)\). Then, the map given in 24 is convex by noting that the map \(\rho_{\mathsf{A}} \mapsto \sqrt{\Gamma_{\mathsf{A}\mathsf{B}}} \rho_{\mathsf{A}} \sqrt{\Gamma_{\mathsf{A}\mathsf{B}}}\) is linear, transpose \((\cdot)^{\top}\) is linear, and \(\langle \Phi |_{\mathsf{A}\bar{\mathsf{A}}} \sqrt{\Gamma_{\mathsf{A}\mathsf{B}}} (\cdot) \sqrt{\Gamma_{\mathsf{A}\mathsf{B}}} \otimes \mathbf{1}_{\bar{\mathsf{A}}} |\Phi \rangle_{\mathsf{A}\bar{\mathsf{A}}}\) is a positive map.
The proof of the case \(\alpha \in (1,2)\) follows similarly by noting that \(\operatorname{Tr}[ \left(\cdot\right)^{1/\alpha} ]\) is concave and non-decreasing for \(\alpha \in (1,2)\). For \(\alpha = 2\), we take the pointwise limit \(\alpha\nearrow 2\) of the concavity for \(\alpha \in (1,2)\).
Lemma 2 (Lieb’s Concavity Theorem [39] & Ando’s Convexity Theorem [40]). Let \(X,Y > 0\) be positive operators. The map \((X,Y)\mapsto X^{p} \otimes Y^{1-p}\) on positive definite operators is jointly convex for \(p \in (-1,0) \cup (1,2)\) and is jointly concave for \(p \in (0,1)\).
◻
Proposition 3 (Additivity for special channels). The following expressions and strong additivity hold for \(I_{\alpha}(\mathscr{N})\).
Isometric channels: \[\begin{align} I_{\alpha}(\mathscr{N}) = \sup_{\rho_{\mathsf{A}}} 2 H_{\frac{2-\alpha}{\alpha}}(\mathsf{A})_{\rho} = \begin{dcases} 2 \log d_{\mathsf{A}} & \alpha \in (0,2], \\ +\infty & \alpha > 2, \end{dcases} \end{align}\] where \(H_{\alpha}(\mathsf{A})_{\rho} \equiv \frac{1}{1-\alpha}\log \operatorname{Tr}[\rho^{\alpha}]\) is the Rényi entropy and \(d_{\mathsf{A}}\) denotes the dimension of the input Hilbert space.
Projective measurement channels with \(|\mathsf{Y}|\) nonzero projectors: \(I_{\alpha}(\mathscr{N}) = \log |\mathsf{Y}|\) for \(\alpha \in (0,2]\).
Covariant quantum channels 2: \(I_{\alpha}(\mathscr{N}) = \frac{1}{1-\alpha}\log d_{\mathsf{A}} + \frac{\alpha}{\alpha-1} \log \operatorname{Tr}\left[ \left( \operatorname{Tr}_{\mathsf{A}}\left[ (\Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}})^{\alpha} \right] \right)^{1/\alpha} \right]\) for \(\alpha \in (0,1)\cup(1,2]\).
Channels with commuting inputs: Suppose every admissible density operator \(\rho_{\mathsf{A}}\) in the input algebra commutes with \(\Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}}\), i.e., \[\begin{align} \label{eq:input95commuting} \left[ \rho_{\mathsf{A}} \otimes \mathbf{1}_{\mathsf{B}}, \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}} \right] = 0_{\mathsf{A}\mathsf{B}}, \end{align}\qquad{(2)}\] we have \[\begin{align} \label{eq:commuting-input-expression} I_{\alpha}(\mathscr{N}) = \sup_{ \rho_{\mathsf{A}} } \frac{\alpha}{\alpha-1} \log \operatorname{Tr}\left[ \left( \operatorname{Tr}_{\mathsf{A}} \left[ \rho_{\mathsf{A}} \cdot \Gamma_{\mathsf{A}\mathsf{B}}^{\alpha} \right] \right)^{1/\alpha} \right], \quad \alpha \in (0,1). \end{align}\qquad{(3)}\]
Remark 4. A trivial class of channels satisfying ?? is the replacer channel \(\mathscr{N}_{\mathsf{A}\to\mathsf{B}}(\rho_{\mathsf{A}}) = \sigma_{\mathsf{B}}\) for all \(\rho_{\mathsf{A}}\) on \(\mathsf{A}\), which has \(I_{\alpha}(\mathscr{N}) = 0\). Another class of channels satisfying ?? is classical-quantum channels \(\mathscr{N}_{\mathsf{X}\to\mathsf{B}}\), whose input states \(\rho_{\mathsf{X}}\) are restricted to diagonal matrices. The corresponding Petz–Rényi information (for \(\alpha \in (0,2]\)) is \[\begin{align} \label{eq:Petz-Renyi95c-q} I_{\alpha}(\mathscr{N}_{\mathsf{X}\to\mathsf{B}}) = \sup_{p_{\mathsf{X}}} \frac{\alpha}{\alpha-1} \log \operatorname{Tr}\left[ \left( \sum_{x\in\mathsf{X}} p_{\mathsf{X}}(x) \mathscr{N}(x)^{\alpha} \right)^{1/\alpha} \right] = \min_{\sigma_{\mathsf{B}}} \sup_{x\in\mathsf{X}} D_{\alpha}(\mathscr{N}(x)\Vert\sigma_{\mathsf{B}}), \end{align}\tag{25}\] where the last term is called Rényi divergence radius and was proved by Mosonyi and Ogawa [41] (see also [42], [43]). The subadditivity of \(I_{\alpha}(\mathscr{N}_{\mathsf{X}\to\mathsf{B}})\) also follows from the min-max expression and the product structure of classical-quantum channels, i.e., \(\mathscr{N}_{\mathsf{X}\to\mathsf{B}}^{\otimes 2}(x_1 x_2) = \mathscr{N}_{\mathsf{X}\to\mathsf{B}}(x_1)\otimes\mathscr{N}_{\mathsf{X}\to\mathsf{B}}(x_2)\). See also the proof by Li and Yang [44].
Proof. Item [app-item:isometric] (isometric channels): For an isometry \(V_{\mathsf{A}\to\mathsf{B}} |a\rangle_{\mathsf{A}} = |v_{a}\rangle_{\mathsf{B}}\), the joint state is \(\sqrt{\rho_{\mathsf{A}}} \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}} \sqrt{\rho_{\mathsf{A}}} = |\Psi\rangle\langle\Psi|_{\mathsf{A}\mathsf{B}}\), where \(|\Psi\rangle_{\mathsf{A}\mathsf{B}} = \sum_a \sqrt{\lambda_a} |a\rangle_{\mathsf{A}}|v_a\rangle_{\mathsf{B}}\). Its \(\alpha\)-power collapses to itself. We calculate \[\begin{align} \left( \operatorname{Tr}_{\mathsf{A}} \left[ \rho_{\mathsf{A}}^{1-\alpha} \left( \sqrt{\rho_{\mathsf{A}}} \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}} \sqrt{\rho_{\mathsf{A}}} \right)^{\alpha} \right] \right)^{1/\alpha} &= \left( \sum_{a} \lambda_a^{1-\alpha} \lambda_a \vert{}v_a\rangle\langle v_a\vert{}_{\mathsf{B}} \right)^{1/\alpha} = \sum_{a} \lambda_a^\frac{2-\alpha}{\alpha} \vert{}v_a\rangle\langle v_a\vert{}_{\mathsf{B}}. \end{align}\]
Item [app-item:projective] (projective measurement channels): For any quantum-classical channel \(\mathscr{M}_{\mathsf{A}\to\mathsf{Y}}(\rho_{\mathsf{A}}) = \sum_{y\in\mathsf{Y}} \operatorname{Tr}[\rho_{\mathsf{A}}M_{\mathsf{A}}^y] |y\rangle\langle y|_{\mathsf{Y}}\), we calculate the joint state as \[\begin{align} \omega_{\mathsf{A}\mathsf{Y}}^{\rho}= \sum_{y\in\mathsf{Y}} \sqrt{\rho_{\mathsf{A}}} (M_{\mathsf{A}}^y)^{\top} \sqrt{\rho_{\mathsf{A}}} \otimes \vert{}y\rangle\langle y\vert{}_{\mathsf{Y}}. \end{align}\] Hence, \[\begin{align} I_{\alpha}(\mathscr{M}_{\mathsf{A}\to\mathsf{Y}}) &= \max_{\rho_{\mathsf{A}}} \frac{\alpha}{\alpha-1} \log \sum_{y\in\mathsf{Y}} \left(\operatorname{Tr}\left[ \rho_{\mathsf{A}}^{1-\alpha} \left(\sqrt{\rho_{\mathsf{A}}} (M_{\mathsf{A}}^y)^{\top} \sqrt{\rho_{\mathsf{A}}}\right)^{\alpha} \right]\right)^{1/\alpha} \\ &= \max_{\rho_{\mathsf{A}}} \frac{\alpha}{\alpha-1} \log \sum_{y\in\mathsf{Y}} \left(\operatorname{Tr}\left[ \rho_{\mathsf{A}}^{1-\alpha} \left(\sqrt{\rho_{\mathsf{A}}} M_{\mathsf{A}}^y\sqrt{\rho_{\mathsf{A}}}\right)^{\alpha} \right]\right)^{1/\alpha}. \label{eq:q-c95formula} \end{align}\tag{26}\] Here, we drop ‘\(\top\)’ after optimization because \(\rho_{\mathsf{A}}\mapsto\rho_{\mathsf{A}}^{\top}\) is a bijection of the state space.
To derive the upper bound on 26 for projective measurements, we invoke the Araki-type trace inequality of Liu–Cheng [45]: \[\begin{align} \label{eq:Liu-Cheng} \operatorname{Tr}\left[f(A) A^s B^s \right] \leq \operatorname{Tr}\left[ f(A) \left(A^{\frac{1}{2}} B A^{\frac{1}{2}} \right)^s \right] & s\in(0,1] \end{align}\tag{27}\] with monotone function \(f(x) = x^{1-\alpha}\), \(s = \alpha \in (0,1)\), \(A\leftarrow \rho_{\mathsf{A}}\), and \(B\leftarrow M_{\mathsf{A}}^y\) to obtain \[\begin{align} I_{\alpha}(\mathscr{M}_{\mathsf{A}\to\mathsf{Y}}) &\leq \max_{\rho_{\mathsf{A}}} \frac{\alpha}{\alpha-1} \log \sum_{y\in\mathsf{Y}} \left( \operatorname{Tr}\left[ \rho_{\mathsf{A}} \left(M_{\mathsf{A}}^y\right)^{\alpha} \right] \right)^{1/\alpha} \\ &\overset{\text{(a)}}{=} \max_{\rho_{\mathsf{A}}} \frac{\alpha}{\alpha-1} \log \sum_{y\in\mathsf{Y}} \left( \operatorname{Tr}\left[ \rho_{\mathsf{A}} M_{\mathsf{A}}^y \right] \right)^{1/\alpha} \\ &= \max_{\rho_{\mathsf{A}}} H_{1/\alpha}(\mathsf{Y})_{\sum_{y}\operatorname{Tr}[\rho_{\mathsf{A}}M_{\mathsf{A}}^y]} \\ &\overset{\text{(b)}}{=} \log |\mathsf{Y}|, \quad \alpha \in (0,1). \end{align}\] where (a) follows from the projective measurement and (b) follows from the dimension bound for Rényi entropies. For \(\alpha \in (1,2]\), we apply 27 again with \(f(x) = x\), \(s = \alpha -1 \in (0,1]\), \(A \leftarrow \sqrt{\rho_{\mathsf{A}}} M_{\mathsf{A}}^y \sqrt{\rho_{\mathsf{A}}}\), and \(B\leftarrow \rho_{\mathsf{A}}^{-1}\) to obtain the same upper bound on \(I_{\alpha}(\mathscr{M}_{\mathsf{A}\to\mathsf{Y}})\).
The lower bound \(I_{\alpha}(\mathscr{M}_{\mathsf{A}\to\mathsf{Y}}) \geq \log |\mathsf{Y}|\) is achieved by choosing \(\rho_{\mathsf{A}} = \frac{1}{|\mathsf{Y}|} \sum_{y\in\mathsf{Y}} |u_y\rangle\langle u_y|_{\mathsf{A}}\) satisfying \(M_{\mathsf{A}}^y |u_y\rangle_{\mathsf{A}} = |u_y\rangle_{\mathsf{A}}\).
Item [app-item:covariant] (covariant channels): The inner trace function in ?? is unitary invariant with respect to the underlying group \(G\). Recalling the convexity (resp. concavity) for \(\alpha \in (0,1)\) (resp. \(\alpha \in (1,2]\)), the optimizer is attained at the depolarized completely mixed state \(\rho_{A}^{\star} = \int_G U_{\mathsf{A}} \rho_{\mathsf{A}} U_{\mathsf{A}}^{\dagger} \, \mathrm{d} U_{\mathsf{A}} = \frac{1}{d_{\mathsf{A}}} \mathbf{1}_{\mathsf{A}}\). Direct calculation proves the claim.
Item [app-item:commuting-input] (commuting input optimizers): The expression ?? directly follows from 22 and the commutation relation ?? .
The superadditivity directly follows from the definition 22 by choosing the product of optimal marginal states, i.e. \(\rho_{\mathsf{A}_1\mathsf{A}_2} = \rho_{\mathsf{A}_1}^\star \otimes \rho_{\mathsf{A}_2}^\star\). Below we prove the subadditivity.
Expressing ?? in terms of the Schatten norm, we have \[\begin{align} {I}_{\alpha}(\mathscr{N}) &= \frac{1}{\alpha-1} \log \inf_{\rho_{\mathsf{A}} \in \mathcal{S}(\mathsf{A})} \left\| \operatorname{Tr}_{\mathsf{A}}\left[ \rho_{\mathsf{A}} (\Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}})^{\alpha} \right] \right\|_{\frac{1}{\alpha}} \\ &\overset{\text{(a)}}{=} \frac{1}{\alpha-1} \log \inf_{\rho_{\mathsf{A}} \in \mathcal{S}(\mathsf{A})} \sup_{Z_{\mathsf{B}}\geq 0, \|Z_{\mathsf{B}}\|_{\frac{1}{1-\alpha}} \leq 1 } \operatorname{Tr}\left[ Z_{\mathsf{B}} \operatorname{Tr}_{\mathsf{A}}\left[ \rho_{\mathsf{A}} (\Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}})^{\alpha} \right] \right] \\ &\overset{\text{(b)}}{=} \frac{1}{\alpha-1} \log \sup_{Z_{\mathsf{B}}\geq 0, \|Z_{\mathsf{B}}\|_{\frac{1}{1-\alpha}} \leq 1 } \inf_{\rho_{\mathsf{A}} \in \mathcal{S}(\mathsf{A})}\operatorname{Tr}\left[ Z_{\mathsf{B}} \operatorname{Tr}_{\mathsf{A}}\left[ \rho_{\mathsf{A}} (\Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}})^{\alpha} \right] \right] \\ &= \frac{1}{\alpha-1} \log \sup_{Z_{\mathsf{B}}\geq 0, \|Z_{\mathsf{B}}\|_{\frac{1}{1-\alpha}} \leq 1 } \inf_{\rho_{\mathsf{A}} \in \mathcal{S}(\mathsf{A})}\operatorname{Tr}\left[ \rho_{\mathsf{A}} \otimes Z_{\mathsf{B}} (\Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}})^{\alpha} \right] \\ &= \frac{1}{\alpha-1} \log \sup_{Z_{\mathsf{B}}\geq 0, \|Z_{\mathsf{B}}\|_{\frac{1}{1-\alpha}} \leq 1 } \lambda_{\min} \left( \operatorname{Tr}_{\mathsf{B}}\left[ Z_{\mathsf{B}} (\Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}})^{\alpha} \right] \right), \label{eq:upper95additive1} \end{align}\tag{28}\] (a) follows from the duality of Schatten \(\frac{1}{\alpha}\) and \(\frac{1}{1-\alpha}\) norms; (b) follows from Sion’s minimax theorem and the bilinearity of the objective function. The subadditivity then follows by choosing \(Z_{\mathsf{B}_1\mathsf{B}_2} = Z_{\mathsf{B}_1}^\star \otimes Z_{\mathsf{B}_2}^\star\), where \(Z_{\mathsf{B}_i}\) is the optimizer in 28 for \(\mathscr{N}_i\). This concludes the proof. ◻
Proposition 5 (Additivity at \(\alpha =2\)). The Petz–Rényi information \(I_{\alpha}(\mathscr{N})\) is strongly additive for \(\alpha = 2\), i.e., \[\begin{align} I_2(\mathscr{N}_1\otimes\mathscr{N}_2) = I_2(\mathscr{N}_1) + I_2(\mathscr{N}_2). \end{align}\]
Proof. It suffices to prove the subadditivity since the superadditivity follows directly from the definition. Define \[\begin{align} Q_2(\mathscr{N}) &\mathrel{\vcenter{:}}= \sup_{\rho_{\mathsf{A}}} \operatorname{Tr}\!\left[ \left( \operatorname{Tr}_{\mathsf{A}}\!\left[ \rho_{\mathsf{A}}^{-1} \left( \sqrt{\rho_{\mathsf{A}}}\, \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}} \sqrt{\rho_{\mathsf{A}}} \right)^2 \right] \right)^{1/2} \right]. \label{eq:def-Q} \end{align}\tag{29}\] Hence, \(I_2(\mathscr{N})=2\log Q_2(\mathscr{N})\). Let \(\Pi_{\rho_{\mathsf{A}}}\) be the projection onto the support of \(\rho_{\mathsf{A}}\). By the cyclic property of trace, we have, \[\begin{align} Q_2(\mathscr{N}) &= \sup_{\rho_{\mathsf{A}}} \operatorname{Tr}\!\sqrt{ \operatorname{Tr}_{\mathsf{A}}\!\left[ \Pi_{\rho_{\mathsf{A}}} \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}} \rho_{\mathsf{A}} \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}} \Pi_{\rho_{\mathsf{A}}} \right]} \\ &\overset{\text{(a)}}{=}\sup_{\rho_{\mathsf{A}}>0, \operatorname{Tr}[\rho_{\mathsf{A}}]=1} \operatorname{Tr}\!\sqrt{ \operatorname{Tr}_{\mathsf{A}}\!\left[ \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}} \rho_{\mathsf{A}} \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}} \right]} \\ &=\max_{\rho_{\mathsf{A}}} \operatorname{Tr}\!\sqrt{ \operatorname{Tr}_{\mathsf{A}}\!\left[ \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}} \rho_{\mathsf{A}} \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}} \right]}, \label{eq:primal-Q} \end{align}\tag{30}\] where in (a) we restricted the optimization to the dense subset of full-rank \(\rho_{\mathsf{A}}\) because the objective function here is continuous in \(\rho_{\mathsf{A}}\).
Recall the Cauchy–Schwartz inequality, we have, for every \(X_{\mathsf{B}}\geq0\), \[\begin{align} \left(\operatorname{Tr}\sqrt{X_{\mathsf{B}}}\right)^2 =\inf_{\sigma_{\mathsf{B}} > 0,\, \operatorname{Tr}[\sigma_{\mathsf{B}}]=1} \operatorname{Tr}[X_{\mathsf{B}}\sigma_{\mathsf{B}}^{-1}]. \label{eq:sqrt-variational} \end{align}\tag{31}\] Hence, \[\begin{align} Q_2(\mathscr{N})^2 &= \max_{\rho_{\mathsf{A}}} \inf_{\sigma_{\mathsf{B}} > 0,\, \operatorname{Tr}[\sigma_{\mathsf{B}}]=1} \operatorname{Tr}\!\left[ \rho_{\mathsf{A}} \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}} (\mathbf{1}_{\mathsf{A}}\otimes\sigma_{\mathsf{B}}^{-1}) \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}} \right] \\ &= \inf_{\sigma_{\mathsf{B}} > 0,\, \operatorname{Tr}[\sigma_{\mathsf{B}}]=1} \max_{\rho_{\mathsf{A}}} \operatorname{Tr}\!\left[ \rho_{\mathsf{A}} \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}} (\mathbf{1}_{\mathsf{A}}\otimes\sigma_{\mathsf{B}}^{-1}) \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}} \right] \\ &= \inf_{\sigma_{\mathsf{B}} > 0,\, \operatorname{Tr}[\sigma_{\mathsf{B}}]=1} \left\| \operatorname{Tr}_{\mathsf{B}}\!\left[ \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}} (\mathbf{1}_{\mathsf{A}}\otimes\sigma_{\mathsf{B}}^{-1}) \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}} \right] \right\|_{\infty}. \end{align}\] Here, the objective function \((\rho_{\mathsf{A}},\sigma_{\mathsf{B}}) \mapsto \operatorname{Tr}_{\mathsf{B}}\!\left[ \rho_{\mathsf{A}} \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}} (\mathbf{1}_{\mathsf{A}}\otimes\sigma_{\mathsf{B}}^{-1}) \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr{N}} \right]\) is linear and convex via the operator convexity of the inverse function. We apply Sion’s minimax theorem by imposing \(\sigma_{\mathsf{B}} \geq \epsilon \mathbf{1}_{\mathsf{B}}\) for the resulting compact convex set and then letting \(\epsilon \searrow 0\).
The multiplicativity of \(Q_2(\mathscr{N})\) then follows by choosing product state \(\sigma_{\mathsf{B}_1\mathsf{B}_2} = \sigma_{\mathsf{B}_1}^\star \otimes \sigma_{\mathsf{B}_2}^\star\) and the multiplicativity of the operator norm \(\|\,\cdot\,\|_\infty\). ◻
This section establishes strict superadditivity for a Fourier measurement (Theorem 8). The one-copy optimization reduces to a diagonal one-dimensional convex problem. The strict two-copy improvement is then proved analytically by a linear-response calculation around the product of the one-copy optimizer toward a diagonal movement. Notably, the witness of the strict superadditivity is only via a classically correlated state on the two-copy system.
Throughout this section, we write \(d = d_{\mathsf{A}}\), and shorthand \(M_y = M_{\mathsf{A}}^y\). Fix \[\begin{align} d\ge 3, \qquad \frac{1}{d}<\lambda<1, \qquad b=\frac{1-\lambda}{d-1}. \end{align}\] Let \[\begin{align} a_0=\lambda, \qquad a_1=\cdots=a_{d-1}=b, \end{align}\] and let \(\omega=\mathrm{e}^{2\pi \mathrm{i}/d}\). On \(\mathsf{A}\cong\mathbb{C}^d\), define \[\begin{align} |v_y\rangle = \sum_{j=0}^{d-1}\sqrt{a_j}\,\omega^{jy}|j\rangle, \qquad y=0,\ldots,d-1. \end{align}\] The measurement \(\mathscr{M}\) consists of the \(d\) rank-one effects and the residual effect: \[\label{eq:Fourier95POVM} \begin{align} M_y&=\frac{1}{d\lambda} |v_y\rangle\!\langle v_y|, \qquad y=0,\ldots,d-1, \\ M_\infty &= \left(1-\frac{b}{\lambda}\right) \sum_{j=1}^{d-1}|j\rangle\!\langle j|. \end{align}\tag{32}\] Indeed, \[\begin{align} \sum_{y=0}^{d-1}M_y = \sum_{j=0}^{d-1}\frac{a_j}{\lambda}|j\rangle\!\langle j|, \end{align}\] and therefore \(\sum_{y=0}^{d-1}M_y+M_\infty=\mathbf{1}\).
Figure 5:
.
For notational simplicity, we define the one-copy functional \[\begin{align} Q_\alpha(\rho;\mathscr{M}) = \sum_y \left( \operatorname{Tr}\left[ \rho^{1-\alpha} \left(\rho^{1/2}M_y^{\top}\rho^{1/2}\right)^\alpha \right] \right)^{1/\alpha}, \quad \alpha \in (0,1), \end{align}\] where the sum includes the residual outcome \(y=\infty\), and define \[\begin{align} \label{eq:K95POVM} K^{(n)}(\alpha) \equiv \inf_{\rho_{\mathsf{A}^n}} Q_{\alpha}(\rho_{\mathsf{A}^n}; \mathscr{M}^{\otimes n}). \end{align}\tag{33}\] The strict superadditivity \[\begin{align} I_{\alpha}(\mathscr{M}^{\otimes 2}) > 2 I_{\alpha}(\mathscr{M}) \end{align}\] is equivalent to \[\begin{align} K^{(2)}(\alpha)<K^{(1)}(\alpha)^2. \end{align}\]
The diagonal reduction follows from covariance and convexity. Let \(Z = \sum_{j=0}^{d-1} \omega^j |j\rangle\langle j|\) be the generalized Pauli phase operator and \[\begin{align} Z^m|j\rangle=\omega^{mj}|j\rangle, \qquad m=0,\ldots,d-1. \end{align}\] The phase operator permutes the transposed Fourier effects \(M_y^{\top}\) and leaves \(M_\infty^{\top}=M_\infty\) fixed. Hence \[\begin{align} Q_\alpha(Z^m\rho Z^{-m};\mathscr{M})=Q_\alpha(\rho;\mathscr{M}). \end{align}\] The phase twirl \[\begin{align} \Delta(\rho) = \frac{1}{d}\sum_{m=0}^{d-1}Z^m \rho Z^{-m} = \sum_{j=0}^{d-1}|j\rangle\langle j|\rho|j\rangle\langle j| \end{align}\] therefore satisfies, by the convexity proven in Proposition 1, \[\begin{align} Q_\alpha(\Delta(\rho);\mathscr{M}) \le \frac{1}{d}\sum_{m=0}^{d-1}Q_\alpha(Z^m \rho Z^{-m};\mathscr{M}) = Q_\alpha(\rho;\mathscr{M}). \end{align}\] Hence, a one-copy optimizer may be chosen diagonal.
For a diagonal state \(p=(p_0,\ldots,p_{d-1})\), the objective is \[\label{eq:diagonal-objective} Q_\alpha(p;\mathscr{M}) = \frac{1}{\lambda} \left(\sum_{j=0}^{d-1}a_jp_j\right)^{\frac{\alpha-1}{\alpha}} \left(\sum_{j=0}^{d-1}a_jp_j^{2-\alpha}\right)^{\frac{1}{\alpha}} + \left(1 - \frac{b}{\lambda}\right)\left(\sum_{j=1}^{d-1}p_j\right)^{1/\alpha}.\tag{34}\] Fix \(t=p_0\). The first factor \(\sum_j a_jp_j=\lambda t+b(1-t)\) and the residual term depend only on \(t\). The remaining tail dependence is through \(\sum_{j=1}^{d-1}p_j^{2-\alpha}\). Since \(2-\alpha>1\), the function \(r\mapsto r^{2-\alpha}\) is strictly convex. Hence, for fixed tail mass \(1-t\), this sum is uniquely minimized by the uniform tail \[\begin{align} p_1=\cdots=p_{d-1}=\frac{1-t}{d-1}. \end{align}\] Therefore \[\begin{align} K^{(1)}(\alpha)=\min_{0\le t\le1} q_\alpha(t), \end{align}\] where \[\begin{align} q_\alpha(t)&=r_\alpha(t)+z_\alpha(t), \\ r_\alpha(t) &= \frac{1}{\lambda}\,\eta(t)^{\frac{\alpha-1}{\alpha}}\, \zeta_\alpha(t)^{\frac{1}{\alpha}}, \\ z_\alpha(t) &= \left(1-\frac{b}{\lambda}\right)(1-t)^{1/\alpha}, \end{align}\] and \[\begin{align} \eta(t)=\lambda t+b(1-t), \qquad \zeta_\alpha(t) = \lambda t^{2-\alpha}+(d-1)b\left(\frac{1-t}{d-1}\right)^{2-\alpha}. \end{align}\] This is the diagonal one-parameter reduction.
Lemma 6 (Convexity of the reduced one-copy objective, POVM). For every \(0<\alpha<1\), the function \[\begin{align} t\mapsto q_{\alpha}(t) = Q_\alpha(\rho_t;\mathscr{M}) \end{align}\] is convex on \([0,1]\).
Proof. Set \(p=1/\alpha>1\) and define \[\begin{align} \phi(z,e) := e\left(\frac{z}{e}\right)^p = z^p e^{1-p}, \qquad z\geq0,\quad e>0. \end{align}\] The function \(\phi\) is the perspective of the convex function \(z\mapsto z^p\), and hence is jointly convex. Moreover, \[\begin{align} \frac{\partial\phi}{\partial z}(z,e) = pz^{p-1}e^{1-p} \geq0, \end{align}\] thereby \(\phi\) is nondecreasing in its first argument.
Notice that \(\eta(t)>0\) is affine and \(\zeta_\alpha(t)\) is convex on \([0,1]\), since \(2-\alpha>1\). Therefore, for every \(s,t\in[0,1]\) and \(\theta\in[0,1]\), \[\begin{align} &r_\alpha\bigl(\theta t+(1-\theta)s\bigr) \nonumber\\ &\quad= \frac{1}{\lambda} \phi\left( \zeta_\alpha\bigl(\theta t+(1-\theta)s\bigr), \eta\bigl(\theta t+(1-\theta)s\bigr) \right) \nonumber\\ &\quad\leq \frac{1}{\lambda} \phi\left( \theta\zeta_\alpha(t)+(1-\theta)\zeta_\alpha(s), \theta\eta(t)+(1-\theta)\eta(s) \right) \nonumber\\ &\quad\leq \theta r_\alpha(t)+(1-\theta)r_\alpha(s). \end{align}\] The first inequality follows from the convexity of \(\zeta_\alpha\), the affinity of \(\eta\), and the monotonicity of \(\phi\) in its first argument; the second follows from the joint convexity of \(\phi\). Hence \(r_\alpha\) is convex.
Finally, since \(1/\alpha>1\), \[\begin{align} z_\alpha(t) = \left(1-\frac{b}{\lambda}\right)(1-t)^{1/\alpha} \end{align}\] is convex. Therefore \(q_\alpha(t)=r_\alpha(t)+z_\alpha(t)\) is convex on \([0,1]\). ◻
Next, we show that the minimizer is an interior point. Indeed, a direct endpoint check gives \[\begin{align} q_\alpha'(0^+)<0, \qquad q_\alpha'(1^-)>0. \end{align}\] Let \(t_\alpha\in(0,1)\) be a global minimizer. Then \[\label{eq:stationarity} q_\alpha'(t_\alpha)=0.\tag{35}\] Moreover, at the uniform point \(t=1/d\), \[\label{eq:uniform-derivative-new} q_\alpha'\!\left(\frac{1}{d}\right) = \frac{d(\lambda-b)}{\alpha \lambda}\,d^{-1/\alpha} \left[1-(d-1)^{1/\alpha-1}\right] <0,\tag{36}\] where the strict inequality uses \(d\ge3\) and \(\alpha<1\). Hence no global minimizer can be the uniform point: \[\label{eq:not-uniform} t_\alpha\ne \frac{1}{d}.\tag{37}\]
Let \[\begin{align} p=p(t_\alpha) = \left(t_\alpha,\frac{1-t_\alpha}{d-1},\ldots,\frac{1-t_\alpha}{d-1}\right), \end{align}\] and define the zero-sum vector \[\begin{align} u=\left(1,-\frac{1}{d-1},\ldots,-\frac{1}{d-1}\right). \end{align}\] For real \(\kappa\), set \[\label{eq:pi-kappa} \pi_\kappa(i,j)=p_ip_j+\kappa \cdot u_i u_j.\tag{38}\] For all sufficiently small \(|\kappa|\), \(\pi_\kappa\) is a bipartite probability distribution. Because \(\sum_i u_i=0\), it has the same one-copy marginals as \(p\). Denote a diagonal state on \(\mathbb{C}^{d\times d}\) by \[\begin{align} \label{eq:two-copy95POVM} \rho_{\alpha}(\kappa) = \sum_{i,j=0}^{d-1}\pi_\kappa(i,j)|ij\rangle\!\langle ij|. \end{align}\tag{39}\] At \(\kappa=0\), this is the product of the one-copy optimizer with itself, i.e., \(\pi_0 = p\otimes p\). Hence, \[\begin{align} Q_\alpha(\rho_{\alpha}(0);\mathscr{M}^{\otimes 2})=K^{(1)}(\alpha)^2. \end{align}\] Define the linear response \[\begin{align} c_1(\alpha) = \left. \frac{\mathrm{d}}{\mathrm{d}\kappa}Q_\alpha(\rho_{\alpha}(\kappa);\mathscr{M}^{\otimes 2}) \right|_{\kappa=0}. \end{align}\] If \(c_1(\alpha)<0\), then for sufficiently small positive \(\kappa\), \[\begin{align} Q_\alpha(\rho_{\alpha}(\kappa);\mathscr{M}^{\otimes 2}) = K^{(1)}(\alpha)^2 + c_{1}(\alpha)\kappa + \mathcal{O}(\kappa^2) <K^{(1)}(\alpha)^2, \end{align}\] and hence \(K^{(2)}(\alpha)<K^{(1)}(\alpha)^2\).
For \(0<t<1\), define \[\begin{align} \xi(t)=\frac{\lambda-b}{\eta(t)} \end{align}\] and \[\begin{align} \theta_\alpha(t) = \frac{1}{\zeta_\alpha(t)} \left[ \lambda t^{1-\alpha} - b \left(\frac{1-t}{d-1}\right)^{1-\alpha} \right]. \end{align}\] Then \[\begin{align} \frac{r_\alpha'(t)}{r_\alpha(t)} = \varphi_\alpha(t), \qquad \varphi_\alpha(t) = \frac{\alpha-1}{\alpha}\xi(t) + \frac{2-\alpha}{\alpha}\theta_\alpha(t), \end{align}\] and \[\begin{align} z_\alpha'(t)=-\frac{z_\alpha(t)}{\alpha(1-t)}. \end{align}\] The stationarity condition 35 is therefore \[\label{eq:rz-stationarity} r_\alpha(t_\alpha)\varphi_\alpha(t_\alpha) = \frac{z_\alpha(t_\alpha)}{\alpha(1-t_\alpha)}.\tag{40}\]
Lemma 7 (Linear response at the one-copy minimizer). At \(t=t_\alpha\), \[\label{eq:c1-formula} c_1(\alpha) = \frac{(\alpha-1)(2-\alpha)}{\alpha} r_\alpha(t_\alpha)^2 \left[\xi(t_\alpha)-\theta_\alpha(t_\alpha)\right]^2.\qquad{(4)}\]
Proof. Differentiate the four two-copy outcome types along the path 38 : Fourier–Fourier, Fourier–residual, residual–Fourier, and residual–residual. At \(\kappa=0\), all four contributions factor into the one-copy terms \(r_\alpha\) and \(z_\alpha\). A direct differentiation gives \[\begin{align} c_1(\alpha) &= r_\alpha^2 \left[ \frac{\alpha-1}{\alpha}\xi^2 + \frac{2-\alpha}{\alpha}\theta_\alpha^2 \right] - \frac{2\varphi_\alpha}{1-t}\,r_\alpha z_\alpha + \frac{z_\alpha^2}{\alpha(1-t)^2}, \end{align}\] where all quantities on the right are evaluated at \(t=t_\alpha\). Using 40 , the last two terms become \(-\alpha r_\alpha^2\varphi_\alpha^2\). Hence \[\begin{align} c_1(\alpha) = r_\alpha^2 \left[ \frac{\alpha-1}{\alpha}\xi^2 + \frac{2-\alpha}{\alpha}\theta_\alpha^2 - \alpha\varphi_\alpha^2 \right]. \end{align}\] Since \[\begin{align} \alpha\varphi_\alpha=(\alpha-1)\xi+(2-\alpha)\theta_\alpha, \end{align}\] the bracket is \[\begin{align} \frac{(\alpha-1)(2-\alpha)}{\alpha}(\xi-\theta_\alpha)^2. \end{align}\] This proves ?? . ◻
The prefactor in ?? is strictly negative for \(0<\alpha<1\). It remains only to show that the square is nonzero. The equality case is explicit: \[\begin{align} \xi(t)=\theta_\alpha(t) \quad\Longleftrightarrow\quad t=\frac{1}{d}. \end{align}\] Indeed, \[\begin{align} \left(\lambda t^{1-\alpha}-b(\tfrac{1-t}{d-1})^{1-\alpha}\right)\eta(t) -(\lambda-b)\zeta_\alpha(t) = \lambda b\left(t^{1-\alpha}-(\tfrac{1-t}{d-1})^{1-\alpha}\right). \end{align}\] We have \(\xi(t)=\theta_\alpha(t)\) if and only if \(t^{1-\alpha}=(\frac{1-t}{d-1})^{1-\alpha}\), equivalently \(t=1/d\). By 37 , the minimizer \(t_\alpha\) is not \(1/d\). Therefore, \[\begin{align} \xi(t_\alpha)\ne\theta_\alpha(t_\alpha), \end{align}\] and ?? gives \[\label{eq:c1-negative} c_1(\alpha)<0.\tag{41}\]
Theorem 8 (Strict superadditivity of Fourier measurements). For every \(d\ge3\), every \(\lambda\in(1/d,1)\), and every \(0<\alpha<1\), the single-heavy Fourier measurement defined in 32 satisfies \[\begin{align} I_{\alpha}(\mathscr{M}^{\otimes 2}) > 2 I_{\alpha}(\mathscr{M}). \end{align}\] The strict improvement is witnessed by the infinitesimal diagonal classically correlated path \(\rho_\alpha(\kappa)\) in 39 .
Remark 9. In this paper, we only focus on the range \(\alpha \in (0,1)\) for \(I_{\alpha}(\mathscr{N})\). However, the proof of Theorem 8 naturally extends to \(\alpha \in (1,2)\).
Proof. By the diagonal reduction, \(K^{(1)}(\alpha)=q_\alpha(t_\alpha)\). The product point \(\rho_\alpha(0)\) satisfies \[\begin{align} Q_\alpha(\rho_\alpha(0);\mathscr{M}^{\otimes 2})=K^{(1)}(\alpha)^2. \end{align}\] The linear-response calculation gives \(c_1(\alpha)<0\). Hence, for sufficiently small positive \(\kappa\), \[\begin{align} Q_\alpha(\rho_\alpha(\kappa);\mathscr{M}^{\otimes 2})<K^{(1)}(\alpha)^2. \end{align}\] Therefore \[\begin{align} K^{(2)}(\alpha)<K^{(1)}(\alpha)^2. \end{align}\] Multiplying the logarithmic inequality by the negative number \(\alpha/(\alpha-1)\) proves the claim. ◻
For strictness, the sign \(c_1(\alpha)<0\) is enough. In computations one may choose \(\kappa\) by the one-dimensional minimization of \(\kappa\mapsto Q_\alpha(\rho_\kappa;\mathscr{M}^{\otimes 2})\). Equivalently, if \[\begin{align} Q_\alpha(\rho_\alpha(\kappa);\mathscr{M}^{\otimes 2}) = K^{(1)}(\alpha)^2+c_1(\alpha)\kappa+c_2(\alpha)\kappa^2+O(\kappa^3) \end{align}\] and \(c_2(\alpha)>0\), the quadratic approximation gives \[\begin{align} \label{eq:kappa95quad} \kappa_{\mathrm{quad}}(\alpha) = -\frac{c_1(\alpha)}{2c_2(\alpha)}. \end{align}\tag{42}\] This is the clean analytic version of the optimized-\(\kappa\) witness. It is especially useful near \(\alpha=1\), where the final gap becomes very small. The proof above does not require resolving that small gap numerically; it only uses the exact sign of the derivative \(c_1(\alpha)\), which remains negative for every fixed \(\alpha<1\).
The following numerical examples illustrate the scale of the linear response with parameters \[\begin{align} d=4, \qquad \lambda=\frac{9}{25}, \qquad b=\frac{16}{75}. \end{align}\]
| \(\alpha\) | \(c_1(\alpha)\) | \(\kappa_{\mathrm{quad}}(\alpha)\) | leading-order gap |
|---|---|---|---|
| \(0.50\) | \(-9.818\times10^{-3}\) | \(2.563\times10^{-3}\) | \(9.13\times10^{-5}\) |
| \(0.60\) | \(-6.651\times10^{-3}\) | \(1.510\times10^{-3}\) | \(2.91\times10^{-5}\) |
| \(0.70\) | \(-3.230\times10^{-3}\) | \(7.570\times10^{-4}\) | \(6.91\times10^{-6}\) |
| \(0.80\) | \(-1.021\times10^{-3}\) | \(3.021\times10^{-4}\) | \(1.04\times10^{-6}\) |
| \(0.90\) | \(-1.308\times10^{-4}\) | \(6.968\times10^{-5}\) | \(5.19\times10^{-8}\) |
For \(\alpha \in [\frac{1}{2},1)\) and the above chosen parameters, one can also choose an explicit \(\kappa(\alpha)\) as \[\begin{align} \label{eq:kappa95closed} \kappa(\alpha) = \frac{7}{250}(1-\alpha)^2 (2-\alpha) t_{\alpha} (1-t_{\alpha}). \end{align}\tag{43}\] Figure 3 plots the above numeric example. However, note that 43 is not a universal witness; one has to consider at least the second-order derivative bound.
The gap tends to zero as \(\alpha\nearrow1\), but strictness does not rely on a floating-point comparison of nearly equal numbers. The analytic certificate is the identity ?? together with \(t_\alpha\ne1/d\).
Let \(\mathscr N_{\mathsf{A}\to\mathsf{B}}\) be the amplitude-damping channel with damping parameter \(0<\gamma<1\), and put \[\begin{align} \eta:=1-\gamma. \end{align}\] Its Choi matrix is \[\begin{align} \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr N} = \begin{pmatrix} 1&0&0&\sqrt{\eta}\\ 0&0&0&0\\ 0&0&\gamma&0\\ \sqrt{\eta}&0&0&\eta \end{pmatrix}. \label{eq:GAD-app} \end{align}\tag{44}\] Equivalently, \[\begin{align} \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr N} = |\phi_0\rangle\!\langle\phi_0| + |\phi_1\rangle\!\langle\phi_1|, \qquad |\phi_0\rangle = |00\rangle+\sqrt{\eta}\,|11\rangle, \qquad |\phi_1\rangle = \sqrt{\gamma}\,|10\rangle. \label{eq:AD-Choi-branches} \end{align}\tag{45}\]
For \(0<\alpha<1\), recall the Petz–Rényi trace functional \[\begin{align} Q_{\alpha}(\rho;\mathscr N) &:= \operatorname{Tr}\left[ \left( \operatorname{Tr}_{\mathsf{A}}\left[ \rho_{\mathsf{A}}^{1-\alpha} \left( \sqrt{\rho_{\mathsf{A}}}\, \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr N}\, \sqrt{\rho_{\mathsf{A}}} \right)^{\alpha} \right] \right)^{1/\alpha} \right], \tag{46} \\ K^{(n)}(\alpha) &:= \inf_{\rho_{\mathsf{A}^n}\in\mathcal{D}(\mathsf{A}^n)} Q_{\alpha}(\rho_{\mathsf{A}^n};\mathscr N^{\otimes n}). \tag{47} \end{align}\] Since \(0<\alpha<1\), the coefficient \(\alpha/(\alpha-1)\) is negative, and hence \[\begin{align} I_{\alpha}(\mathscr N^{\otimes n}) = \frac{\alpha}{\alpha-1}\log K^{(n)}(\alpha). \label{eq:AD-I-via-K} \end{align}\tag{48}\] Consequently, it is enough to prove \[\begin{align} K^{(2)}(\alpha) < \bigl(K^{(1)}(\alpha)\bigr)^2. \label{eq:AD-target-K} \end{align}\tag{49}\]
Let \[\begin{align} Z:=|0\rangle\!\langle0|-|1\rangle\!\langle1|. \end{align}\] The amplitude-damping Choi matrix satisfies \[\begin{align} (Z_{\mathsf{A}}\otimes Z_{\mathsf{B}}) \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr N} (Z_{\mathsf{A}}\otimes Z_{\mathsf{B}}) = \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr N}. \label{eq:AD-Z-Choi} \end{align}\tag{50}\]
For convenience, define the positive operator \[\begin{align} T_{\alpha}^{\mathscr N}(\rho) := \operatorname{Tr}_{\mathsf{A}}\left[ \rho_{\mathsf{A}}^{1-\alpha} \left( \sqrt{\rho_{\mathsf{A}}}\, \Gamma_{\mathsf{A}\mathsf{B}}^{\mathscr N}\, \sqrt{\rho_{\mathsf{A}}} \right)^{\alpha} \right], \label{eq:AD-T-def} \end{align}\tag{51}\] so that \[\begin{align} Q_{\alpha}(\rho;\mathscr N) = \operatorname{Tr}\left[ T_{\alpha}^{\mathscr N}(\rho)^{1/\alpha} \right]. \end{align}\]
Lemma 10 (Diagonal reduction for amplitude damping). For every \(0<\alpha<1\), \[\begin{align} K^{(1)}(\alpha) = \min_{0\leq t\leq1} Q_{\alpha}(\rho_t;\mathscr N), \qquad \rho_t:=\mathrm{diag}(t,1-t). \label{eq:AD-diagonal-reduction} \end{align}\qquad{(5)}\]
Proof. For every state \(\rho\), 50 gives \[\begin{align} \sqrt{Z\rho Z}\, \Gamma^{\mathscr N}\, \sqrt{Z\rho Z} = (Z_{\mathsf{A}}\otimes Z_{\mathsf{B}}) \left( \sqrt{\rho}\, \Gamma^{\mathscr N}\, \sqrt{\rho} \right) (Z_{\mathsf{A}}\otimes Z_{\mathsf{B}}). \end{align}\] Using also \((Z\rho Z)^{1-\alpha}=Z\rho^{1-\alpha}Z\), we obtain \[\begin{align} T_{\alpha}^{\mathscr N}(Z\rho Z) = Z_{\mathsf{B}}T_{\alpha}^{\mathscr N}(\rho)Z_{\mathsf{B}}. \end{align}\] Therefore \[\begin{align} Q_{\alpha}(Z\rho Z;\mathscr N) = Q_{\alpha}(\rho;\mathscr N). \label{eq:AD-Z-invariance} \end{align}\tag{52}\]
Consider the computational-basis pinching \[\begin{align} \Delta_Z(\rho) := \frac{1}{2}\left(\rho+Z\rho Z\right). \end{align}\] For a qubit, \(\Delta_Z(\rho)\) is diagonal. By Proposition 1, the map \(\rho\mapsto Q_{\alpha}(\rho;\mathscr N)\) is convex for \(0<\alpha<1\). Hence, using 52 , \[\begin{align} Q_{\alpha}(\Delta_Z(\rho);\mathscr N) &\leq \frac{1}{2} Q_{\alpha}(\rho;\mathscr N) + \frac{1}{2} Q_{\alpha}(Z\rho Z;\mathscr N) \notag\\ &= Q_{\alpha}(\rho;\mathscr N). \end{align}\] Thus every state can be replaced by a diagonal state without increasing the objective. Since diagonal states are themselves admissible, this proves ?? . The minimum is attained by finite-dimensional continuity and compactness of the state space. ◻
Set \[\begin{align} u:=1-t, \qquad \ell(t):=t+\eta u=\eta+\gamma t. \label{eq:AD-u-ell} \end{align}\tag{53}\] For the diagonal state \(\rho_t\), define \[\begin{align} |\psi_t\rangle := \sqrt t\,|00\rangle+\sqrt{\eta u}\,|11\rangle. \end{align}\] Using 45 , we have \[\begin{align} \sqrt{\rho_t}\, \Gamma^{\mathscr N}\, \sqrt{\rho_t} = |\psi_t\rangle\!\langle\psi_t| + \gamma u\,|10\rangle\!\langle10|. \label{eq:AD-one-copy-blocks} \end{align}\tag{54}\] The two summands have orthogonal supports, and \(\langle\psi_t|\psi_t\rangle=\ell(t)\). Therefore \[\begin{align} \left( \sqrt{\rho_t}\, \Gamma^{\mathscr N}\, \sqrt{\rho_t} \right)^\alpha = \ell(t)^{\alpha-1} |\psi_t\rangle\!\langle\psi_t| + (\gamma u)^\alpha |10\rangle\!\langle10|. \end{align}\] After multiplying by \(\rho_t^{1-\alpha}\) and tracing out \(\mathsf{A}\), we obtain \[\begin{align} T_{\alpha}^{\mathscr N}(\rho_t) = w_0(t)|0\rangle\!\langle0| + w_1(t)|1\rangle\!\langle1|, \label{eq:AD-one-copy-T} \end{align}\tag{55}\] where \[\begin{align} w_0(t) &= t^{2-\alpha}\ell(t)^{\alpha-1} + \gamma^\alpha(1-t), \tag{56} \\ w_1(t) &= \eta(1-t)^{2-\alpha}\ell(t)^{\alpha-1}. \tag{57} \end{align}\] Consequently, \[\begin{align} q_{\alpha}(t) &:= Q_{\alpha}(\rho_t;\mathscr N) = w_0(t)^{1/\alpha} + w_1(t)^{1/\alpha}. \label{eq:AD-one-copy-q} \end{align}\tag{58}\]
Lemma 11 (Convexity and interiority of the one-copy minimizer). For every \(0<\gamma<1\) and \(0<\alpha<1\), the function \(q_{\alpha}\) is convex on \([0,1]\). Moreover, \[\begin{align} q_{\alpha}'(0+)<0, \qquad q_{\alpha}'(1-)>0. \end{align}\] In particular, every minimizer \(t_\alpha\) of \(q_\alpha\) belongs to \((0,1)\).
Proof. Since \(2-\alpha>1\), the function \(r\mapsto r^{2-\alpha}\) is convex on \([0,\infty)\). Its perspective \[\begin{align} (r,z) \mapsto z\left(\frac{r}{z}\right)^{2-\alpha} = r^{2-\alpha}z^{\alpha-1} \end{align}\] is therefore jointly convex for \(r\geq0\) and \(z>0\). Because \(t\), \(1-t\), and \(\ell(t)\) are affine in \(t\), both \[\begin{align} t^{2-\alpha}\ell(t)^{\alpha-1}, \qquad (1-t)^{2-\alpha}\ell(t)^{\alpha-1} \end{align}\] are convex. Thus \(w_0\) and \(w_1\) are nonnegative convex functions. Since \(r\mapsto r^{1/\alpha}\) is convex and increasing for \(0<\alpha<1\), the function \(q_\alpha=w_0^{1/\alpha}+w_1^{1/\alpha}\) is convex.
A direct evaluation of the one-sided derivatives gives \[\begin{align} q_{\alpha}'(0+) = -\frac{2-\alpha}{\alpha} <0, \label{eq:AD-q-derivative-zero} \end{align}\tag{59}\] and \[\begin{align} q_{\alpha}'(1-) = \frac{ 2-\alpha+(\alpha-1)\gamma-\gamma^\alpha }{\alpha}. \label{eq:AD-q-derivative-one} \end{align}\tag{60}\] To see that the latter is positive, define \[\begin{align} h_\alpha(\gamma) := 2-\alpha+(\alpha-1)\gamma-\gamma^\alpha. \end{align}\] Then \[\begin{align} h_\alpha(1)=0, \qquad h_\alpha'(\gamma) = \alpha-1-\alpha\gamma^{\alpha-1} <0. \end{align}\] Hence \(h_\alpha(\gamma)>h_\alpha(1)=0\) whenever \(0<\gamma<1\). Neither endpoint can therefore minimize \(q_\alpha\), proving the claim. ◻
Fix any minimizer \(t_\alpha\in(0,1)\), and abbreviate \[\begin{align} t:=t_\alpha, \qquad u:=1-t, \qquad \ell:=\eta+\gamma t, \qquad w_i:=w_i(t),\quad i\in\{0,1\}. \label{eq:AD-minimizer-abbreviations} \end{align}\tag{61}\] Then \[\begin{align} K^{(1)}(\alpha) = q_\alpha(t) = w_0^{1/\alpha}+w_1^{1/\alpha}. \label{eq:AD-K1} \end{align}\tag{62}\]
Consider the diagonal two-copy state \[\begin{align} \rho_\alpha(\kappa) &= \rho_t\otimes\rho_t + \kappa\, \mathrm{diag}(1,-1,-1,1) \notag\\ &= \mathrm{diag}\left( p_{00}(\kappa), p_{01}(\kappa), p_{10}(\kappa), p_{11}(\kappa) \right), \label{eq:two-copy95GAD} \end{align}\tag{63}\] where \[\begin{align} p_{00}(\kappa)&=t^2+\kappa, & p_{01}(\kappa)&=tu-\kappa, \notag\\ p_{10}(\kappa)&=tu-\kappa, & p_{11}(\kappa)&=u^2+\kappa. \label{eq:AD-p-kappa} \end{align}\tag{64}\] This is a full-rank density operator whenever \[\begin{align} -\min\{t^2,u^2\}<\kappa<tu. \label{eq:AD-kappa-interval} \end{align}\tag{65}\] Moreover, the perturbation preserves both one-copy marginals: \[\begin{align} \operatorname{Tr}_{\mathsf{A}_2}\rho_\alpha(\kappa) = \operatorname{Tr}_{\mathsf{A}_1}\rho_\alpha(\kappa) = \rho_t. \label{eq:AD-fixed-marginals} \end{align}\tag{66}\] Thus positive \(\kappa\) introduces a classical correlation without changing either marginal.
For product states and product channels, the functional is multiplicative: \[\begin{align} Q_\alpha(\rho\otimes\sigma;\mathscr N\otimes\mathscr M) = Q_\alpha(\rho;\mathscr N) Q_\alpha(\sigma;\mathscr M). \label{eq:AD-Q-product} \end{align}\tag{67}\] Indeed, the operator \(T_\alpha\) in 51 factors as a tensor product, and both the \(1/\alpha\)-power and the trace factor. Consequently, \[\begin{align} Q_\alpha(\rho_\alpha(0);\mathscr N^{\otimes2}) = Q_\alpha(\rho_t;\mathscr N)^2 = \bigl(K^{(1)}(\alpha)\bigr)^2. \label{eq:AD-product-value} \end{align}\tag{68}\]
The decomposition 45 gives four two-copy branches \[\begin{align} |\phi_r\rangle_{\mathsf{A}_1\mathsf{B}_1} \otimes |\phi_s\rangle_{\mathsf{A}_2\mathsf{B}_2}, \qquad r,s\in\{0,1\}. \end{align}\] For the diagonal state \(\rho_\alpha(\kappa)\), define \[\begin{align} |\psi_{rs}(\kappa)\rangle := \left( \sqrt{\rho_\alpha(\kappa)_{\mathsf{A}_1\mathsf{A}_2}} \otimes\mathbf{1}_{\mathsf{B}_1\mathsf{B}_2} \right) |\phi_r\rangle_{\mathsf{A}_1\mathsf{B}_1} |\phi_s\rangle_{\mathsf{A}_2\mathsf{B}_2}. \end{align}\] The four vectors \(|\psi_{rs}(\kappa)\rangle\) have mutually orthogonal supports. Hence \[\begin{align} \sqrt{\rho_\alpha(\kappa)}\, (\Gamma^{\mathscr N})^{\otimes2}\, \sqrt{\rho_\alpha(\kappa)} = \sum_{r,s=0}^1 |\psi_{rs}(\kappa)\rangle \!\langle\psi_{rs}(\kappa)| \end{align}\] is an orthogonal sum of rank-one operators.
Let \[\begin{align} \lambda_{rs}(\kappa) := \langle\psi_{rs}(\kappa)|\psi_{rs}(\kappa)\rangle. \end{align}\] Explicitly, \[\begin{align} \lambda_{00}(\kappa) &= p_{00} + \eta(p_{01}+p_{10}) + \eta^2p_{11} = \ell^2+\gamma^2\kappa, \tag{69} \\ \lambda_{01}(\kappa) &= \gamma(p_{01}+\eta p_{11}) = \gamma(u\ell-\gamma\kappa), \tag{70} \\ \lambda_{10}(\kappa) &= \gamma(p_{10}+\eta p_{11}) = \gamma(u\ell-\gamma\kappa), \tag{71} \\ \lambda_{11}(\kappa) &= \gamma^2p_{11} = \gamma^2(u^2+\kappa). \tag{72} \end{align}\] Here and below, the argument \(\kappa\) of the \(p_{ij}\)’s is suppressed.
Since \[\begin{align} \bigl( |\psi_{rs}\rangle\!\langle\psi_{rs}| \bigr)^\alpha = \lambda_{rs}^{\alpha-1} |\psi_{rs}\rangle\!\langle\psi_{rs}|, \end{align}\] the operator after tracing out the two input systems is diagonal: \[\begin{align} T_\alpha^{\mathscr N^{\otimes2}} \bigl(\rho_\alpha(\kappa)\bigr) = \sum_{i,j=0}^1 w_{ij}(\kappa)|ij\rangle\!\langle ij|. \label{eq:AD-two-copy-T} \end{align}\tag{73}\] The four diagonal entries are \[\begin{align} w_{00}(\kappa) &= p_{00}^{2-\alpha}\lambda_{00}^{\alpha-1} + \gamma p_{01}^{2-\alpha}\lambda_{01}^{\alpha-1} + \gamma p_{10}^{2-\alpha}\lambda_{10}^{\alpha-1} + \gamma^2p_{11}^{2-\alpha}\lambda_{11}^{\alpha-1}, \tag{74} \\ w_{01}(\kappa) &= \eta p_{01}^{2-\alpha}\lambda_{00}^{\alpha-1} + \gamma\eta p_{11}^{2-\alpha}\lambda_{10}^{\alpha-1}, \tag{75} \\ w_{10}(\kappa) &= \eta p_{10}^{2-\alpha}\lambda_{00}^{\alpha-1} + \gamma\eta p_{11}^{2-\alpha}\lambda_{01}^{\alpha-1}, \tag{76} \\ w_{11}(\kappa) &= \eta^2p_{11}^{2-\alpha}\lambda_{00}^{\alpha-1}. \tag{77} \end{align}\] For example, the four terms in \(w_{00}\) correspond, respectively, to no jumps, a jump in the second use, a jump in the first use, and jumps in both uses. The two terms in \(w_{01}\) correspond to the input \(01\) with no jump and the input \(11\) with a jump in the first use.
It follows that \[\begin{align} Q_\alpha(\rho_\alpha(\kappa);\mathscr N^{\otimes2}) = \sum_{i,j=0}^1 w_{ij}(\kappa)^{1/\alpha}. \label{eq:AD-Q-two-copy-weights} \end{align}\tag{78}\] At \(\kappa=0\), \[\begin{align} \lambda_{00}(0)&=\ell^2, & \lambda_{01}(0)&=\gamma u\ell, \notag\\ \lambda_{10}(0)&=\gamma u\ell, & \lambda_{11}(0)&=\gamma^2u^2, \end{align}\] and the product structure gives \[\begin{align} w_{ij}(0)=w_iw_j, \qquad i,j\in\{0,1\}. \label{eq:AD-w-product} \end{align}\tag{79}\]
Define \[\begin{align} a_0 &:= \gamma t^{2-\alpha}\ell^{\alpha-2} - \gamma^\alpha, & a_1 &:= \gamma\eta u^{2-\alpha}\ell^{\alpha-2} = \frac{\gamma w_1}{\ell}, \tag{80} \\ b_0 &:= t^{1-\alpha}\ell^{\alpha-1} - \gamma^\alpha, & b_1 &:= -\eta u^{1-\alpha}\ell^{\alpha-1} = -\frac{w_1}{u}. \tag{81} \end{align}\]
Lemma 12 (Factorized first variation). The one-copy derivatives satisfy \[\begin{align} w_i'(t) = (\alpha-1)a_i+(2-\alpha)b_i, \qquad i\in\{0,1\}. \label{eq:split-direct} \end{align}\qquad{(6)}\] Moreover, along the correlated path 63 , \[\begin{align} \left. \frac{\mathrm d}{\mathrm d\kappa} w_{ij}(\kappa) \right|_{\kappa=0} = (\alpha-1)a_i a_j + (2-\alpha)b_i b_j, \qquad i,j\in\{0,1\}. \label{eq:AD-factorized-wdot} \end{align}\qquad{(7)}\] Equivalently, \[\begin{align} \left. \frac{\mathrm d}{\mathrm d\kappa} \begin{pmatrix} w_{00}(\kappa)&w_{01}(\kappa)\\ w_{10}(\kappa)&w_{11}(\kappa) \end{pmatrix} \right|_{\kappa=0} = (\alpha-1) \begin{pmatrix}a_0\\a_1\end{pmatrix} \begin{pmatrix}a_0&a_1\end{pmatrix} + (2-\alpha) \begin{pmatrix}b_0\\b_1\end{pmatrix} \begin{pmatrix}b_0&b_1\end{pmatrix}. \label{eq:wdot-direct} \end{align}\qquad{(8)}\]
Proof. We give a factorized derivation of the identity.
Let \[\begin{align} r_0:=t, \qquad r_1:=u, \qquad d_0:=1, \qquad d_1:=-1. \end{align}\] Thus \[\begin{align} \frac{\mathrm d}{\mathrm dt}r_x=d_x, \qquad p_{xy}(\kappa)=r_xr_y+\kappa d_xd_y. \label{eq:AD-correlation-factorization} \end{align}\tag{82}\]
For a single channel use, let \(s=0\) denote the no-jump branch and \(s=1\) the jump branch. Denote by \(c_s(x)\) the squared branch amplitude for input \(x\): \[\begin{align} c_0(0)=1, \qquad c_0(1)=\eta, \qquad c_1(0)=0, \qquad c_1(1)=\gamma. \end{align}\] The corresponding branch eigenvalues and their variations in the direction \(d=(1,-1)\) are \[\begin{align} \lambda_s &:= \sum_{x=0}^1 r_xc_s(x), & \delta\lambda_s &:= \sum_{x=0}^1 d_xc_s(x). \end{align}\] Thus \[\begin{align} \lambda_0=\ell, \qquad \lambda_1=\gamma u, \qquad \delta\lambda_0=\gamma, \qquad \delta\lambda_1=-\gamma. \label{eq:AD-single-branch-data} \end{align}\tag{83}\]
The input-branch pairs producing output \(i\) are \[\begin{align} \mathcal{E}_0 = \{(0,0),(1,1)\}, \qquad \mathcal{E}_1 = \{(1,0)\}. \end{align}\] Accordingly, \[\begin{align} w_i = \sum_{(x,s)\in\mathcal{E}_i} r_x^{2-\alpha} c_s(x) \lambda_s^{\alpha-1}. \label{eq:AD-w-branch-form} \end{align}\tag{84}\] Differentiating 84 with respect to \(t\) gives \[\begin{align} w_i' &= (\alpha-1) \sum_{(x,s)\in\mathcal{E}_i} r_x^{2-\alpha} c_s(x) \lambda_s^{\alpha-2} \delta\lambda_s \notag\\ &\quad+ (2-\alpha) \sum_{(x,s)\in\mathcal{E}_i} d_xr_x^{1-\alpha} c_s(x) \lambda_s^{\alpha-1}. \label{eq:AD-one-copy-factorized-derivative} \end{align}\tag{85}\] The first sum equals \(a_i\), and the second equals \(b_i\). Evaluating the three possible pairs \((0,0),(1,1),(1,0)\) gives exactly 80 –81 , proving ?? .
For two channel uses, define the branch-pair eigenvalue \[\begin{align} \lambda_{s_1s_2}^{(2)}(\kappa) := \sum_{x,y=0}^1 p_{xy}(\kappa)c_{s_1}(x)c_{s_2}(y). \end{align}\] Using 82 , we find \[\begin{align} \lambda_{s_1s_2}^{(2)}(0) &= \lambda_{s_1}\lambda_{s_2}, \tag{86} \\ \left. \frac{\mathrm d}{\mathrm d\kappa} \lambda_{s_1s_2}^{(2)}(\kappa) \right|_{\kappa=0} &= \delta\lambda_{s_1}\delta\lambda_{s_2}. \tag{87} \end{align}\] The two-copy output weight has the branch representation \[\begin{align} w_{ij}(\kappa) = \sum_{\substack{(x,s_1)\in\mathcal{E}_i\\ (y,s_2)\in\mathcal{E}_j}} p_{xy}(\kappa)^{2-\alpha} c_{s_1}(x)c_{s_2}(y) \left( \lambda_{s_1s_2}^{(2)}(\kappa) \right)^{\alpha-1}. \label{eq:AD-wij-branch-form} \end{align}\tag{88}\] Differentiating at \(\kappa=0\), the variation of the branch eigenvalue contributes \[\begin{align} &(\alpha-1) \sum_{\substack{(x,s_1)\in\mathcal{E}_i\\ (y,s_2)\in\mathcal{E}_j}} (r_xr_y)^{2-\alpha} c_{s_1}(x)c_{s_2}(y) (\lambda_{s_1}\lambda_{s_2})^{\alpha-2} \delta\lambda_{s_1}\delta\lambda_{s_2} \notag\\ &\qquad= (\alpha-1)a_i a_j. \end{align}\] The variation of the input probabilities contributes \[\begin{align} &(2-\alpha) \sum_{\substack{(x,s_1)\in\mathcal{E}_i\\ (y,s_2)\in\mathcal{E}_j}} d_xd_y(r_xr_y)^{1-\alpha} c_{s_1}(x)c_{s_2}(y) (\lambda_{s_1}\lambda_{s_2})^{\alpha-1} \notag\\ &\qquad= (2-\alpha)b_i b_j. \end{align}\] This proves ?? . ◻
Define \[\begin{align} x_\alpha &:= w_0^{1/\alpha-1}a_0 + w_1^{1/\alpha-1}a_1, \tag{89} \\ y_\alpha &:= w_0^{1/\alpha-1}b_0 + w_1^{1/\alpha-1}b_1. \tag{90} \end{align}\] Since \(t=t_\alpha\) is an interior minimizer, stationarity and ?? give \[\begin{align} 0 = q_\alpha'(t) = \frac{1}{\alpha} \left[ (\alpha-1)x_\alpha + (2-\alpha)y_\alpha \right]. \end{align}\] Thus \[\begin{align} (\alpha-1)x_\alpha + (2-\alpha)y_\alpha = 0. \label{eq:xy-stationarity-direct} \end{align}\tag{91}\]
Using 79 and Lemma 12, we obtain \[\begin{align} &\left. \frac{\mathrm d}{\mathrm d\kappa} Q_\alpha(\rho_\alpha(\kappa);\mathscr N^{\otimes2}) \right|_{\kappa=0} \notag\\ &\quad= \frac{1}{\alpha} \sum_{i,j=0}^1 (w_iw_j)^{1/\alpha-1} \left[ (\alpha-1)a_ia_j + (2-\alpha)b_ib_j \right] \notag\\ &\quad= \frac{1}{\alpha} \left[ (\alpha-1)x_\alpha^2 + (2-\alpha)y_\alpha^2 \right]. \label{eq:response-pre-direct} \end{align}\tag{92}\] Eliminating \(y_\alpha\) by 91 gives \[\begin{align} \left. \frac{\mathrm d}{\mathrm d\kappa} Q_\alpha(\rho_\alpha(\kappa);\mathscr N^{\otimes2}) \right|_{\kappa=0} = \frac{\alpha-1}{\alpha(2-\alpha)} x_\alpha^2. \label{eq:response-x-direct} \end{align}\tag{93}\]
It remains to prove that \(x_\alpha\neq0\). Suppose, to the contrary, that \(x_\alpha=0\). Since \(2-\alpha>0\), 91 then gives \(y_\alpha=0\). Therefore the strictly positive vector \[\begin{align} z := \left( w_0^{1/\alpha-1}, w_1^{1/\alpha-1} \right) \end{align}\] is orthogonal to both \((a_0,a_1)\) and \((b_0,b_1)\). In two dimensions this would imply that these latter two vectors are linearly dependent.
On the other hand, direct simplification gives \[\begin{align} a_0b_1-a_1b_0 &= \eta\gamma^\alpha u^{1-\alpha}\ell^{\alpha-2} - \eta\gamma t^{1-\alpha}u^{1-\alpha} \ell^{2\alpha-3} \notag\\ &= \eta\gamma\, t^{1-\alpha}u^{1-\alpha} \ell^{\alpha-2} \left[ (\gamma t)^{\alpha-1} - \ell^{\alpha-1} \right]. \label{eq:det-direct} \end{align}\tag{94}\] All factors outside the brackets are strictly positive. Moreover, \[\begin{align} 0<\gamma t<\ell \end{align}\] and \(\alpha-1<0\), so \[\begin{align} (\gamma t)^{\alpha-1} > \ell^{\alpha-1}. \end{align}\] Hence \[\begin{align} a_0b_1-a_1b_0>0, \end{align}\] contradicting linear dependence. Therefore \(x_\alpha\neq0\).
Since \(0<\alpha<1\), we conclude from 93 that \[\begin{align} \left. \frac{\mathrm d}{\mathrm d\kappa} Q_\alpha(\rho_\alpha(\kappa);\mathscr N^{\otimes2}) \right|_{\kappa=0} <0. \label{eq:AD-negative-response} \end{align}\tag{95}\]
Theorem 13 (Strict superadditivity of amplitude damping). Let \(\mathscr N\) be the amplitude-damping channel 44 with \(0<\gamma<1\). Then \[\begin{align} I_\alpha(\mathscr N^{\otimes2}) > 2I_\alpha(\mathscr N) \qquad \forall\,0<\alpha<1. \end{align}\] The strict improvement is witnessed by the diagonal, classically correlated path \(\rho_\alpha(\kappa)\) in 63 .
Proof. By 95 , for every sufficiently small \(\kappa>0\), \[\begin{align} Q_\alpha(\rho_\alpha(\kappa);\mathscr N^{\otimes2}) < Q_\alpha(\rho_\alpha(0);\mathscr N^{\otimes2}) = \bigl(K^{(1)}(\alpha)\bigr)^2. \end{align}\] Therefore \[\begin{align} K^{(2)}(\alpha) &\leq Q_\alpha(\rho_\alpha(\kappa);\mathscr N^{\otimes2}) < \bigl(K^{(1)}(\alpha)\bigr)^2. \end{align}\] Finally, because \(\alpha/(\alpha-1)<0\), \[\begin{align} I_\alpha(\mathscr N^{\otimes2}) &= \frac{\alpha}{\alpha-1}\log K^{(2)}(\alpha) \notag\\ &> \frac{\alpha}{\alpha-1} \log\bigl(K^{(1)}(\alpha)^2\bigr) = 2I_\alpha(\mathscr N). \end{align}\] ◻
Except for the special channels (Proposition 3) and \(\alpha = 1,2\), the Petz–Rényi information \(I_{\alpha}(\mathscr{N})\) can be strictly superadditive as shown in Sections 10 and 11. This means that the achievable random coding exponent \(E_{\mathrm{r}}(R;\mathscr{N})\) has a multi-letter expression in general.
Define the regularized Petz channel information as \[\begin{align} I_{\alpha}^{\infty}(\mathscr{N}) &\coloneq \sup_{n\in\mathbb{N}} \frac{1}{n} I_{\alpha}\left(\mathscr{N}^{\otimes n}\right) = \lim_{n\to\infty} \frac{1}{n} I_{\alpha}\left(\mathscr{N}^{\otimes n}\right), \end{align}\] where the last equality follows from the superadditive sequence and Fekete’s lemma.
Lemma 14 (, [46], [33]). Let \(\rho\) and \(\sigma\) be density operators. Then, \[\begin{align} D_{2-\frac{1}{\alpha}}(\rho\Vert\sigma) \leq \widetilde{D}_{\alpha}(\rho\Vert\sigma), \quad \forall\, \alpha \geq \tfrac{1}{2}. \end{align}\]
Equivalently, \[\begin{align} D_{\alpha}(\rho\Vert\sigma) \leq \widetilde{D}_{\frac{1}{2-\alpha}}(\rho\Vert\sigma), \quad \forall \alpha \in [0,2]. \end{align}\]
Proposition 15 (Single-letter bounds). Let \(\mathscr{N}_{\mathsf{A}\to\mathsf{B}}\) be a quantum channel. We have \[\begin{align} I_{\alpha}(\mathscr{N}) \leq I_{\alpha}^{\infty}(\mathscr{N}) \leq \widetilde{I}_{\frac{1}{2-\alpha}}(\mathscr{N}) \leq {I}_{\frac{1}{2-\alpha}}(\mathscr{N}), \quad \forall\,\alpha \in [0,2]. \end{align}\]
Proof. For any integer \(n\) and \(\alpha \in [0,2]\), Lemma 14 implies that \[\begin{align} I_{\alpha}(\mathscr{N}) \leq \frac{1}{n}I_{\alpha}(\mathscr{N}^{\otimes n}) \leq \frac{1}{n}\widetilde{I}_{\frac{1}{2-\alpha}}(\mathscr{N}^{\otimes n}) = \widetilde{I}_{\frac{1}{2-\alpha}}(\mathscr{N}), \end{align}\] where the last equality follows from the additivity of \(\widetilde{I}_{\beta}(\mathscr{N})\), \(\beta = \frac{1}{2-\alpha}\in[\frac{1}{2},1)\) [19] and \(\beta = \frac{1}{2-\alpha}\geq 1\) [17]. Taking \(n\to\infty\) completes the proof. ◻
The single-letter upper bound in terms of \(\widetilde{I}_{\frac{1}{2-\alpha}}(\mathscr{N})\) is a double-state optimization; see 23 . One may further relax it to a more computationally feasible one-state optimization of the Petz form \({I}_{\frac{1}{2-\alpha}}(\mathscr{N})\). Below, we show that the sandwiched Rényi information \(\widetilde{I}_{\alpha}(\mathscr{N})\) admits a one-state optimization expression for quantum-classical channels.
Proposition 16 (One-state optimization for quantum-classical channels). Let \(\alpha\in[\frac{1}{2},1)\cup(1,+\infty)\) and let \(\mathscr{M}=\{M_y\}_{y\in\mathsf{Y}}\) be a finite POVM. Define \[\widetilde{Q}_y(\rho):= \operatorname{Tr}\left[ \left( \rho^{1/(2\alpha)} M_y^{\top} \rho^{1/(2\alpha)} \right)^\alpha \right].\] Then \[\widetilde{I}_\alpha(\mathscr{M}) = \begin{cases} \displaystyle \frac{\alpha}{\alpha-1} \log \inf_{\rho} \sum_{y\in\mathsf{Y}} \widetilde{Q}_y(\rho)^{1/\alpha}, & \alpha\in[\frac{1}{2},1), \\[10pt] \displaystyle \frac{\alpha}{\alpha-1} \log \sup_{\rho} \sum_{y\in\mathsf{Y}} \widetilde{Q}_y(\rho)^{1/\alpha}, & \alpha\in(1,+\infty). \end{cases} \label{eq:K-reduction}\qquad{(9)}\] For fixed \(\rho\), an optimizing auxiliary output distribution is \[q_y^\star(\rho) = \frac{\widetilde{Q}_y(\rho)^{1/\alpha}}{\sum_{\bar{y}\in\mathsf{Y}} \widetilde{Q}_{\bar{y}}(\rho)^{1/\alpha}}, \label{eq:sigma-opt}\qquad{(10)}\] with the convention that \(q_y^\star(\rho)=0\) whenever \(\widetilde{Q}_y(\rho)=0\).
Proof. Write \(\sigma_{\mathsf{Y}}=\sum_y q_y|y\rangle\langle y|_{\mathsf{Y}}\). Since the output is classical, we have \[\begin{align} &\operatorname{Tr}\left[ \left( (\rho\otimes\sigma_{\mathsf{Y}})^{\frac{1-\alpha}{2\alpha}} \omega_{\mathsf{R}\mathsf{Y}}^{\rho} (\rho\otimes\sigma_{\mathsf{Y}})^{\frac{1-\alpha}{2\alpha}} \right)^\alpha \right] \\ &\qquad = \sum_y q_y^{1-\alpha} \operatorname{Tr}\left[ \left( \rho^{\frac{1-\alpha}{2\alpha}} \rho^{1/2}M_y^{\top}\rho^{1/2} \rho^{\frac{1-\alpha}{2\alpha}} \right)^\alpha \right] \\ &\qquad = \sum_y q_y^{1-\alpha}\widetilde{Q}_y(\rho). \end{align}\] For \(\alpha\in[\frac{1}{2},1)\), since \(1/(\alpha-1)<0\), minimizing the divergence over \(q\) is equivalent to maximizing \(\sum_yq_y^{1-\alpha}\widetilde{Q}_y(\rho)\) over the probability simplex. For \(\alpha>1\), since \(1/(\alpha-1)>0\), minimizing the divergence is equivalent to minimizing the same expression. In the latter case, the support condition requires \(q_y>0\) whenever \(\widetilde{Q}_y(\rho)>0\); terms for which \(\widetilde{Q}_y(\rho)=0\) may be omitted.
The scalar objective is concave in \(q\) for \(\alpha<1\) and convex in \(q\) for \(\alpha>1\). In both cases, the optimality condition gives \[q_y\propto\widetilde{Q}_y(\rho)^{1/\alpha}\] for \(y:\widetilde{Q}_y(\rho)>0\) and the optimal value is \[\left( \sum_y\widetilde{Q}_y(\rho)^{1/\alpha} \right)^\alpha.\] Hence, for either \(\alpha\in[\frac{1}{2},1)\) or \(\alpha>1\), \[\inf_{\sigma_{\mathsf{Y}}} \widetilde{D}_\alpha \left( \omega_{\mathsf{R}\mathsf{Y}}^{\rho} \middle\| \rho\otimes\sigma_{\mathsf{Y}} \right) = \frac{\alpha}{\alpha-1} \log \sum_y\widetilde{Q}_y(\rho)^{1/\alpha}.\] For \(\alpha<1\), the coefficient \(\alpha/(\alpha-1)\) is negative, so the outer supremum over \(\rho\) becomes the infimum in ?? . For \(\alpha>1\), this coefficient is positive, so the outer supremum remains the supremum in ?? . ◻