Metric Properties: From \(S\)-Divergence to Quantum Jensen Divergence


Abstract

We extend the trace-logarithmic \(S\)-divergence from matrices to tracial \(C^*\)-algebras and finite von Neumann algebras, and show that its square root defines a metric on the invertible positive cone. We also prove an integral representation of the quantum Jensen–Shannon divergence in terms of shifted trace-log distances, implying metricity of its square root on the full positive cone in the same tracial framework. In the matrix case, we answer two questions of Virosztek [1] on Hilbertianity. Finally, we show that symmetric quantum Jensen divergences generated by non-affine operator convex functions yield metrics in the tracial setting via a Nevanlinna–Stieltjes type representation of the derivative, which generalizes a result of Carlen, Lieb and Seiringer.

1 Introduction↩︎

Throughout this paper, let \(M_n^+(\mathbb{C})\) denote the cone of all \(n\times n\) positive semidefinite matrices, and let \(M_n^{++}(\mathbb{C})\) denote the cone of all \(n\times n\) positive definite matrices.

Let \((\mathcal{A},\tau)\) be a unital \(C^*\)-algebra equipped with a faithful tracial state \[\tau(\mathbf{1})=1,\qquad \tau(x^*x)=0 \;\Rightarrow\;x=0,\qquad \text{and}\qquad \tau(xy)=\tau(yx)\;\;\text{for all }x,y\in\mathcal{A}.\] We also use the associated noncommutative \(L^2\)-norm \(\|x\|_{2,\tau}:=\tau(x^*x)^{1/2},\) whenever \(\tau\) is understood from context. We write \(\mathcal{A}_{+}\) for the positive cone of \(\mathcal{A}\), and \(\mathcal{A}_{++}\) for its invertible positive cone. We also write \[\mathcal{S}_\tau:=\{\rho\in\mathcal{A}_{+}:\tau(\rho)=1\},\qquad \mathcal{S}_{\tau,++}:=\mathcal{S}_\tau\cap \mathcal{A}_{++}.\]

Let \((\mathcal{M},\tau)\) be a finite von Neumann algebra with a faithful normal tracial state \(\tau\), and denote by \(\mathcal{M}^\times\) the group of invertible elements of \(\mathcal{M}\). Write \(\mathcal{M}_{\mathrm{sa}}=\{x\in\mathcal{M}:\;x^*=x\}\), and denote by \(\mathcal{M}_{+}\) and \(\mathcal{M}_{++}\) the positive cone and the invertible positive cone of \(\mathcal{M}\), respectively.

1.1 The \(S\)-divergence↩︎

The \(S\)-divergence was introduced by Sra [2] in the study of the open convex cone \(M_n^{++}(\mathbb{C})\). In contrast, the affine-invariant Riemannian distance \[\delta_R(A,B)=\bigl\|\log\bigl(B^{-1/2}AB^{-1/2}\bigr)\bigr\|_F, \qquad A,B\in M_n^{++}(\mathbb{C}),\] where \(\|\cdot\|_F\) denotes the Frobenius norm, arises from the nonpositively curved Riemannian geometry of \(M_n^{++}(\mathbb{C})\) (see, e.g., [3] and [4]). However, outside this specific geometric framework, there is in general no canonical “natural” choice of distance on \(M_n^{++}(\mathbb{C})\).

Motivated by ideas from convex optimization and information geometry [5], [6], Sra proposed the symmetric trace-logarithmic expression \[\label{eq:sra} \delta_S^2(A,B) =\log\det\!\Bigl(\frac{A+B}{2}\Bigr)-\frac{1}{2}\log\det A-\frac{1}{2}\log\det B, \qquad A,B\in M_n^{++}(\mathbb{C}),\tag{1}\] also known as the Jensen–Bregman LogDet divergence (or symmetric Stein divergence); see [2], [7], [8]. A key feature proved in [2] is that \(\delta_S:=\sqrt{\delta_S^2}\) satisfies the triangle inequality on \(M_n^{++}(\mathbb{C})\), hence defines a genuine metric.

Theorem 1 (Sra). Let \(A,B\in M_n^{++}(\mathbb{C})\). Define \(\delta_S\) by 1 . Then \(\delta_S\) is a metric on \(M_n^{++}(\mathbb{C})\). Moreover, it does not admit any Hilbert space embedding for any \(n\ge 2\).

Besides its intrinsic geometric interest, \(\delta_S\) is computationally attractive and has been used effectively in applications such as similarity search for covariance descriptors [6], which further motivates the study of trace-logarithmic metrics beyond finite dimensions.

There is a substantial literature extending log-determinant divergences to infinite-dimensional operator settings. For example, Minh [9], [10] developed infinite-dimensional LogDet divergences for unitized trace-class and Hilbert–Schmidt perturbations via extended Fredholm-type determinants. From the viewpoint of finite von Neumann algebras, the Fuglede–Kadison determinant \(\Delta\) [11] provides a canonical replacement for \(\det\), and “distance-like” functionals built from \(\Delta\) and functional calculus have been considered in connection with isometry problems (see, e.g., Gaál–Nagy–Szokol [12]).

In the matrix case, \(\log\det X=\mathop{\mathrm{Tr}}(\log X)\), suggesting a natural extension to tracial \(C^*\)-algebras and finite von Neumann algebras by replacing \(\mathop{\mathrm{Tr}}\) with a faithful tracial state. This raises the question of whether the \(S\)-divergence remains a metric in the \(C^*\)-algebraic setting. In this paper, we answer this in the affirmative.

Theorem 2. Let \((\mathcal{A},\tau)\) be a unital \(C^*\)-algebra with a faithful tracial state. For \(A,B\in\mathcal{A}_{++}\) define \[\label{eq:d95tau} d^2_\tau(A,B) :=\tau\!\left(\log\!\left(\frac{A+B}{2}\right)\right)-\frac{1}{2}\tau(\log A)-\frac{1}{2}\tau(\log B).\qquad{(1)}\] Then \(d_\tau(A,B)\) is a metric on \(\mathcal{A}_{++}\). In particular, its restriction to \(\mathcal{S}_{\tau,++}\) is a metric.

We call \(d_\tau\) the trace-log distance (or \(S\)-divergence metric) associated with \((\mathcal{A},\tau)\).

1.2 Quantum Jensen–Shannon divergence↩︎

Let \(\eta(x)=x\log x\) for \(x>0\) and \(\eta(0)=0\) (continuous extension). The quantum Jensen–Shannon divergence (QJSD) is a quantum analogue of the classical Jensen–Shannon divergence, systematically introduced by Majtey–Lamberti–Prato [13]. A standard matrix form is \[\label{eq:qjsd-def-matrix} J_{\eta}(A,B) :=\frac{1}{2}\mathop{\mathrm{Tr}}(\eta(A))+\frac{1}{2}\mathop{\mathrm{Tr}}(\eta(B))-\mathop{\mathrm{Tr}}\!\left(\eta\!\left(\frac{A+B}{2}\right)\right), \qquad A,B\in M_n^{+}(\mathbb{C}).\tag{2}\] The QJSD is nonnegative, symmetric, and bounded; compared with the quantum relative entropy it is often regarded as smoother in applications. It has found broad applications beyond quantum theory [14][16], including complex network theory [17], [18], pattern recognition [19], graph theory [20], and chemical physics [21].

Virosztek [1] proved that the square root of the quantum Jensen–Shannon divergence defines a genuine metric on the full cone of positive matrices, in every dimension.

Theorem 3 (Virosztek). Let \(A,B\in M_n^{+}(\mathbb{C})\). Define \(J_\eta(A,B)\) by 2 . Then \(\sqrt{J_{\eta}(A,B)}\) is a metric on \(M_n^{+}(\mathbb{C})\).

Having established metricity, Virosztek [1] further asked whether this metric is induced by a Hilbert space embedding.

Problem 4 (Virosztek). Is the metric induced by the quantum Jensen–Shannon divergence Hilbertian on \(M_n^+(\mathbb{C})\) for \(n\ge 2\)?

A related question [1] concerns the physically important subset of density matrices.

Problem 5 (Virosztek). Is the QJSD metric Hilbertian on the space of \(n\times n\) density matrices for \(n\ge 3\)?

Our next theorem answers both questions in the negative.

Theorem 6. Let \(\sqrt{J_{\eta}}\) be the QJSD metric on \(M_n^{+}(\mathbb{C})\) defined by 2 .

  1. If \(n\ge 2\), then \(\sqrt{J_{\eta}}\) is not Hilbertian on \(M_n^{++}(\mathbb{C})\). Consequently, it is not Hilbertian on \(M_n^{+}(\mathbb{C})\).

  2. If \(n\ge 3\), then \(\sqrt{J_{\eta}}\) is not Hilbertian on the density matrices \(\mathcal{D}_n:=\{\rho\in M_n^{+}(\mathbb{C}):\mathop{\mathrm{Tr}}(\rho)=1\}.\)

In addition to these finite-dimensional results, we also show that Virosztek’s metricity theorem extends to the setting of tracial \(C^*\)-algebras.

Theorem 7. Let \((\mathcal{A},\tau)\) be a unital \(C^*\)-algebra with a faithful tracial state, and let \(\eta(x)=x\log x\) with \(\eta(0)=0\). For \(A,B\in\mathcal{A}_{+}\) define \[J_{\tau,\eta}(A,B):=\frac{1}{2}\tau(\eta(A))+\frac{1}{2}\tau(\eta(B))-\tau\!\left(\eta\!\left(\frac{A+B}{2}\right)\right).\] Then the square root \(\sqrt{J_{\tau,\eta}(A,B)}\) is a metric on \(\mathcal{A}_{+}\). In particular, its restriction to the \(\tau\)-state space \(\mathcal{S}_\tau\) is a metric.

The QJSD fits into a broader family of Jensen-type divergences generated by convex (and, in the noncommutative setting, operator convex) functions. This viewpoint will be useful both conceptually and technically in later sections.

1.3 Quantum Jensen \(f\)-divergences and trace Jensen gaps↩︎

For a continuous function \(f:[0,\infty)\to\mathbb{R}\), the associated trace Jensen gap on \(M_n^{+}(\mathbb{C})\) is defined by \[J_f(A,B) :=\frac{1}{2}\,\mathrm{Tr}\bigl(f(A)\bigr)+\frac{1}{2}\,\mathrm{Tr}\bigl(f(B)\bigr) -\mathrm{Tr}\!\left(f\!\left(\frac{A+B}{2}\right)\right), \qquad A,B\in M_n^{+}(\mathbb{C}),\] whenever the terms are well-defined (this is automatic under bounded functional calculus). To ensure that \(J_f\ge 0\), we additionally assume that \(f\) is operator convex and not affine (otherwise, \(J_f\equiv 0\)).

Let \(\mathcal{S}(\mathbb{C}^2)\) denote the space of \(2\times 2\) density matrices. On the qubit state space \(\mathcal{S}(\mathbb{C}^2)\), the following strengthening was proved by Carlen–Lieb–Seiringer; see [1].

Theorem 8 (Carlen–Lieb–Seiringer). For an operator convex function \(f:[0,\infty)\to\mathbb{R}\), the square root of the symmetric quantum Jensen \(f\)-divergence defined by \[J_f(\rho,\sigma) :=\frac{1}{2}\bigl(\mathop{\mathrm{Tr}}f(\rho)\bigr)+\frac{1}{2}\bigl(\mathop{\mathrm{Tr}}f(\sigma)\bigr) -\mathop{\mathrm{Tr}}f\!\left(\frac{\rho+\sigma}{2}\right), \qquad (\rho,\sigma\in \mathcal{S}(\mathbb{C}^2)),\] is a true metric on \(\mathcal{S}(\mathbb{C}^2)\); moreover, it admits a Hilbert space embedding.

Beyond the qubit setting, however, such an embedding phenomenon cannot hold in full generality, see [1]. Indeed, the cone of \(2\times 2\) positive definite matrices equipped with the Jensen divergence corresponding to the operator convex function \(x\mapsto -\log x\) is known not to admit any Hilbert space embedding (Theorem 1). This indicates that the Carlen–Lieb–Seiringer theorem is optimal: one cannot, in general, go beyond the qubit state space when seeking Hilbertianity simultaneously for all operator convex generators. (To be precise, \(x\mapsto -\log x\) is operator convex only on \((0,\infty)\) and not on \([0,\infty)\), but the integral representation \[\log x=\int_{0}^{\infty}\left(\frac{1}{1+t}-\frac{1}{x+t}\right)\,dt\] shows that, from the viewpoint of negative definite kernels, it behaves similarly to operator convex functions on \([0,\infty)\).)

An important and challenging problem is whether the square root of the symmetric quantum Jensen \(f\)-divergence defines a metric on \(M_n(\mathbb{C})^{+}\). For \(f(x)=x\log x\), it reduces to Theorem 3. We next provide a general mechanism showing that operator convex Jensen generators still yield metrics (without claiming Hilbertianity in general) in the tracial \(C^*\)-algebra setting.

Theorem 9. Let \((\mathcal{A},\tau)\) be a unital \(C^*\)-algebra with a faithful tracial state. Let \(f:(0,\infty)\to\mathbb{R}\) be an operator convex function which is not affine, and assume that \(f\) extends continuously to \([0,\infty)\) with finite \(f(0)\). For \(A,B\in\mathcal{A}_{+}\) define \[J_{\tau,f}(A,B) :=\frac{1}{2}\tau\bigl(f(A)\bigr)+\frac{1}{2}\tau\bigl(f(B)\bigr) -\tau\!\left(f\!\left(\frac{A+B}{2}\right)\right).\] Then the square root \(\sqrt{J_{\tau,f}}\) defines a metric on \(\mathcal{A}_{+}\). More precisely, there exist \(b\ge 0\) and a positive Borel measure \(\nu\) on \((0,\infty)\) satisfying \(\int_{(0,\infty)}\frac{1}{1+t}\,\,\mathrm{d}\nu(t)<\infty\) such that for all \(A,B\in\mathcal{A}_{+}\), \[\label{eq:opconvex-decomp-intro} J_{\tau,f}(A,B) =\frac{b}{8}\,\left\|A-B\right\|_{2,\tau}^{2} +\int_{(0,\infty)} t\,d_\tau(A+t\mathbf{1},B+t\mathbf{1})^2\,\,\mathrm{d}\nu(t).\qquad{(2)}\]

As an immediate corollary of Theorem 9, we have

Corollary 10. Let \(f:(0,\infty)\to\mathbb{R}\) be a operator convex function which is not affine, and assume that \(f\) extends continuously to \([0,\infty)\) with finite \(f(0)\). Then the square root of the symmetric quantum Jensen \(f\)-divergence defined by \[J_f(A,B) :=\frac{1}{2}\bigl(\mathop{\mathrm{Tr}}f(A)\bigr)+\frac{1}{2}\bigl(\mathop{\mathrm{Tr}}f(B)\bigr) -\mathop{\mathrm{Tr}}f\!\left(\frac{A+B}{2}\right), \qquad (A,B\in M_n^{+}(\mathbb{C})),\] is a true metric on \(M_n^{+}(\mathbb{C})\).

Organization of this paper. In Section 2, we recall the Fuglede–Kadison determinant and generalized singular value functions. In Sections 46 we prove the triangle inequality for \(d_\tau\) on \(\mathcal{M}_{++}\): we first embed the scalar divergence into a Hilbert space (Section 4), then establish a rearrangement lower bound using the Fack–Kosaki inequality (Section 5), and finally combine these ingredients with Minkowski’s inequality (Section 6). Section 7 transfers the von Neumann result back to tracial \(C^*\)-algebras via the GNS envelope and completes the proof of Theorem 2. Section 8 proves an integral representation for QJSD and deduces Theorem 7. Section 9 establishes the non-Hilbertianity results in Theorem 6 and provides an explicit numerical witness. Finally, Section 10 discusses an operator-convex extension mechanism for metricity of \(\sqrt{J_{\tau,f}}\).

Acknowledgments. The author is also grateful to his advisor, Lajos Molnár, for suggesting this problem. This work is supported by the China Scholarship Council, the Young Elite Scientists Sponsorship Program for PhD Students (China Association for Science and Technology), and the Fundamental Research Funds for the Central Universities at Xian Jiaotong University (Grant No. xzy022024045).

2 Von Neumann preliminaries: determinants and rearrangements↩︎

2.1 Fuglede–Kadison determinant↩︎

For an invertible element \(x\in \mathcal{M}^\times\), the Fuglede–Kadison determinant is defined by \[\Delta(x):=\exp\big(\tau(\log|x|)\big)\in (0,\infty).\] In the original work [11] (for finite factors) one shows that \(\Delta:\mathcal{M}^\times\to(0,\infty)\) is a continuous group homomorphism, hence multiplicative (see [11]): \[\label{eq:FK-mult} \Delta(xy)=\Delta(x)\Delta(y)\qquad (x,y\in \mathcal{M}^\times).\tag{3}\] The same conclusions extend to general finite von Neumann algebras with faithful normal tracial states, e.g.by reduction to factors via central decomposition; see [22], [23].

When \(x\in\mathcal{M}^\times\), we have \(\log\Delta(x)=\tau(\log|x|)\). Moreover, for \(X\in\mathcal{M}_{++}\) and \(r\in\mathbb{R}\), \[\label{eq:Delta-power} \Delta(X^r)=\Delta(X)^r,\tag{4}\] since \(\log\Delta(X^r)=\tau(\log(X^r))=r\,\tau(\log X)=r\log\Delta(X)\).

2.2 Generalized \(s\)-numbers and a trace formula↩︎

We recall \(\tau\)-measurability and generalized \(s\)-numbers in the sense of Fack–Kosaki [24]. Let \(\mathcal{M}\subset B(H)\) be a von Neumann algebra acting on a complex Hilbert space \(H\), and let \(\tau:\mathcal{M}_{+}\to[0,\infty]\) be a faithful normal trace. In our applications \(\tau\) will be finite and normalized, i.e.\(\tau(\mathbf{1})=1\).

Let \(T\) be a densely-defined closed operator on \(H\) (possibly unbounded) affiliated with \(\mathcal{M}\), and write \(\mathop{\mathrm{Dom}}(T)\) for its domain. Following [24], \(T\) is called \(\tau\)-measurable if for every \(\varepsilon>0\) there exists a projection \(E\in\mathcal{M}\) such that \[E(H)\subset\mathop{\mathrm{Dom}}(T)\qquad\text{and}\qquad \tau(\mathbf{1}-E)\le\varepsilon.\] Since \(\tau\) is finite in our setting, the algebra \(\widetilde{\mathcal{M}}\) of \(\tau\)-measurable operators coincides with the set of all densely-defined closed operators affiliated with \(\mathcal{M}\); see [24].

If \(T\) is affiliated with \(\mathcal{M}\), write \(T=U|T|\) for its polar decomposition and let \(E^{|T|}\) denote the spectral measure of \(|T|\). For \(\lambda\ge 0\), set \[d_{|T|}(\lambda):=\tau\!\left(E^{|T|}\big((\lambda,\infty)\big)\right).\] The generalized \(s\)-numbers of \(T\) are defined for \(t>0\) by \[\label{eq:mu-def} \mu_t(T):=\inf\{\lambda\ge 0:\;d_{|T|}(\lambda)\le t\},\tag{5}\] see [24]. For a bounded positive operator \(X\in\mathcal{M}_{+}\) we write \(\mu_X(t):=\mu_t(X)\) for \(t\in(0,1]\).

Assume now that \(\tau(\mathbf{1})=1\).

Remark 11 (Normalization of the trace). Throughout Sections 26 we work with a tracial state, so \(\tau(\mathbf{1})=1\) and the generalized singular value parameter \(t\) naturally ranges in \((0,1]\). If instead \(\tau\) is a finite faithful normal trace with \(\tau(\mathbf{1})=\theta\in(0,\infty)\), one may normalize it by \(\tau_0:=\theta^{-1}\tau\). All rearrangement identities then hold with \(\tau_0\) (and hence with integration range \((0,1]\)), or equivalently one may keep \(\tau\) and replace \(\int_0^1(\cdot)\,dt\) by \(\int_0^\theta(\cdot)\,dt\). Since metricity is unchanged by rescaling the trace (it only rescales the distance by a constant factor), we restrict to \(\tau(\mathbf{1})=1\) for notational simplicity.

Remark 12 (Endpoint behavior of \(\mu_X\)). Assume \(\tau(\mathbf{1})=1\) and let \(X\in\mathcal{M}_{+}\) be bounded. Then \(t\mapsto \mu_X(t)\) is decreasing and right-continuous on \((0,1]\). Moreover, one always has \(\mu_X(1)=0\) (since \(d_X(\lambda)\le 1\) for all \(\lambda\ge 0\)).

When \(X\in\mathcal{M}_{++}\), there exists \(m>0\) such that \(\mu_X(t)\ge m\) for all \(t\in(0,1)\). Accordingly, integrals of the form \(\int_0^1 \varphi(\mu_X(t))\,dt\) are understood as Lebesgue integrals on \((0,1)\), and the value at \(t=1\) is irrelevant.

Let \(f:[0,\infty)\to\mathbb{R}\) be continuous and increasing with \(f(0)=0\). For every \(\tau\)-measurable operator \(T\), Fack–Kosaki proved the trace formula \[\label{eq:FK-Cor28} \tau\big(f(|T|)\big)=\int_0^\infty f\big(\mu_t(T)\big)\,\,\mathrm{d}t,\tag{6}\] see [24]. Moreover, if \(X\in\mathcal{M}_{+}\) is bounded, then \(\mu_t(X)=0\) for all \(t>\tau(\mathbf{1})=1\) (see [24]), so 6 reduces to \[\label{eq:trace-rearr} \tau(f(X))=\int_0^1 f\big(\mu_X(t)\big)\,\,\mathrm{d}t.\tag{7}\]

Remark 13 (Shifting constants). Assume \(\tau(\mathbf{1})=1\) and let \(X\in\mathcal{M}_{+}\) be bounded. If \(f:[0,\infty)\to\mathbb{R}\) is continuous and increasing (not necessarily satisfying \(f(0)=0\)), then \[\label{eq:trace-rearr-shifted} \tau(f(X))=\int_0^1 f\big(\mu_X(t)\big)\,\,\mathrm{d}t.\tag{8}\]

Finally, if \(X\in\mathcal{M}_{++}\) then \(X\ge m\mathbf{1}\) for some \(m>0\), hence \(\mu_X(t)\ge m\) for all \(t\in(0,1)\) (cf.Remark 12). Define a continuous increasing function \(g_m:[0,\infty)\to\mathbb{R}\) by \[g_m(u):= \begin{cases} 0, & 0\le u\le m,\\[0.2em] \log(u/m), & u\ge m. \end{cases}\] Then \(g_m(0)=0\) and, since \(\sigma(X)\subset[m,\|X\|]\), functional calculus gives \[g_m(X)=\log X-(\log m)\mathbf{1}, \qquad g_m(\mu_X(t))=\log(\mu_X(t))-\log m\quad \text{for a.e. }t\in(0,1).\] Applying 7 to \(g_m\) yields \[\tau(\log X)-\log m =\tau(g_m(X)) =\int_0^1 g_m(\mu_X(t))\,\,\mathrm{d}t =\int_0^1 \bigl(\log(\mu_X(t))-\log m\bigr)\,\,\mathrm{d}t,\] and therefore \[\label{eq:trace-log-mu} \tau(\log X)=\int_{(0,1)} \log\big(\mu_X(t)\big)\,\,\mathrm{d}t,\qquad X\in\mathcal{M}_{++}.\tag{9}\] Here the integral is understood in the Lebesgue sense; the integrand is finite for \(t\in(0,1)\) since \(\mu_X(t)\ge m>0\) on \((0,1)\) (cf.Remark 12).

Fack–Kosaki inequality. A key inequality we will use is [24]: for \(X,Y\in\mathcal{M}_{+}\) and every continuous convex increasing function \(f:[0,\infty)\to\mathbb{R}\), \[\label{eq:FK} \int_0^u f\big(\mu_{X+Y}(t)\big)\,\,\mathrm{d}t \le \int_0^u f\big(\mu_X(t)+\mu_Y(t)\big)\,\,\mathrm{d}t, \qquad 0<u\le 1.\tag{10}\]

3 The trace-log distance on \(\mathcal{M}_{++}\)↩︎

3.1 Arithmetic and geometric means↩︎

For \(A,B\in\mathcal{M}_{++}\) define the arithmetic mean and geometric mean by \[A\mathbin{\nabla}B:=\frac{A+B}{2}, \qquad A\mathbin{\#}B:=A^{1/2}\big(A^{-1/2}BA^{-1/2}\big)^{1/2}A^{1/2}.\] The geometric mean is symmetric and satisfies the arithmetic–geometric mean inequality \[\label{eq:ag} A\mathbin{\#}B \le A\mathbin{\nabla}B.\tag{11}\]

3.2 Equivalent forms of the distance↩︎

Lemma 14 (Geometric term as an averaged trace-log). For \(A,B\in\mathcal{M}_{++}\), \[\tau(\log(A\mathbin{\#}B))=\frac{1}{2}\tau(\log A)+\frac{1}{2}\tau(\log B).\] Consequently, \[\label{eq:alt-form} d_\tau(A,B)^2 =\tau\!\left(\log(A\mathbin{\nabla}B)\right)-\frac{1}{2}\tau(\log A)-\frac{1}{2}\tau(\log B) =\tau\!\big(\log(A\mathbin{\nabla}B)-\log(A\mathbin{\#}B)\big).\qquad{(3)}\]

Proof. Using \(\tau(\log X)=\log\Delta(X)\) and multiplicativity 3 , \[\Delta(A\mathbin{\#}B)=\Delta(A^{1/2})\,\Delta\!\big((A^{-1/2}BA^{-1/2})^{1/2}\big)\,\Delta(A^{1/2}).\] Since \(\Delta(A^{1/2})^2=\Delta(A)\) and \(\Delta(X^{1/2})=\Delta(X)^{1/2}\) on \(\mathcal{M}_{++}\) (by 4 ), we get \[\Delta(A\mathbin{\#}B)=\Delta(A)\cdot \Delta(A^{-1/2}BA^{-1/2})^{1/2}.\] Again by multiplicativity, \[\Delta(A^{-1/2}BA^{-1/2}) =\Delta(A^{-1/2})\,\Delta(B)\,\Delta(A^{-1/2}) =\Delta(A)^{-1}\Delta(B).\] Hence \(\Delta(A\mathbin{\#}B)=\Delta(A)^{1/2}\Delta(B)^{1/2}\), and taking \(\log\) yields the claimed identity. Substituting into ?? gives ?? . ◻

3.3 Basic metric axioms (except triangle inequality)↩︎

Proposition 15 (Nonnegativity, symmetry, definiteness). For all \(A,B\in\mathcal{M}_{++}\):

  1. \(d_\tau(A,B)\ge 0\).

  2. \(d_\tau(A,B)=d_\tau(B,A)\).

  3. \(d_\tau(A,B)=0\) if and only if \(A=B\).

Proof. Symmetry is immediate from ?? . Nonnegativity follows from 11 and operator monotonicity of \(\log\): \(\log(A\mathbin{\#}B)\le \log(A\mathbin{\nabla}B)\), hence \(d_\tau(A,B)^2\ge 0\) by ?? .

For definiteness, assume \(d_\tau(A,B)=0\). Then \(\log(A\mathbin{\nabla}B)-\log(A\mathbin{\#}B)\ge 0\) and its trace is \(0\), so by faithfulness it must be \(0\). Thus \(A\mathbin{\nabla}B=A\mathbin{\#}B\). Let \(X:=A^{-1/2}BA^{-1/2}\in\mathcal{M}_{++}\). Conjugating gives \((\mathbf{1}+X)/2=X^{1/2}\), so \((X^{1/2}-\mathbf{1})^2=0\). Since \(X^{1/2}-\mathbf{1}\) is self-adjoint, \(X^{1/2}=\mathbf{1}\), hence \(X=\mathbf{1}\) and \(A=B\). The converse is immediate. ◻

3.4 Congruence invariance and scaling↩︎

Lemma 16 (Congruence invariance). For every invertible \(S\in \mathcal{M}^\times\) and all \(A,B\in\mathcal{M}_{++}\), \[d_\tau(S^*AS,S^*BS)=d_\tau(A,B).\]

Proof. By ?? , it suffices to show that for \(X\in\mathcal{M}_{++}\), \[\label{eq:cong-trace-log} \tau(\log(S^*XS))=\tau(\log X)+2\log\Delta(S).\tag{12}\] Indeed, applying 12 to \(X=A\mathbin{\nabla}B\), \(X=A\), and \(X=B\) makes the additive constants cancel.

To prove 12 , use \(\tau(\log Y)=\log\Delta(Y)\) and multiplicativity: \[\tau(\log(S^*XS))=\log\Delta(S^*XS)=\log\Delta(S^*)+\log\Delta(X)+\log\Delta(S).\] Since \(\Delta(S^*)=\Delta(S)\) (via polar decomposition and traciality), 12 follows. ◻

Lemma 17 (Positive scalar homogeneity). For every \(c>0\) and all \(A,B\in\mathcal{M}_{++}\), \[d_\tau(cA,cB)=d_\tau(A,B).\]

Proof. Use \(\log(cX)=(\log c)\mathbf{1}+\log X\) in ?? and cancel the constant terms. ◻

Remark 18 (Reduction to a base point). By Lemma 16, for \(A,B\in\mathcal{M}_{++}\), \[d_\tau(A,B)=d_\tau\big(\mathbf{1},\,A^{-1/2}BA^{-1/2}\big).\] Thus the triangle inequality reduces to bounds for \(d_\tau(\mathbf{1},\cdot)\) and comparisons along rearrangements.

4 Scalar divergence as a Hilbert space distance↩︎

Define for \(x,y>0\) the scalar divergence \[\label{eq:scalar-delta} \delta_s(x,y)^2 := \log\!\Big(\frac{x+y}{2}\Big)-\frac{1}{2}\log x-\frac{1}{2}\log y, \qquad \delta_s(x,y):=\sqrt{\delta_s(x,y)^2}.\tag{13}\]

Lemma 19 (Integral representation). For \(x,y>0\), \[\label{eq:integral-repr} \delta_s(x,y)^2 =\frac{1}{2}\int_0^\infty \big(e^{-rx/2}-e^{-ry/2}\big)^2\,\frac{\,\mathrm{d}r}{r}.\qquad{(4)}\]

Proof. Use the classical Laplace representation for \(\log u\) (valid for all \(u>0\)), \[\log u=\int_0^\infty \frac{e^{-s}-e^{-us}}{s}\,\,\mathrm{d}s,\] as an improper integral. Apply it to \(u=(x+y)/2\), \(u=x\), and \(u=y\), combine the terms, and factor the numerator.

Convergence check. As \(r\downarrow 0\) we have \(e^{-rx/2}-e^{-ry/2}=\frac{r}{2}(y-x)+O(r^2)\), hence \(\big(e^{-rx/2}-e^{-ry/2}\big)^2/r = O(r)\), which is integrable near \(0\). As \(r\to\infty\) the integrand decays exponentially. Therefore the improper integral in ?? is finite. ◻

Corollary 20 (Scalar triangle inequality). The function \(\delta_s\) is a metric on \((0,\infty)\).

Proof. Fix \(x_0>0\) and define \(\Phi_{x_0}:(0,\infty)\to L^2((0,\infty),\,\mathrm{d}r/r)\) by \[\Phi_{x_0}(x)(r):=\frac{1}{\sqrt2}\,\big(e^{-rx/2}-e^{-r x_0/2}\big).\] Lemma 19 gives \[\delta_s(x,y)^2=\|\Phi_{x_0}(x)-\Phi_{x_0}(y)\|_{L^2((0,\infty),\,\mathrm{d}r/r)}^2,\] so \(\delta_s\) is the pullback of a Hilbert space metric and satisfies the triangle inequality. For definiteness, assume \(\delta_s(x,y)=0\). Then by ?? , \(e^{-rx/2}=e^{-ry/2}\) for a.e.\(r>0\). If \(x\neq y\), the function \(r\mapsto e^{-rx/2}-e^{-ry/2}\) is real-analytic on \((0,\infty)\) and not identically zero, hence its zero set is discrete and cannot have positive measure. Therefore \(x=y\). ◻

5 A rearrangement lower bound for \(\tau(\log(X+Y))\)↩︎

Proposition 21 (Rearrangement lower bound). Let \(X,Y\in\mathcal{M}_{++}\). Then \(X+Y\in\mathcal{M}_{++}\) and \[\label{eq:log-sum-lower} \tau(\log(X+Y)) \ge \int_{(0,1)} \log\big(\mu_X(t)+\mu_Y(t)\big)\,\,\mathrm{d}t.\qquad{(5)}\]

Proof. Since \(X,Y\) are invertible, there exist \(m_X,m_Y>0\) such that \(X\ge m_X\mathbf{1}\) and \(Y\ge m_Y\mathbf{1}\). Then \(X+Y\ge (m_X+m_Y)\mathbf{1}\), hence \(X+Y\) is invertible.

Set \(m:=m_X+m_Y\). Define \[f_m(s):= \begin{cases} -\log m+1, & 0\le s\le m,\\[0.3em] -\log s + \dfrac{s}{m}, & s\ge m. \end{cases}\] Remark (constant shift). If one prefers a formulation of 10 assuming \(f(0)=0\), apply it to \(\widetilde{f}_m:=f_m-f_m(0)\). Since we take \(u=1\), the constant term cancels on both sides. Then \(f_m\) is continuous, convex, and increasing on \([0,\infty)\). Apply 10 with \(u=1\) and \(f=f_m\): \[\int_0^1 f_m(\mu_{X+Y}(t))\,\,\mathrm{d}t \le \int_0^1 f_m(\mu_X(t)+\mu_Y(t))\,\,\mathrm{d}t.\] Since \(X\ge m_X\mathbf{1}\), we have \(E^X((\lambda,\infty))=\mathbf{1}\) for every \(0\le \lambda<m_X\), hence \(d_X(\lambda)=\tau(\mathbf{1})=1\) for \(\lambda<m_X\). By 5 this implies \(\mu_X(t)\ge m_X\) for all \(t\in(0,1)\). Similarly \(\mu_Y(t)\ge m_Y\) and \(\mu_{X+Y}(t)\ge m_X+m_Y=m\) for all \(t\in(0,1)\). For a.e.\(t\in(0,1)\) we have \(\mu_{X+Y}(t)\ge m\) and \(\mu_X(t)+\mu_Y(t)\ge m\), hence \(f_m(z)=-\log z+\frac{z}{m}\) on the relevant range. Moreover, by 7 applied to \(f(s)=s\), \[\int_0^1 \mu_{X+Y}(t)\,\,\mathrm{d}t=\tau(X+Y)=\tau(X)+\tau(Y) =\int_0^1\big(\mu_X(t)+\mu_Y(t)\big)\,\,\mathrm{d}t.\] Hence the linear terms \(\frac{1}{m}(\cdot)\) cancel, and we obtain \[\int_0^1 \log\big(\mu_{X+Y}(t)\big)\,\,\mathrm{d}t \ge \int_0^1 \log\big(\mu_X(t)+\mu_Y(t)\big)\,\,\mathrm{d}t.\] Finally, apply 9 to \(X+Y\) to obtain ?? . ◻

6 Triangle inequality for the trace-log distance on \(\mathcal{M}_{++}\)↩︎

6.1 \(d_\tau(\mathbf{1},\cdot)\) as an \(L^2\)-norm along rearrangements↩︎

Lemma 22 (\(d_\tau(\mathbf{1},\cdot)\) via generalized singular values). For every \(X\in\mathcal{M}_{++}\), \[d_\tau(\mathbf{1},X)^2=\int_{(0,1)} \delta_s\big(1,\mu_X(t)\big)^2\,\,\mathrm{d}t,\] where the integrand is taken for a.e. \(t\in(0,1)\), hence \[d_\tau(\mathbf{1},X)=\big\|\delta_s(1,\mu_X)\big\|_{L^2(0,1)}.\]

Proof. By ?? with \(A=\mathbf{1}\), \[d_\tau(\mathbf{1},X)^2=\tau\!\left(\log\!\Big(\frac{\mathbf{1}+X}{2}\Big)\right)-\frac{1}{2}\,\tau(\log X).\] Since \(u\mapsto \log\!\big(\frac{1+u}{2}\big)\) is continuous and increasing on \([0,\infty)\), we may apply 8 to obtain \[\tau\!\left(\log\!\Big(\frac{\mathbf{1}+X}{2}\Big)\right) =\int_0^1 \log\!\Big(\frac{1+\mu_X(t)}{2}\Big)\,\,\mathrm{d}t.\] Moreover, since \(X\in\mathcal{M}_{++}\), 9 gives \[\tau(\log X)=\int_0^1 \log\big(\mu_X(t)\big)\,\,\mathrm{d}t.\] Substituting these into the expression for \(d_\tau(\mathbf{1},X)^2\) yields \[d_\tau(\mathbf{1},X)^2 =\int_0^1\left(\log\!\Big(\frac{1+\mu_X(t)}{2}\Big)-\frac{1}{2}\log\big(\mu_X(t)\big)\right)\,\mathrm{d}t =\int_0^1 \delta_s\bigl(1,\mu_X(t)\bigr)^2\,\,\mathrm{d}t,\] by the definition 13 of the scalar divergence. Since \(X\in\mathcal{M}_{++}\), there exists \(m>0\) such that \(X\ge m\mathbf{1}\), hence \(\mu_X(t)\ge m\) for all \(t\in(0,1)\). Therefore \(\delta_s(1,\mu_X(t))\) is finite for a.e.\(t\in(0,1)\) (cf.Remark 12). ◻

6.2 A lower bound for \(d_\tau(S,T)\)↩︎

Lemma 23 (Fack–Kosaki lower bound for \(d_\tau(S,T)\)). For all \(S,T\in\mathcal{M}_{++}\), \[\label{eq:dST-lower} d_\tau(S,T)^2 \ge \int_0^1 \delta_s\big(\mu_S(t),\mu_T(t)\big)^2\,\,\mathrm{d}t,\qquad{(6)}\] where the integrand is taken for a.e. \(t\in(0,1)\). Hence \[\big\|\delta_s(\mu_S,\mu_T)\big\|_{L^2(0,1)}\le d_\tau(S,T).\]

Proof. Using ?? and \(\log(\frac{S+T}{2})=\log(S+T)-(\log 2)\mathbf{1}\), Proposition 21 yields \[\tau\!\left(\log\!\Big(\frac{S+T}{2}\Big)\right) \ge \int_0^1 \log\!\Big(\frac{\mu_S(t)+\mu_T(t)}{2}\Big)\,\,\mathrm{d}t.\] Moreover, 9 gives \(\tau(\log S)=\int_0^1\log\mu_S\) and similarly for \(T\). Subtracting and comparing with 13 gives ?? . ◻

6.3 Triangle inequality↩︎

Proposition 24 (Triangle inequality on \(\mathcal{M}_{++}\)). For all \(A,B,C\in\mathcal{M}_{++}\), \[d_\tau(A,C)\le d_\tau(A,B)+d_\tau(B,C).\]

Proof. By congruence invariance (Lemma 16) with \(S=A^{-1/2}\), it suffices to prove \[\label{eq:triangle-reduced} d_\tau(\mathbf{1},T)\le d_\tau(\mathbf{1},S)+d_\tau(S,T)\qquad(S,T\in\mathcal{M}_{++}).\tag{14}\] Fix \(S,T\in\mathcal{M}_{++}\). By Lemma 22, \[d_\tau(\mathbf{1},T)=\|\delta_s(1,\mu_T)\|_{L^2(0,1)}, \qquad d_\tau(\mathbf{1},S)=\|\delta_s(1,\mu_S)\|_{L^2(0,1)}.\] By Lemma 23, \(\|\delta_s(\mu_S,\mu_T)\|_{L^2(0,1)}\le d_\tau(S,T)\).

Since \(t\mapsto \mu_S(t)\) and \(t\mapsto \mu_T(t)\) are decreasing and right-continuous on \((0,1)\), they are measurable. As \(\delta_s\) is continuous on \((0,\infty)^2\), the functions \(t\mapsto \delta_s(1,\mu_S(t))\) and \(t\mapsto \delta_s(\mu_S(t),\mu_T(t))\) are measurable; moreover they belong to \(L^2(0,1)\) by Lemma 22 and Lemma 23. Since \(\delta_s\) is a scalar metric (Corollary 20), for a.e.\(t\in(0,1)\), \[\delta_s(1,\mu_T(t))\le \delta_s(1,\mu_S(t))+\delta_s(\mu_S(t),\mu_T(t)).\] Taking \(L^2\)-norms and applying Minkowski’s inequality yields 14 . ◻

Theorem 25 (Trace-log metric on a finite von Neumann algebra). Let \((\mathcal{M},\tau)\) be a finite von Neumann algebra with a faithful normal tracial state. Then \(d_\tau\) is a metric on \(\mathcal{M}_{++}\).

Proof. Proposition 15 gives symmetry, nonnegativity, and definiteness. Proposition 24 gives the triangle inequality. ◻

7 Transfer to tracial \(C^*\)-algebras via the GNS von Neumann envelope↩︎

Lemma 26 (The induced trace on the von Neumann envelope). Let \((\mathcal{A},\tau)\) be a unital \(C^*\)-algebra with a faithful tracial state and let \((H_\tau,\pi_\tau,\xi_\tau)\) be the GNS triple. Set \(\mathcal{M}:=\pi_\tau(\mathcal{A})''\subset B(H_\tau)\) and define \(\tilde{\tau}(x):=\langle x\xi_\tau,\xi_\tau\rangle\) for \(x\in\mathcal{M}\). Then \(\tilde{\tau}\) is a faithful normal tracial state on \(\mathcal{M}\).

Proof. Normality is immediate since \(\tilde{\tau}\) is the restriction to \(\mathcal{M}\) of a vector state on \(B(H_\tau)\).

For traciality, for \(a,b\in\mathcal{A}\) we have \[\tilde{\tau}(\pi_\tau(a)\pi_\tau(b))=\langle \pi_\tau(ab)\xi_\tau,\xi_\tau\rangle=\tau(ab)=\tau(ba)=\tilde{\tau}(\pi_\tau(b)\pi_\tau(a)).\] Since \(\pi_\tau(\mathcal{A})\) is strongly dense in \(\mathcal{M}\) and \(\tilde{\tau}\) is normal, this extends to all of \(\mathcal{M}\).

For faithfulness, let \(x\in\mathcal{M}_{+}\) and assume \(\tilde{\tau}(x)=0\). Then \[\|x^{1/2}\xi_\tau\|^2=\langle x\xi_\tau,\xi_\tau\rangle=\tilde{\tau}(x)=0,\] hence \(x^{1/2}\xi_\tau=0\).

We claim that \(\xi_\tau\) is separating for \(\mathcal{M}\). To see this, recall that \(H_\tau\) is the completion of \(\mathcal{A}\) with inner product \(\langle a,b\rangle=\tau(b^*a)\) and \(\xi_\tau\) is the class of \(\mathbf{1}\). Define the right action on the dense subspace \(\pi_\tau(\mathcal{A})\xi_\tau\) by \[\rho_\tau(a)\,\pi_\tau(b)\xi_\tau:=\pi_\tau(ba)\xi_\tau\qquad(a,b\in\mathcal{A}).\] For \(a,b\in\mathcal{A}\) we have \[\|\rho_\tau(a)\pi_\tau(b)\xi_\tau\|^2 =\|\pi_\tau(ba)\xi_\tau\|^2 =\tau\big((ba)^*(ba)\big) =\tau\big(a^*b^*ba\big) =\tau\big(b^*b\,aa^*\big) \le \|a\|^2\,\tau(b^*b).\] Well-definedness. If \(\pi_\tau(b)\xi_\tau=0\), then \(\|\pi_\tau(b)\xi_\tau\|^2=\tau(b^*b)=0\), hence \(b=0\) by faithfulness of \(\tau\). Thus the definition of \(\rho_\tau(a)\) on \(\pi_\tau(\mathcal{A})\xi_\tau\) is unambiguous.

Thus \(\rho_\tau(a)\) extends to a bounded operator on \(H_\tau\) with \(\|\rho_\tau(a)\|\le\|a\|\). A direct computation on the dense subspace \(\pi_\tau(\mathcal{A})\xi_\tau\) shows that \(\rho_\tau(a)\) commutes with each \(\pi_\tau(c)\): for \(a,b,c\in\mathcal{A}\), \(\pi_\tau(c)\rho_\tau(a)\pi_\tau(b)\xi_\tau=\pi_\tau(cba)\xi_\tau=\rho_\tau(a)\pi_\tau(c)\pi_\tau(b)\xi_\tau\). Hence \(\rho_\tau(\mathcal{A})\subset \pi_\tau(\mathcal{A})' \subset \mathcal{M}'\). Moreover, \[\rho_\tau(\mathcal{A})\xi_\tau=\{\pi_\tau(a)\xi_\tau:\;a\in\mathcal{A}\}\] is dense in \(H_\tau\), hence \(\xi_\tau\) is cyclic for \(\mathcal{M}'\). Therefore \(\xi_\tau\) is separating for \(\mathcal{M}\): if \(y\in\mathcal{M}\) and \(y\xi_\tau=0\), then \(y z\xi_\tau = z y\xi_\tau=0\) for all \(z\in\mathcal{M}'\), so \(y\) vanishes on the dense subspace \(\mathcal{M}'\xi_\tau\) and thus \(y=0\).

Applying this to \(y=x^{1/2}\) yields \(x^{1/2}=0\), hence \(x=0\). ◻

Proof of Theorem 2. Let \((\mathcal{A},\tau)\) be a unital \(C^*\)-algebra with faithful trace and let \(\mathcal{M}=\pi_\tau(\mathcal{A})''\) be the von Neumann envelope. By Lemma 26, \(\tilde{\tau}\) is a faithful normal tracial state on \(\mathcal{M}\).

If \(A\in\mathcal{A}_{++}\) then \(\pi_\tau(A)\in\mathcal{M}_{++}\) and continuous functional calculus is respected: \(\pi_\tau(\log A)=\log(\pi_\tau(A))\), and similarly for \(\log(\frac{A+B}{2})\). Therefore for \(A,B\in\mathcal{A}_{++}\), \[d_\tau(A,B)^2 = \tilde{\tau}\!\left(\log\!\left(\frac{\pi_\tau(A)+\pi_\tau(B)}{2}\right)\right) -\frac{1}{2}\tilde{\tau}(\log\pi_\tau(A))-\frac{1}{2}\tilde{\tau}(\log\pi_\tau(B)) = d_{\tilde{\tau}}(\pi_\tau(A),\pi_\tau(B))^2.\] Thus \(d_\tau(A,B)=d_{\tilde{\tau}}(\pi_\tau(A),\pi_\tau(B))\). Since \(d_{\tilde{\tau}}\) is a metric on \(\mathcal{M}_{++}\) by Theorem 25, its restriction to \(\pi_\tau(\mathcal{A}_{++})\) is a metric; faithfulness of \(\pi_\tau\) transfers definiteness back to \(\mathcal{A}_{++}\). ◻

8 Quantum Jensen–Shannon divergence and an integral representation↩︎

Throughout this section we work first in a finite von Neumann algebra \((\mathcal{M},\tau)\) and then transfer to \(\mathcal{A}\) via Section 7. Let \(\eta(x)=x\log x\) with \(\eta(0)=0\).

8.1 A differentiation identity↩︎

Lemma 27 (Differentiation formula). Let \(A\in\mathcal{M}_{+}\). The map \(t\mapsto \tau(\eta(A+t\mathbf{1}))\) is continuously differentiable on \((0,\infty)\) and \[\frac{\,\mathrm{d}}{\,\mathrm{d}t}\tau(\eta(A+t\mathbf{1}))=\tau(\log(A+t\mathbf{1})+\mathbf{1})\qquad(t>0).\]

Proof. Let \(E^A\) be the spectral measure of \(A\) and define the finite measure \(\nu_A(\Omega)=\tau(E^A(\Omega))\) on \([0,\|A\|]\). Then \(\tau(f(A))=\int f(\lambda)\,\,\mathrm{d}\nu_A(\lambda)\) for bounded Borel \(f\). For \(t>0\), \[\tau(\eta(A+t\mathbf{1}))=\int_{[0,\|A\|]} (\lambda+t)\log(\lambda+t)\,\,\mathrm{d}\nu_A(\lambda).\] Fix \(0<t_0<t_1<\infty\). For \(t\in[t_0,t_1]\) and \(\lambda\in[0,\|A\|]\), the derivative \(\partial_t\big((\lambda+t)\log(\lambda+t)\big)=\log(\lambda+t)+1\) is bounded in absolute value by a constant depending only on \(t_0,t_1,\|A\|\). Therefore differentiation under the integral sign is justified by dominated convergence on \([t_0,t_1]\). The same domination argument applied to \(\partial_t(\log(\lambda+t)+1)=1/(\lambda+t)\) shows that the derivative depends continuously on \(t\) on any compact subinterval of \((0,\infty)\). Hence \(t\mapsto \tau(\eta(A+t\mathbf{1}))\) is \(C^1\) on \((0,\infty)\). ◻

8.2 Vanishing at infinity↩︎

Lemma 28 (Vanishing at infinity). For all \(A,B\in\mathcal{M}_{+}\), \[\lim_{t\to\infty} J_{\tau,\eta}(A+t\mathbf{1},B+t\mathbf{1})=0.\]

Proof. Using \(\eta(cX)=c\,\eta(X)+c(\log c)\,X\) and cancellation of affine terms in \(J_{\tau,\eta}\), one checks \(J_{\tau,\eta}(cA,cB)=c\,J_{\tau,\eta}(A,B)\) for \(c>0\). Thus \[J_{\tau,\eta}(A+t\mathbf{1},B+t\mathbf{1})=t\,J_{\tau,\eta}\!\left(\mathbf{1}+\frac{A}{t},\,\mathbf{1}+\frac{B}{t}\right).\] Let \(\phi(r):=\eta(1+r)-r=(1+r)\log(1+r)-r\) for \(r\ge 0\). Fix \(r_0\in(0,1)\). Since \(\phi\) is \(C^2\) on \([0,r_0]\) and \(\phi(0)=\phi'(0)=0\), there exists \(C>0\) such that \[0\le \phi(r)\le C r^2\qquad (0\le r\le r_0).\] Choose \(t_0>0\) so that \(\|A\|/t_0\le r_0\) and \(\|B\|/t_0\le r_0\). Then for \(t\ge t_0\), functional calculus gives \[\|\phi(A/t)\|\le C\|A\|^2/t^2,\qquad \|\phi(B/t)\|\le C\|B\|^2/t^2,\qquad \Big\|\phi\Big(\frac{A+B}{2t}\Big)\Big\|\le C\|A+B\|^2/(4t^2).\] Using \(\eta(\mathbf{1}+X)=X+\phi(X)\) and the fact that affine terms cancel in \(J_{\tau,\eta}\), we get \[J_{\tau,\eta}\Big(\mathbf{1}+\frac{A}{t},\,\mathbf{1}+\frac{B}{t}\Big) =\frac{1}{2}\tau\!\Big(\phi\Big(\frac{A}{t}\Big)\Big) +\frac{1}{2}\tau\!\Big(\phi\Big(\frac{B}{t}\Big)\Big) -\tau\!\Big(\phi\Big(\frac{A+B}{2t}\Big)\Big).\] Since \(\tau\) is a state, \(|\tau(X)|\le \|X\|\), hence \[\Big|J_{\tau,\eta}\Big(\mathbf{1}+\frac{A}{t},\,\mathbf{1}+\frac{B}{t}\Big)\Big| \le \frac{1}{2}\Big\|\phi\Big(\frac{A}{t}\Big)\Big\| +\frac{1}{2}\Big\|\phi\Big(\frac{B}{t}\Big)\Big\| +\Big\|\phi\Big(\frac{A+B}{2t}\Big)\Big\| \le \frac{C'}{t^2}.\] Therefore \(J_{\tau,\eta}(A+t\mathbf{1},B+t\mathbf{1})=t\,J_{\tau,\eta}(\mathbf{1}+A/t,\mathbf{1}+B/t)=O(t^{-1})\to 0\). ◻

8.3 Integral representation and metricity↩︎

For \(t>0\) and \(A,B\in\mathcal{M}_{+}\) set \[d_{\tau,t}(A,B):=d_\tau(A+t\mathbf{1},B+t\mathbf{1}).\]

Proposition 29 (Integral representation). For all \(A,B\in\mathcal{M}_{+}\), \[\label{eq:qjsd-integral} J_{\tau,\eta}(A,B)=\int_0^\infty d_{\tau,t}(A,B)^2\,\,\mathrm{d}t.\qquad{(7)}\]

Proof. Define \(F(t):=J_{\tau,\eta}(A+t\mathbf{1},B+t\mathbf{1})\) for \(t\ge 0\). For \(t>0\), Lemma 27 gives \[F'(t) =\frac{1}{2}\tau(\log(A+t\mathbf{1})+\mathbf{1})+\frac{1}{2}\tau(\log(B+t\mathbf{1})+\mathbf{1}) -\tau\!\left(\log\!\left(\frac{A+B}{2}+t\mathbf{1}\right)+\mathbf{1}\right).\] The \(\mathbf{1}\)-terms cancel, and the remaining expression equals \(-d_{\tau,t}(A,B)^2\) by ?? . Hence \(F'(t)=-d_{\tau,t}(A,B)^2\) for \(t>0\).

Fix \(\varepsilon>0\). For \(R>\varepsilon\), \[F(R)-F(\varepsilon)=\int_\varepsilon^R F'(t)\,\,\mathrm{d}t=-\int_\varepsilon^R d_{\tau,t}(A,B)^2\,\,\mathrm{d}t.\] Letting \(R\to\infty\) and using Lemma 28 gives \[F(\varepsilon)=\int_\varepsilon^\infty d_{\tau,t}(A,B)^2\,\,\mathrm{d}t.\] The map \(t\mapsto d_{\tau,t}(A,B)^2\) is Borel measurable on \((0,\infty)\) (indeed, it is norm-continuous on \([\varepsilon,\infty)\) for every \(\varepsilon>0\) by continuous functional calculus for \(\log\)).

Moreover, \(F(t)=J_{\tau,\eta}(A+t\mathbf{1},B+t\mathbf{1})\) is continuous at \(t=0\) since \(\eta\) extends continuously to \([0,\infty)\) and functional calculus is norm-continuous for continuous functions.

Since \(d_{\tau,t}(A,B)^2\ge 0\), the functions \(t\mapsto \mathbb{1}_{(\varepsilon,\infty)}(t)\,d_{\tau,t}(A,B)^2\) increase pointwise to \(d_{\tau,t}(A,B)^2\) as \(\varepsilon\downarrow 0\). Hence, by the monotone convergence theorem, \[\lim_{\varepsilon\downarrow 0}\int_\varepsilon^\infty d_{\tau,t}(A,B)^2\,\,\mathrm{d}t =\int_0^\infty d_{\tau,t}(A,B)^2\,\,\mathrm{d}t.\] Finally, since \(\eta\) extends continuously to \([0,\infty)\) with \(\eta(0)=0\), functional calculus is norm-continuous for the map \(X\mapsto \eta(X)\) on bounded sets. Hence \(t\mapsto \tau(\eta(A+t\mathbf{1}))\), \(t\mapsto \tau(\eta(B+t\mathbf{1}))\), and \(t\mapsto \tau\!\big(\eta(\tfrac{A+B}{2}+t\mathbf{1})\big)\) are continuous at \(t=0\). Therefore \(F(\varepsilon)\to F(0)=J_{\tau,\eta}(A,B)\) as \(\varepsilon\downarrow 0\), and combining this with the monotone convergence step above yields ?? . ◻

Proposition 30 (QJSD is a metric in the von Neumann setting). Let \(A,B\in\mathcal{M}_{+}\) and set \(D_{\tau,\eta}:=\sqrt{J_{\tau,\eta}}\). Then \(D_{\tau,\eta}\) is a metric on \(\mathcal{M}_{+}\).

Proof. Symmetry is clear. Nonnegativity follows from ?? . If \(D_{\tau,\eta}(A,B)=0\), then \(d_{\tau,t}(A,B)=0\) for a.e.\(t>0\), hence \(A+t\mathbf{1}=B+t\mathbf{1}\) for such \(t\) and therefore \(A=B\).

For the triangle inequality, let \(A,B,C\in\mathcal{M}_{+}\). For each \(t>0\), \(d_{\tau,t}\) is the trace-log metric on \(\mathcal{M}_{++}\) applied to \((A+t\mathbf{1},B+t\mathbf{1},C+t\mathbf{1})\), hence by Theorem 25, \[d_{\tau,t}(A,C)\le d_{\tau,t}(A,B)+d_{\tau,t}(B,C).\] Taking \(L^2(0,\infty)\) norms and applying Minkowski yields \[D_{\tau,\eta}(A,C)=\|d_{\tau,\bullet}(A,C)\|_{L^2(0,\infty)} \le \|d_{\tau,\bullet}(A,B)\|_{L^2}+\|d_{\tau,\bullet}(B,C)\|_{L^2} = D_{\tau,\eta}(A,B)+D_{\tau,\eta}(B,C).\] ◻

Proof of Theorem 7. As in Section 7, pass to the von Neumann envelope \(\mathcal{M}=\pi_\tau(\mathcal{A})''\) with induced faithful normal trace \(\tilde{\tau}\). Functional calculus gives \(\pi_\tau(\eta(A))=\eta(\pi_\tau(A))\) and \(\tilde{\tau}(\pi_\tau(\eta(A)))=\tau(\eta(A))\). Hence \(J_{\tau,\eta}(A,B)=J_{\tilde{\tau},\eta}(\pi_\tau(A),\pi_\tau(B))\). By Proposition 30, \(D_{\tilde{\tau},\eta}\) is a metric; restricting to \(\pi_\tau(\mathcal{A}_{+})\) and using faithfulness of \(\pi_\tau\) yields that \(D_{\tau,\eta}\) is a metric on \(\mathcal{A}_{+}\). ◻

9 Is the QJSD metric Hilbertian? Negative results in matrix dimensions↩︎

9.1 Hilbertian metrics and conditional negative definiteness↩︎

Definition 31 (Conditionally negative definite kernel / negative type). Let \(X\) be a set and \(K:X\times X\to\mathbb{R}\) be symmetric with \(K(x,x)=0\). We say \(K\) is conditionally negative definite (CND) if for every \(m\in\mathbb{N}\), every \(x_1,\dots,x_m\in X\), and every \(c_1,\dots,c_m\in\mathbb{R}\) with \(\sum_{i=1}^m c_i=0\), one has \[\sum_{i,j=1}^m c_i c_j\,K(x_i,x_j)\le 0.\] A metric \(d\) on \(X\) is Hilbertian if there exist a Hilbert space \(\mathcal{H}\) and an embedding \(\Phi:X\to\mathcal{H}\) such that \(d(x,y)=\|\Phi(x)-\Phi(y)\|_{\mathcal{H}}\) for all \(x,y\in X\).

Lemma 32 (CND kernels produce Hilbert embeddings). Let \(K\) be a CND kernel on \(X\). Then there exist a Hilbert space \(\mathcal{H}\) and a map \(\Phi:X\to\mathcal{H}\) such that \[K(x,y)=\|\Phi(x)-\Phi(y)\|_{\mathcal{H}}^{2}\qquad(x,y\in X).\] Consequently, a metric space \((X,d)\) is Hilbertian if and only if \(d^2\) is CND.

Proof. Fix \(x_0\in X\). Let \(\mathcal{V}\) be the real vector space of finitely supported functions \(u:X\to\mathbb{R}\) with \(\sum_{x\in X}u(x)=0\). Define a symmetric bilinear form on \(\mathcal{V}\) by \[\langle u,v\rangle_0:=-\frac{1}{2}\sum_{x,y\in X}u(x)v(y)\,K(x,y).\] CND of \(K\) implies \(\langle u,u\rangle_0\ge 0\) for all \(u\in\mathcal{V}\). Quotient by the null space and complete to obtain a Hilbert space \(\mathcal{H}\). For \(x\in X\) let \(u_x:=\delta_x-\delta_{x_0}\in\mathcal{V}\) and set \(\Phi(x):=[u_x]\in\mathcal{H}\). A direct expansion shows \(\|\Phi(x)-\Phi(y)\|^2=K(x,y)\). The final assertion follows by applying this to \(K=d^2\) and observing that squared Hilbert distances are CND. ◻

Lemma 33 (Exponentials and CND). Let \(K:X\times X\to\mathbb{R}\) be symmetric with \(K(x,x)=0\). If \(K\) is CND, then for every \(\beta>0\) the kernel \(\exp(-\beta K)\) is positive definite. Conversely, if \(\exp(-\beta K)\) is positive definite for all \(\beta>0\), then \(K\) is CND.

Proof. This is a standard equivalence due to Schoenberg; see [25] (or [26]). Assume first that \(K\) is CND. By Lemma 32, \(K(x,y)=\|\Phi(x)-\Phi(y)\|^2\) for some Hilbert embedding. For \(u,v\in\mathcal{H}\), \[e^{-\beta\|u-v\|^2}=e^{-\beta\|u\|^2}e^{-\beta\|v\|^2}e^{2\beta\langle u,v\rangle}\] and \(e^{2\beta\langle u,v\rangle}\) is positive definite as a power series with positive coefficients (tensor powers). Multiplying by the positive diagonal factors preserves positive definiteness. Hence \(\exp(-\beta K)\) is positive definite.

Conversely, assume \(\exp(-\beta K)\) is positive definite for all \(\beta>0\). Fix \(x_1,\dots,x_m\in X\) and \(c_1,\dots,c_m\in\mathbb{R}\) with \(\sum_i c_i=0\). Define \[F(\beta):=\sum_{i,j=1}^m c_i c_j\,e^{-\beta K(x_i,x_j)}\ge 0\qquad(\beta>0),\] and note \(F(0)=\sum_{i,j}c_ic_j=(\sum_i c_i)^2=0\). Thus \(F'(0^+)\ge 0\). But \[F'(\beta)=-\sum_{i,j} c_i c_j\,K(x_i,x_j)\,e^{-\beta K(x_i,x_j)},\] so \(F'(0^+)=-\sum_{i,j}c_i c_j K(x_i,x_j)\ge 0\), i.e.\(\sum_{i,j}c_ic_jK(x_i,x_j)\le 0\). Hence \(K\) is CND. ◻

9.2 Non-Hilbertianity on \(M_n^{++}(\mathbb{C})\) for \(n\ge 2\)↩︎

In this subsection let \(\mathcal{A}=M_n(\mathbb{C})\) and let \(\tau=\frac{1}{n}\mathop{\mathrm{Tr}}\).

Proposition 34 (QJSD is not Hilbertian on \(M_n^{++}(\mathbb{C})\) for \(n\ge 2\)). Let \(n\ge 2\). The QJSD metric \(D_{\tau,\eta}\) on \(M_n^{++}(\mathbb{C})\) is not* Hilbertian. Consequently, \(D_{\tau,\eta}\) is not Hilbertian on the full cone \(M_n^{+}(\mathbb{C})\).*

Proof. Recall that a metric space \((X,d)\) is Hilbertian if and only if \(d^2\) is CND (Lemma 32). Thus it suffices to show that the kernel \(J_{\tau,\eta}\) is not CND.

Step 1: the case \(n=2\). By Example 1, there exist \(X_1,\dots,X_m\in M_2^{++}(\mathbb{C})\) and \(c_1,\dots,c_m\in\mathbb{R}\) with \(\sum_{i=1}^m c_i=0\) such that \[\sum_{i,j=1}^m c_ic_j\,J_{\tau_2,\eta}(X_i,X_j)>0, \qquad \text{where }\;\tau_2=\tfrac12\mathop{\mathrm{Tr}}\text{ on }M_2(\mathbb{C}).\] Hence \(J_{\tau_2,\eta}\) is not CND on \(M_2^{++}(\mathbb{C})\), and therefore the metric \(D_{\tau_2,\eta}=\sqrt{J_{\tau_2,\eta}}\) is not Hilbertian on \(M_2^{++}(\mathbb{C})\).

Step 2: embedding into \(M_n^{++}(\mathbb{C})\) for \(n\ge 2\). Let \(n\ge 2\) and write \(\tau_n=\tfrac1n\mathop{\mathrm{Tr}}\) on \(M_n(\mathbb{C})\). Define the block embedding \(\iota:M_2^{++}(\mathbb{C})\to M_n^{++}(\mathbb{C})\) by \[\iota(X):=X\oplus \mathbf{1}_{n-2}.\] Since \(\eta(1)=1\cdot\log 1=0\) and \(\eta\) acts blockwise on block-diagonal matrices, we have \[\eta(\iota(X))=\eta(X)\oplus 0_{n-2}, \qquad \eta\!\left(\frac{\iota(X)+\iota(Y)}{2}\right) = \eta\!\left(\frac{X+Y}{2}\right)\oplus 0_{n-2}.\] Therefore, for all \(X,Y\in M_2^{++}(\mathbb{C})\), \[J_{\tau_n,\eta}(\iota(X),\iota(Y)) = \frac{1}{n}\mathop{\mathrm{Tr}}\!\left( \frac{1}{2}\eta(X)+\frac{1}{2}\eta(Y)-\eta\!\left(\frac{X+Y}{2}\right) \right) = \frac{2}{n}\,J_{\tau_2,\eta}(X,Y).\] Now set \(Y_i:=\iota(X_i)\in M_n^{++}(\mathbb{C})\) and use the same coefficients \(c_i\) as in Step 1: \[\sum_{i,j=1}^m c_ic_j\,J_{\tau_n,\eta}(Y_i,Y_j) = \frac{2}{n}\sum_{i,j=1}^m c_ic_j\,J_{\tau_2,\eta}(X_i,X_j) >0.\] Hence \(J_{\tau_n,\eta}\) is not CND on \(M_n^{++}(\mathbb{C})\), so \(D_{\tau_n,\eta}\) is not Hilbertian on \(M_n^{++}(\mathbb{C})\).

Finally, if \(D_{\tau_n,\eta}\) were Hilbertian on \(M_n(\mathbb{C})_+\), there would exist an embedding \(\Phi:M_n(\mathbb{C})_+\to\mathcal{H}\) into a Hilbert space with \(D_{\tau_n,\eta}(X,Y)=\|\Phi(X)-\Phi(Y)\|\). Restricting \(\Phi\) to the subset \(M_n^{++}(\mathbb{C})\) would make \(D_{\tau_n,\eta}\) Hilbertian on \(M_n^{++}(\mathbb{C})\), contradicting the conclusion above. ◻

9.3 Non-Hilbertianity on density matrices for \(n\ge 3\)↩︎

Let \[\mathcal{D}_n:=\{\rho\in M_n^{+}(\mathbb{C}):\mathop{\mathrm{Tr}}(\rho)=1\}\] be the density matrices.

Proposition 35 (QJSD is not Hilbertian on \(\mathcal{D}_n\) for \(n\ge 3\)). For every \(n\ge 3\), the QJSD metric \(D_{\tau,\eta}=\sqrt{J_{\tau,\eta}}\) is not Hilbertian on \(\mathcal{D}_n\).

Proof. Let \(\tau_n=\frac{1}{n}\mathop{\mathrm{Tr}}\) on \(M_n(\mathbb{C})\). By Lemma 32, a metric is Hilbertian if and only if its square is conditionally negative definite (CND). Thus it suffices to show that the kernel \(J_{\tau_n,\eta}\) is not CND on \(\mathcal{D}_n\).

Step 1: a CND violation already on \(\mathcal{D}_3\). In Example 1 we exhibit density matrices \(\rho_1,\dots,\rho_5\in\mathcal{D}_3\) and integers \(c_1,\dots,c_5\) with \(\sum_i c_i=0\) such that \[\sum_{i,j=1}^5 c_ic_j\,J_{\tau_3,\eta}(\rho_i,\rho_j) >0.\] Hence \(J_{\tau_3,\eta}\) is not CND on \(\mathcal{D}_3\), and therefore \(D_{\tau_3,\eta}\) is not Hilbertian on \(\mathcal{D}_3\).

Step 2: embedding \(\mathcal{D}_3\) into \(\mathcal{D}_n\) for \(n>3\). For \(n>3\), embed \(\mathcal{D}_3\) into \(\mathcal{D}_n\) via \[\iota(\rho):=\rho\oplus 0_{n-3}\qquad(\rho\in\mathcal{D}_3).\] Since \(\eta(0)=0\) and \(\eta\) acts blockwise on block-diagonal matrices, we have for all \(\rho,\sigma\in\mathcal{D}_3\), \[J_{\tau_n,\eta}(\iota(\rho),\iota(\sigma)) =\frac{1}{n}\mathop{\mathrm{Tr}}\!\left(\frac{1}{2}\eta(\rho)+\frac{1}{2}\eta(\sigma)-\eta\!\left(\frac{\rho+\sigma}{2}\right)\right) =\frac{3}{n}\,J_{\tau_3,\eta}(\rho,\sigma).\] Therefore the same CND violation on \(\mathcal{D}_3\) persists on \(\mathcal{D}_n\). This proves that \(D_{\tau_n,\eta}\) is not Hilbertian on \(\mathcal{D}_n\) for every \(n\ge 3\). ◻

Example 1 (Explicit CND violation with integer matrices (certified)). Let \(\tau_2=\frac{1}{2}\mathop{\mathrm{Tr}}\) on \(M_2(\mathbb{C})\) and \(\eta(x)=x\log x\) with \(\eta(0)=0\). Consider the following \(2\times 2\) real symmetric integer matrices: \[X_1=\begin{pmatrix}2&1\\[2pt]1&1\end{pmatrix},\quad X_2=\begin{pmatrix}9&2\\[2pt]2&1\end{pmatrix},\quad X_3=\begin{pmatrix}2&1\\[2pt]1&7\end{pmatrix},\quad X_4=\begin{pmatrix}8&5\\[2pt]5&8\end{pmatrix},\quad X_5=\begin{pmatrix}8&8\\[2pt]8&9\end{pmatrix}.\] They all lie in \(M_2^{++}(\mathbb{C})\) since \[\det(X_1)=1,\;\det(X_2)=5,\;\det(X_3)=13,\;\det(X_4)=39,\;\det(X_5)=8.\] Let \[c=(c_1,\dots,c_5)=(-10,\;10,\;10,\;-20,\;10),\qquad \sum_{i=1}^5 c_i=0.\]

(i) Violation on \(M_2^{++}(\mathbb{C})\). Define \[S_2:=\sum_{i,j=1}^5 c_ic_j\,J_{\tau_2,\eta}(X_i,X_j).\] A directed-rounding interval-arithmetic computation yields the rigorous enclosure \[S_2\in\bigl[\,9.8113517061626681,\;9.8113517062313917\,\bigr],\] hence \(S_2>0\) and \(J_{\tau_2,\eta}\) is not CND on \(M_2^{++}(\mathbb{C})\).

(ii) Violation on \(\mathcal{D}_3\). Fix \(T=40\) (note that \(T>\max_i\mathop{\mathrm{Tr}}(X_i)=17\)) and define density matrices in \(M_3(\mathbb{C})\) by the block embedding \[\rho_i:=\frac{1}{T}X_i\;\oplus\;\left(1-\frac{\mathop{\mathrm{Tr}}(X_i)}{T}\right)\in\mathcal{D}_3, \qquad i=1,\dots,5.\] (For instance, \(\rho_1=\frac{1}{40}\!\begin{pmatrix}2&1\\1&1\end{pmatrix}\oplus\frac{37}{40}\), \(\rho_4=\frac{1}{40}\!\begin{pmatrix}8&5\\5&8\end{pmatrix}\oplus\frac{3}{5}\), etc.) Let \(\tau_3=\frac{1}{3}\mathop{\mathrm{Tr}}\) and set \[S_3:=\sum_{i,j=1}^5 c_ic_j\,J_{\tau_3,\eta}(\rho_i,\rho_j).\] The same certified interval method gives the enclosure \[S_3\in\bigl[\,0.1601205745051536,\;0.16012057450782458\,\bigr],\] so \(S_3>0\). Therefore \(J_{\tau_3,\eta}\) is not CND on \(\mathcal{D}_3\).

Reproducibility. All inputs in this example are given by integers (and, after the density-matrix embedding, rationals with fixed denominator), so the verification can be carried out without any floating-point parsing ambiguity. A fully rigorous directed-rounding interval-arithmetic workflow and pseudo-code are provided in Appendix 11.

Remark 36 (Reproducibility and certification). The positivity statements in Example 1 can be verified in a fully rigorous way by directed-rounding interval arithmetic.

(1) Exact input data. All matrix entries are integers, and the coefficients \(c_i\) are integers with \(\sum_i c_i=0\). Moreover, in the density-matrix embedding we fix \(T=40\), so every entry of \(\rho_i\) is a rational number with denominator \(T\) (and the last diagonal entry is \(1-\mathop{\mathrm{Tr}}(X_i)/T\)). Thus there is no floating-point parsing ambiguity: all inputs can be treated as exact rationals before interval evaluation.

(2) Reduction to one-dimensional certified operations. For a \(2\times2\) real symmetric matrix \(X=\begin{pmatrix}a&b\\ b&d\end{pmatrix}\), the eigenvalues are given in closed form by \[\lambda_\pm(X)=\frac{(a+d)\pm\sqrt{(a-d)^2+4b^2}}{2}.\] Hence \[\mathop{\mathrm{Tr}}(\eta(X))=\eta(\lambda_+(X))+\eta(\lambda_-(X)),\] so computing \(J_{\tau_2,\eta}(X_i,X_j)\) reduces to interval evaluation of addition, multiplication, square root, and logarithm in one real variable (with directed rounding). For the block-diagonal states \(\rho_i\in\mathcal{D}_3\), one additionally evaluates the scalar term \(\eta\!\left(1-\mathop{\mathrm{Tr}}(X_i)/T\right)\) and the analogous term for \((\rho_i+\rho_j)/2\); no higher-dimensional spectral computation is needed.

(3) Certified positivity. With directed rounding, the computed intervals provably enclose the exact values of \(S_2\) and \(S_3\); a strictly positive lower endpoint therefore certifies the CND violation. This workflow can be implemented in any standard verified interval package (e.g.Arb, INTLAB, or Julia IntervalArithmetic).

10 Jensen generators yielding metrics↩︎

Definition 37 (Operator convex). A function \(f:(0,\infty)\to\mathbb{R}\) is operator convex if for every \(n\in\mathbb{N}\), all \(X,Y\in M_n^{++}(\mathbb{C})\) and \(\lambda\in[0,1]\), \[f(\lambda X+(1-\lambda)Y)\le \lambda f(X)+(1-\lambda)f(Y)\] in the Loewner order.

Lemma 38 (An integral representation for \(f'\)). Let \(f:(0,\infty)\to\mathbb{R}\) be operator convex. Then \(f'\) is operator monotone, and admits a Nevanlinna–Stieltjes type representation. More precisely, there exist constants \(a\in\mathbb{R}\), \(b\ge 0\) and a positive Borel measure \(\nu\) on \((0,\infty)\) with \(\int_{(0,\infty)}\frac{1}{1+t}\,\,\mathrm{d}\nu(t)<\infty\) such that \[\label{eq:eta-prime-rep} f'(x)=a+bx+\int_{(0,\infty)}\frac{x}{x+t}\,\,\mathrm{d}\nu(t),\qquad x>0.\qquad{(8)}\]

Proof. By [27] (see also [28]), if \(f\) is operator convex on \((0,\infty)\) then \(f\) is differentiable and, for each \(t_0>0\), the divided difference \[g_{t_0}(t)= \begin{cases} \dfrac{f(t)-f(t_0)}{t-t_0}, & t\neq t_0,\\[0.4em] f'(t_0), & t=t_0, \end{cases}\] is operator monotone on \((0,\infty)\). In particular, \(h:=f'\) is operator monotone on \((0,\infty)\).

We now invoke the Stieltjes (Löwner) representation for operator monotone functions on \((0,\infty)\); see, e.g., [27] (equivalently, the integral representation in [27] can be rewritten in this form). Thus there exist \(a\in\mathbb{R}\), \(b\ge 0\) and a positive Borel measure \(\nu\) on \((0,\infty)\) such that \[h(x)=a+bx+\int_{(0,\infty)}\frac{x}{x+t}\,\,\mathrm{d}\nu(t),\qquad x>0.\] Moreover, \(h\) is finite-valued on \((0,\infty)\) and the integral in ?? is finite for every \(x>0\). In particular, evaluating at \(x=1\) gives \[\int_{(0,\infty)}\frac{1}{1+t}\,\,\mathrm{d}\nu(t) = h(1)-a-b <\infty,\] which is exactly the stated integrability condition. ◻

Proof of Theorem 9. Work first in a finite von Neumann algebra \((\mathcal{M},\tau)\) and fix \(A,B\in\mathcal{M}_{+}\). By Lemma 38, \(f'\) has the form ?? . Integrating from \(1\) to \(x\) yields \[f(x)=\alpha+\beta x+\frac{b}{2}x^2+\int_{(0,\infty)}\Bigl((x-1)-t\log\!\Bigl(\frac{x+t}{1+t}\Bigr)\Bigr)\,\,\mathrm{d}\nu(t), \qquad x>0,\] for suitable constants \(\alpha,\beta\in\mathbb{R}\).

Uniform convergence on bounded \(x\)-intervals. Fix \(R>0\) and consider \[k_t(x):=(x-1)-t\log\!\Big(\frac{x+t}{1+t}\Big),\qquad x\ge 0,\;t>0.\] First, note that for every \(L>0\) one has \(\nu((0,L])<\infty\). Indeed, since \(\frac{1}{1+t}\ge \frac{1}{1+L}\) on \((0,L]\), we get \[\nu((0,L])\le (1+L)\int_{(0,\infty)}\frac{1}{1+t}\,\,\mathrm{d}\nu(t)<\infty.\]

Next, for \(x\in[0,R]\) and \(t>0\) we write \[\log\!\Big(\frac{x+t}{1+t}\Big)=\log\!\Big(1+\frac{x-1}{1+t}\Big).\] By the mean value theorem for \(\log\), there exists \(\theta=\theta(x,t)\in(0,1)\) such that \[\log\!\Big(1+\frac{x-1}{1+t}\Big)=\frac{x-1}{t+1+\theta(x-1)}.\] Therefore \[k_t(x) =(x-1)\Bigl(1-\frac{t}{t+1+\theta(x-1)}\Bigr) =\frac{(x-1)\bigl(1+\theta(x-1)\bigr)}{t+1+\theta(x-1)}.\] Let \(M:=\max\{1,R-1\}\), so \(|x-1|\le M\) for \(x\in[0,R]\). For \(t\ge 2M\) we have \(t+1+\theta(x-1)\ge t+1-M\ge t/2\), hence \[|k_t(x)|\le \frac{2M(1+M)}{t}\le \frac{C_R}{1+t}\qquad (x\in[0,R],\;t\ge 2M).\] On the remaining range \(t\in(0,2M]\), we argue by a compactness extension. Define \(k_0(x):=x-1\) for \(x\ge 0\). Then for \(t>0\), \[k_t(x)-k_0(x)=-t\log\!\Big(\frac{x+t}{1+t}\Big).\] Fix \(R>0\). For \(x\in[0,R]\) and \(t\in(0,1]\) we have \[\Big|\log\!\Big(\frac{x+t}{1+t}\Big)\Big| \le |\log(x+t)|+|\log(1+t)| \le |\log t|+\log(R+1)+\log 2,\] hence \[\sup_{x\in[0,R]}|k_t(x)-k_0(x)| \le t\big(|\log t|+\log(R+1)+\log 2\big)\xrightarrow[t\downarrow 0]{}0.\] Therefore \((t,x)\mapsto k_t(x)\) extends continuously to the compact set \([0,2M]\times[0,R]\) (by setting \(k_0(x)=x-1\)), and hence is bounded there. Consequently, there exists a constant \(C_R'>0\) such that \[|k_t(x)|\le C_R' \qquad (x\in[0,R],\;0<t\le 2M).\]

Combining these bounds yields an integrable dominating function: \[\sup_{x\in[0,R]} |k_t(x)| \le C_R'\,\mathbb{1}_{(0,2M]}(t)+\frac{C_R}{1+t}\,\mathbb{1}_{(2M,\infty)}(t),\] and the right-hand side is \(\nu\)-integrable because \(\nu((0,2M])<\infty\) and \(\int_{(0,\infty)}\frac{1}{1+t}\,\,\mathrm{d}\nu(t)<\infty\). Hence \(x\mapsto \int k_t(x)\,\,\mathrm{d}\nu(t)\) converges absolutely and uniformly on \([0,R]\). Consequently, for each fixed \(R>0\) the scalar function \[g_R(x):=\int_{(0,\infty)} k_t(x)\,\,\mathrm{d}\nu(t),\qquad x\in[0,R],\] is well-defined and continuous on \([0,R]\), and the convergence is uniform on \([0,R]\).

Operator-valued integral and interchange with \(\tau\). Now fix \(A,B\in\mathcal{M}_+\) and choose \(R>\max\{\|A\|,\|B\|,\|\tfrac{A+B}{2}\|\}\). By the domination obtained above, there exists \(h_R\in L^1((0,\infty),\nu)\) such that \[\sup_{x\in[0,R]}|k_t(x)|\le h_R(t)\qquad (t>0).\] Hence for each \(X\in\{A,B,\tfrac{A+B}{2}\}\) the map \(t\mapsto k_t(X)\) is Bochner integrable in operator norm and \[g_R(X)=\int_{(0,\infty)} k_t(X)\,\,\mathrm{d}\nu(t)\in\mathcal{M}, \qquad \text{with}\quad \tau(g_R(X))=\int_{(0,\infty)} \tau(k_t(X))\,\,\mathrm{d}\nu(t),\] since \(|\tau(k_t(X))|\le \|k_t(X)\|\le h_R(t)\) and \(h_R\) is \(\nu\)-integrable.

Therefore we may compute the Jensen gap by interchanging \(\tau\) with the \(t\)-integral. Affine terms cancel in Jensen gaps, and one obtains \[J_{\tau,f}(A,B)=\frac{b}{2}\,J_{\tau,x^2}(A,B)+\int_{(0,\infty)}J_{\tau,k_t}(A,B)\,\,\mathrm{d}\nu(t),\] where \(k_t(x):=(x-1)-t\log\!\big(\frac{x+t}{1+t}\big)\). Here \(J_{\tau,x^2}(A,B)=\frac{1}{4}\left\|A-B\right\|_{2,\tau}^2\). Moreover, \(k_t(x)\) differs from \(-t\log(x+t)\) by an affine function in \(x\) (depending on \(t\)), hence \[J_{\tau,k_t}(A,B)=t\,d_\tau(A+t\mathbf{1},B+t\mathbf{1})^2.\] This gives ?? .

Define a measure space \(\Omega:=\{0\}\cup(0,\infty)\) where \(\{0\}\) carries counting measure and \((0,\infty)\) carries \(\nu\). For \(X,Y\in\mathcal{M}_+\) set \[w_{X,Y}(0):=\sqrt{\frac{b}{8}}\;\left\|X-Y\right\|_{2,\tau}, \qquad w_{X,Y}(t):=\sqrt{t}\,d_\tau(X+t\mathbf{1},Y+t\mathbf{1})\quad (t>0).\] Then ?? reads \[J_{\tau,f}(X,Y)=\|w_{X,Y}\|_{L^2(\Omega)}^2.\] For each \(t\in\Omega\), \(w_{A,C}(t)\le w_{A,B}(t)+w_{B,C}(t)\) by the triangle inequalities for \(\left\|\cdot\right\|_{2,\tau}\) and for \(d_\tau(\cdot,\cdot)\) (applied to the shifted pairs when \(t>0\)). Applying Minkowski’s inequality in \(L^2(\Omega)\) yields the triangle inequality for \(\sqrt{J_{\tau,f}}\). Since \(f\) is not affine, either \(b>0\) or \(\nu\neq 0\), so \(\sqrt{J_{\tau,f}}(A,B)=0\) forces \(A=B\).

Finally, transfer the metricity from \(\mathcal{M}\) back to \(\mathcal{A}\) via the GNS von Neumann envelope as in Section 7. ◻

Corollary 39. Let \(f:[0,\infty)\to\mathbb{R}\) be operator convex and not affine. Then \(\sqrt{J_f}\) defines a metric on the density matrices \(\mathcal{D}_n\) for every \(n\).

Proof. Apply Theorem 9 to \(\mathcal{A}=M_n(\mathbb{C})\) with \(\tau=\frac{1}{n}\mathop{\mathrm{Tr}}\) and restrict to \(\mathcal{D}_n\subset \mathcal{A}_{+}\). ◻

11 Certified verification of Example 1↩︎

This appendix records a concrete interval-arithmetic workflow which rigorously encloses the quantities \(S_2\) and \(S_3\) in Example 1. The key point is that for \(2\times 2\) real symmetric matrices, eigenvalues are available in closed form, so the whole computation reduces to one-dimensional interval operations (addition, multiplication, square root, logarithm) with directed rounding.

11.1 Closed-form eigenvalues for \(2\times 2\) blocks↩︎

For \(X=\begin{pmatrix}a&b\\ b&d\end{pmatrix}\in M_2^{++}(\mathbb{C})\) (real symmetric), set \[\mathop{\mathrm{Tr}}(X)=a+d,\qquad \det(X)=ad-b^2,\qquad \Delta=(a-d)^2+4b^2.\] Then the eigenvalues are \[\lambda_\pm(X)=\frac{\mathop{\mathrm{Tr}}(X)\pm \sqrt{\Delta}}{2}.\] Hence for \(\eta(x)=x\log x\), \[\mathop{\mathrm{Tr}}(\eta(X))=\eta(\lambda_+(X))+\eta(\lambda_-(X)).\]

11.2 Interval-arithmetic pseudo-code↩︎

Below is pseudo-code in a language-agnostic style; any standard certified interval package (e.g.INTLAB, Arb, or Julia IntervalArithmetic) can implement the same steps.

        Input:
        Matrices X_i (2x2), coefficients c_i (sum c_i = 0), entropy eta(x)=x*log(x)
        Normalized trace tau_2 = (1/2)Tr on M_2
        
        Interval primitives (directed rounding):
        +, -, *, /, sqrt(·), log(·) on intervals
        
        function eigvals_2x2_interval(a,b,d):
        tr   = a + d
        disc = (a - d)^2 + 4*b^2
        s    = sqrt(disc)
        lam_plus  = (tr + s)/2
        lam_minus = (tr - s)/2
        return (lam_plus, lam_minus)
        
        function Tr_eta_2x2_interval(X):
        (lam_plus, lam_minus) = eigvals_2x2_interval(X.a, X.b, X.d)
        return eta(lam_plus) + eta(lam_minus)
        
        function J_tau2_interval(X,Y):
        Z = (X + Y)/2
        # J_{tau_2,eta}(X,Y) = (1/2)tau_2(eta(X)) + (1/2)tau_2(eta(Y)) - tau_2(eta(Z))
        # with tau_2=(1/2)Tr this equals:
        return (1/4) * ( Tr_eta_2x2_interval(X) + Tr_eta_2x2_interval(Y)
        - 2*Tr_eta_2x2_interval(Z) )
        
        Compute S_2:
        S2 = 0
        for i=1..m:
        for j=1..m:
        S2 += c_i * c_j * J_tau2_interval(X_i, X_j)
        output rigorous enclosure interval(S2)
        
        For S_3 (block-diagonal density matrices):
        choose T > 2*max_i Tr(X_i)
        rho_i = blockdiag( X_i/T, 1 - Tr(X_i)/T ) in D_3
        use Tr(eta(rho_i)) = Tr(eta(X_i/T)) + eta(1 - Tr(X_i)/T)
        and similarly for (rho_i + rho_j)/2.
        With tau_3 = (1/3)Tr, define J_tau3 analogously and compute S_3.

In the present paper we treat all decimal inputs as exact rationals at the interval endpoints, so the certification is independent of binary floating-point parsing.

References↩︎

[1]
D. Virosztek, The metric property of the quantum Jensen–Shannon divergence, Adv. Math.380(2021), 107595. .
[2]
S. Sra, Positive definite matrices and the \(S\)-divergence, Proc. Amer. Math. Soc.144(2016), 2787–2797. .
[3]
M. R. Bridson and A. Haefliger, Metric Spaces of Non-Positive Curvature, Grundlehren der Mathematischen Wissenschaften, vol. 319, Springer-Verlag, Berlin, 1999.
[4]
R. Bhatia, Positive Definite Matrices, Princeton Series in Applied Mathematics, Princeton University Press, Princeton, NJ, 2007. .
[5]
S. Boyd and L. Vandenberghe, Convex Optimization, Cambridge University Press, Cambridge, 2004.
[6]
A. Cherian, S. Sra, A. Banerjee, and N. Papanikolopoulos, Efficient similarity search for covariance matrices via the Jensen–Bregman LogDet divergence, in Proc. IEEE Int. Conf. Computer Vision (ICCV), 2011, pp. 2399–2406.
[7]
P. Chen, Y. Chen, and M. Rao, Metrics defined by Bregman divergences, Commun. Math. Sci.6(2008), no. 4, 927–948. .
[8]
Z. Chebbi and M. Moakher, Means of Hermitian positive-definite matrices based on the log-determinant \(\alpha\)-divergence function, Linear Algebra Appl.436(2012), no. 7, 1872–1889. .
[9]
H. Q. Minh, Infinite-dimensional log-determinant divergences between positive definite trace class operators, Linear Algebra Appl.528(2017), 331–383. .
[10]
H. Q. Minh, Infinite-dimensional log-determinant divergences between positive definite Hilbert–Schmidt operators, Positivity24(2020), 631–662. .
[11]
B. Fuglede and R. V. Kadison, Determinant theory in finite factors, Ann. of Math. (2)55(1952), 520–530. .
[12]
M. Gaál, G. Nagy, and P. Szokol, Isometries on positive definite operators with unit Fuglede–Kadison determinant, Taiwanese J. Math.23(2019), no. 6, 1423–1433. .
[13]
A. P. Majtey, P. W. Lamberti, and D. P. Prato, Jensen–Shannon divergence as a measure of distinguishability between mixed quantum states, Phys. Rev. A72(2005), 052310. .
[14]
J. Dajka, J. Łuczka, and P. Hänggi, Distance between quantum states in the presence of initial qubit-environment correlations: a comparative study, Phys. Rev. A84(2011), 032120. .
[15]
C. Radhakrishnan, M. Parthasarathy, S. Jambulingam, and T. Byrnes, Distribution of quantum coherence in multipartite systems, Phys. Rev. Lett.116(2016), 150504. .
[16]
W. Roga, M. Fannes, and K. Życzkowski, Universal bounds for the Holevo quantity, coherent information, and the Jensen–Shannon divergence, Phys. Rev. Lett.105(2010), 040505. .
[17]
M. De Domenico, V. Nicosia, A. Arenas, and V. Latora, Structural reducibility of multilayer networks, Nat. Commun.6(2015), 6864. .
[18]
M. De Domenico, M. A. Porter, and A. Arenas, MuxViz: a tool for multilayer analysis and visualization of networks, J. Complex Netw.3(2015), 159–176. .
[19]
L. Bai, L. Rossi, A. Torsello, and E. R. Hancock, A quantum Jensen–Shannon graph kernel for unattributed graphs, Pattern Recognit.48(2015), 344–355. .
[20]
L. Rossi, A. Torsello, E. R. Hancock, and R. C. Wilson, Characterizing graph symmetries through quantum Jensen–Shannon divergence, Phys. Rev. E88(2013), 032806. .
[21]
J. Antolı́n, J. C. Angulo, and S. López-Rosa, Fisher and Jensen–Shannon divergences: quantitative comparisons among distributions. Application to position and momentum atomic densities, J. Chem. Phys.130(2009), 074110. .
[22]
J. Dixmier, Les algèbres d’opérateurs dans l’espace hilbertien (Algèbres de von Neumann), Gauthier-Villars, Paris, 1957.
[23]
P. de la Harpe, Fuglede–Kadison determinant: theme and variations, Proc. Natl. Acad. Sci. USA110(2013), no. 52, 20800–20806. .
[24]
T. Fack and H. Kosaki, Generalized \(s\)-numbers of \(\tau\)-measurable operators, Pacific J. Math.123(1986), no. 2, 269–300. .
[25]
I. J. Schoenberg, Metric spaces and completely monotone functions, Ann. of Math. (2)39(1938), 811–841. .
[26]
C. Berg, J. P. R. Christensen, and P. Ressel, Harmonic Analysis on Semigroups, Graduate Texts in Mathematics, vol. 100, Springer-Verlag, New York, 1984. .
[27]
F. Hansen, The fast track to Löwner’s theorem, Linear Algebra Appl.438(2013), no. 11, 4557–4571. .
[28]
F. Bendat and S. Sherman, Monotone and convex operator functions, Trans. Amer. Math. Soc.79(1955), 58–71. .