June 03, 2026
We propose a bipartite entanglement measure \(E(\rho)\) defined as the minimal order-1 quantum Wasserstein distance from a state to the set of separable states. Owing to the universal data-processing inequality of the Wasserstein metric, the measure satisfies all fundamental axioms within a single geometric framework. A Lipschitz dual formulation yields explicit lower bounds for pure and mixed states, a sharp constant for two-qubit systems, and an expected value for Haar-random pure states. We further establish a quantitative connection to entanglement witnesses: any negative witness expectation value certifies a lower bound on \(E\), and the dual variational bound is exactly the maximal violation achievable by a Lipschitz-1 witness. The approach naturally provides subadditivity, trace-distance estimates, and bounds on local observables, while pointing toward large-deviation conjectures. This work furnishes a versatile paradigm at the interface of entanglement theory, optimal transport, and experimental entanglement detection.
Quantifying quantum entanglement is a central challenge in quantum information theory [1]. A proper entanglement measure must satisfy a set of fundamental requisites: convexity, vanishing on separable states, invariance under local unitary operations, monotonicity under local operations and classical communication (LOCC), and continuity [1]–[3]. Over the past decades numerous entanglement measures have been introduced, including the entanglement of formation [4], [5] based on the convex roof construction, the distillable entanglement [3], [6], the relative entropy of entanglement [7], [8], the negativity [9], the geometric measure [10], and many others [11]–[13]. Among them, distance-based measures, defined as the minimal distance from a state to the set of separable states, form a particularly natural class. Well-known examples use the relative entropy [7], the Bures metric [14], the trace distance [15], [16], or the Hilbert–Schmidt distance [17]. While intuitive, these distances do not automatically guarantee LOCC monotonicity; the proof often requires specific properties of the chosen distance and may fail for even slightly modified choices (e.g., the Hilbert–Schmidt distance is not LOCC monotone [13]). Moreover, explicit computable lower bounds are rare and usually rely on case-by-case constructions such as entanglement witnesses [18]–[20] or correlation tensors [21], [22].A recurring obstacle in distance-based entanglement measures is the verification of LOCC monotonicity. While the relative entropy of entanglement relies on the joint convexity of relative entropy, most metric-based proposals do not share this property. For instance, the Hilbert–Schmidt distance fails to be LOCC monotone [13], and even the trace distance requires careful handling of optimisation over all separable states. In contrast, the order‑1 quantum Wasserstein distance \(W_1\) recently introduced by De Palma et al. [23] satisfies the data-processing inequality \[W_1(\Lambda(\rho),\Lambda(\sigma)) \le W_1(\rho,\sigma)\] for every CPTP map \(\Lambda\). This property, which is the exact metric analogue of LOCC monotonicity, makes \(W_1\) an ideal building block: defining \[E(\rho) = \inf_{\sigma\in\mathrm{SEP}} W_1(\rho,\sigma)\] immediately endows \(E\) with all five fundamental requirements without any additional ad-hoc arguments. To the best of our knowledge, this is the first entanglement measure whose complete set of axioms follows directly from a single geometric data-processing inequality.
A recent breakthrough in quantum information geometry is the introduction of quantum Wasserstein distances [23]–[25]. In particular, the order-1 quantum Wasserstein distance \(W_1\) of De Palma et al. [23] possesses the key data-processing inequality \[\label{eq:dpi} W_1(\Lambda(\rho),\Lambda(\sigma))\le W_1(\rho,\sigma)\tag{1}\] for every completely positive trace-preserving (CPTP) map \(\Lambda\), which is the exact metric analogue of the LOCC monotonicity requirement. The distance is built from a cost operator based on a complete set of local observables and a dual Lipschitz seminorm \(\|\cdot\|_{\rm Lip}\) that involves commutators with local operators. The data-processing inequality makes \(W_1\) an ideal building block for an entanglement measure: defining \(E(\rho)=\inf_{\sigma\in\mathrm{SEP}} W_1(\rho,\sigma)\) automatically yields a functional that inherits all five fundamental requirements from the geometric features of \(W_1\), as we prove in a single compact proposition (Proposition 2 of the original paper).
In this work we introduce and systematically study the Wasserstein entanglement measure \(E\). Our principal technical tool is the dual representation \(W_1(\rho,\sigma)=\mathop{\rm max}_{\|\!H\!\|_{\rm Lip}\le1}({\rm Tr}[H\rho]-{\rm Tr}[H\sigma])\), which together with Sion’s minimax theorem gives the variational formula \(E(\rho)=\mathop{\rm max}_{\|\!H\!\|_{\rm Lip}\le1}({\rm Tr}[H\rho]-\alpha(H))\), where \(\alpha(H)=\mathop{\rm max}_{\sigma\in\mathrm{SEP}}{\rm Tr}[H\sigma]\) is the support functional of the separable set. This dual formulation, whose validity is rooted in the works of De Palma et al. [23], [26], allows us to derive concrete lower bounds by choosing suitable test observables \(H\). For a pure state \(|\psi\rangle\) with Schmidt coefficients \(\{\mu_i\}\), we choose \(H_0=|\Phi^+\rangle\langle\Phi^+|\), the projector onto the maximally entangled state, and prove \(\|H_0\|_{\rm Lip}\le1\), leading to the lower bound \[\label{eq:purebound} E(|\psi\rangle\langle\psi|)\ge\frac{1}{d}\Bigl(\Bigl(\sum_{i=1}^d\sqrt{\mu_i}\Bigr)^2-1\Bigr),\tag{2}\] with equality for the maximally entangled state. This bound immediately yields an expected value for Haar-random pure states: \(\mathbb{E}[E(\rho)]\ge\frac{\pi}{4}(1-1/d)\), showing that large amounts of entanglement are typical. For mixed states we work within the Bloch expansion with respect to the local orthonormal bases \(\{F_i^A\},\{F_j^B\}\): \[\rho=\frac{I}{d_Ad_B}+\frac{1}{d_B}\sum_i s_i F_i^A\otimes I+\frac{1}{d_A}\sum_j t_j I\otimes F_j^B+\sum_{i,j}r_{ij}F_i^A\otimes F_j^B,\] and denote by \(R=(r_{ij})\) the correlation matrix.By relating the Lipschitz seminorm to the entrywise \(\ell_1\) norm of the coefficient matrix, we obtain a universal lower bound expressed in terms of the rank \(r\) of the correlation matrix \(R\) \[E(\rho) \ge \frac{1}{2\sqrt{r\,(d_A^2-1)(d_B^2-1)}} \bigl( \|R\|_{\mathrm{Tr}} - c_A c_B \bigr),\] where \(c_A = \sqrt{1-1/d_A}\) and \(c_B = \sqrt{1-1/d_B}\). This bound, though not necessarily tight in all dimensions, provides an explicit constant that depends only on the local dimensions and the rank of \(R\). Remarkably, for the standard Pauli basis in a two-qubit system we obtain the sharp estimate \(\|H\|_{\mathrm{Lip}} \le 3\|\mathbf{H}\|_{\infty}\), which leads to the explicit constant \(1/3\). Furthermore, we bridge the abstract geometric measure to quantities that are directly accessible in experiments. By exploiting the dual formulation, we prove that any entanglement witness \(W\) with Lipschitz seminorm \(\|W\|_{\mathrm{Lip}}\le L\) yields a quantitative lower bound \(E(\rho)\ge -\operatorname{Tr}[W\rho]/L\) (Theorem 5). In particular, a measured violation \(\operatorname{Tr}[W\rho]=-\varepsilon<0\) immediately certifies \(E(\rho)\ge \varepsilon/L\). We then show that the dual lower bound 8 is itself realized by an optimal Lipschitz-1 witness (Proposition 3), endowing the variational principle with a transparent operational meaning and enabling semidefinite programming approaches for estimating \(E\). Our geometric approach further yields a collection of cross-disciplinary bounds. We prove subadditivity under tensor products \(E(\rho_1\otimes\rho_2)\le E(\rho_1)+E(\rho_2)\), a trace-distance bound \(\inf_{\sigma\in\mathrm{SEP}}\|\rho-\sigma\|_1\le 2E(\rho)\), and a bound for local observables \(|\langle A\otimes B\rangle_\rho-\mathop{\rm max}_{\sigma\in\mathrm{SEP}}\langle A\otimes B\rangle_\sigma|\le 2E(\rho)\). Moreover, we discuss the asymptotic behaviour of the overlap \(\inf_{\sigma_n\in\mathrm{SEP}_n} \operatorname{Tr}[\rho^{\otimes n}\sigma_n]\), which points towards a quantum Sanov theorem with rate function \(E(\rho)\); a full proof of the equality is left as an open problem. From a broader perspective, the Wasserstein entanglement measure embeds entanglement theory into the mature field of optimal transport and non-commutative metric geometry [24], [27], [28]. The commutator condition \(\|[X,H]\|\le1\) is a quantum analogue of the classical \(1\)-Lipschitz condition, and the state space equipped with \(W_1\) becomes a quantum metric space. This connection suggests new bridges to free probability [29], [30], quantum large deviations [31], [32], and convex geometry [33]. We also formulate several conjectures, including a quantum Talagrand-type inequality and a conjectured relation to hyperfiniteness of von Neumann algebras, which hint at deep links with the modular theory and the geometry of entangled states in infinite-dimensional systems [34]. The paper is structured as follows. In Section II we introduce the cost operator and define \(W_1\) and \(E\). Section III proves the five fundamental properties in a single proposition. Lower bounds for pure and mixed states are derived in Sections IV and V, respectively. Section VI collects further cross-disciplinary bounds, including the connection to entanglement witnesses (Theorems 5 and 3), subadditivity, trace-distance estimates, and a large-deviation conjecture. Section VII discusses the physical significance and mathematical outlook, and we conclude in Section VIII.
In this section we introduce the mathematical setting of our theory. We first specify a set of local observables and construct the cost operator that defines the quantum Wasserstein distance. Then we state the distance itself and our new entanglement measure. Finally we derive a Lipschitz‑dual formulation that will be the central tool in all later proofs.
Let \(\mathcal{H}_A=\mathbb{C}^{d_A}\), \(\mathcal{H}_B=\mathbb{C}^{d_B}\) and \(\mathcal{H}_{AB}=\mathcal{H}_A\otimes\mathcal{H}_B\). For system \(A\) we choose a family \(\{F_i^A\}_{i=1}^{d_A^2-1}\) of traceless Hermitian matrices satisfying \[\label{eq:orth} \mathop{\rm Tr}(F_i^A F_j^A)=\delta_{ij},\qquad \|F_i^A\|\le 1\quad\forall\,i.\tag{3}\] Such a family can be obtained by a suitable rescaling of the generalised Gell‑Mann matrices; for instance the Pauli matrices divided by \(\sqrt2\) work for qubits. A similar family \(\{F_j^B\}_{j=1}^{d_B^2-1}\) is fixed for system \(B\). On \(\mathcal{H}_{AB}\) we define \[X_i^A = F_i^A\otimes I_B,\qquad X_j^B = I_A\otimes F_j^B .\] Set \({\cal K}_0=\{X_i^A\}\cup\{X_j^B\}\). The cost operator on \(\mathcal{H}_{AB}\otimes\mathcal{H}_{AB}\) is \[\label{eq:C} C = \sum_{X\in{\cal K}_0} (X\otimes I_{AB} - I_{AB}\otimes X)^2 .\tag{4}\]
For two states \(\rho,\sigma\in{\cal D}(\mathcal{H}_{AB})\) denote by \({\cal C}(\rho,\sigma)\) the set of quantum couplings (density operators on \(\mathcal{H}_{AB}\otimes\mathcal{H}_{AB}\) with marginals \(\rho,\sigma\)). The quantum Wasserstein distance of order \(1\) is \[W_1(\rho,\sigma)=\Bigl(\inf_{\tau\in{\cal C}(\rho,\sigma)}\mathop{\rm Tr}[C\tau]\Bigr)^{1/2}.\]
Definition 1. The Wasserstein entanglement measure* is \[{\cal E}(\rho)=\inf_{\sigma\in\mathcal{SEP}(\mathcal{H}_{AB})} W_1(\rho,\sigma),\] where \(\mathcal{SEP}(\mathcal{H}_{AB})\) denotes the set of separable states.*
Having defined the distance, we now introduce a convenient Lipschitz seminorm that allows us to exploit convex duality. For any self‑adjoint \(H\in\mathcal{B}(\mathcal{H}_{AB})\) define \[\label{eq:lipdef} \|H\|_{\mathrm{Lip}}= \mathop{\rm max}\bigl\{ \sup_{A:\|A\|\le1}\|[A\otimes I_B , H]\|,\; \sup_{B:\|B\|\le1}\|[I_A\otimes B , H]\| \bigr\}.\tag{5}\] Because \({\cal K}_0\subset\{A\otimes I:\|A\|\le1\}\cup\{I\otimes B:\|B\|\le1\}\), the condition \(\|H\|_{\mathrm{Lip}}\le1\) implies \(\|[X,H]\|\le1\) for all \(X\in{\cal K}_0\). As proven in the seminal works on quantum Wasserstein distances [23], [26], the order‑1 distance satisfies the general dual representation \[\label{eq:W1dualexact} W_1(\rho,\sigma) = \sup\bigl\{ \left|\operatorname{Tr}[H(\rho-\sigma)]\right| : \|[L_r, H]\| \le 1 \text{ for all } L_r\in\mathcal{K}_0 \bigr\}.\tag{6}\] Our Lipschitz seminorm is defined via a larger set of commutators, hence \(\|H\|_{\mathrm{Lip}}\le 1\) guarantees that \(H\) is admissible in the above dual. Consequently we obtain \[\label{eq:W1dual} W_1(\rho,\sigma) \ge \mathop{\rm max}_{H:\|H\|_{\mathrm{Lip}}\le 1} \bigl( \operatorname{Tr}[H\rho] - \operatorname{Tr}[H\sigma] \bigr).\tag{7}\] Taking the infimum over \(\sigma\in\mathcal{SEP}\) and exchanging \(\inf\) and \(\mathop{\rm max}\) (justified by compactness and convexity) we arrive at the fundamental inequality \[\label{eq:emax} {\cal E}(\rho) \ge \mathop{\rm max}_{H:\|H\|_{\mathrm{Lip}}\le 1} \Bigl( \operatorname{Tr}[H\rho] - \mathop{\rm max}_{\sigma\in\mathcal{SEP}}\operatorname{Tr}[H\sigma] \Bigr).\tag{8}\] Define the support functional of the separable set \[\label{ah} \alpha(H) = \mathop{\rm max}_{\sigma\in\mathcal{SEP}}\operatorname{Tr}[H\sigma].\tag{9}\] Then \({\cal E}(\rho) \ge \mathop{\rm max}_{H:\|H\|_{\mathrm{Lip}}\le 1}\bigl( \operatorname{Tr}[H\rho] - \alpha(H) \bigr)\).
Proposition 2. The functional \({\cal E}\) satisfies:
\({\cal E}(\rho)=0\) for every \(\rho\in\mathcal{SEP}\);
convexity: \({\cal E}\bigl(\lambda\rho_1+(1-\lambda)\rho_2\bigr)\le \lambda{\cal E}(\rho_1)+(1-\lambda){\cal E}(\rho_2)\), \(0\le\lambda\le1\);
local unitary invariance: \({\cal E}\bigl((U_A\otimes U_B)\rho(U_A^\dagger\otimes U_B^\dagger)\bigr)={\cal E}(\rho)\) for all unitaries \(U_A,U_B\);
LOCC monotonicity: \({\cal E}(\Lambda(\rho))\le{\cal E}(\rho)\) for any LOCC channel \(\Lambda\);
continuity: \({\cal E}\) is continuous with respect to the trace norm on the finite‑dimensional state space.
Proof. (i) If \(\rho\in\mathcal{SEP}\), choose \(\sigma=\rho\) in Definition 1; \(W_1(\rho,\rho)=0\), hence \({\cal E}(\rho)=0\). (ii) The distance \(W_1\) is induced by a norm (the Lipschitz dual norm), therefore it is convex in each argument. For any \(\sigma\in\mathcal{SEP}\), \[W_1(\lambda\rho_1+(1-\lambda)\rho_2,\sigma)\le \lambda W_1(\rho_1,\sigma)+(1-\lambda)W_1(\rho_2,\sigma).\] Taking the infimum over \(\sigma\) yields the convexity of \({\cal E}\). (iii) The cost operator \(C\) in 4 is built from the orthonormal bases \(\{F_i^A\}\) and \(\{F_j^B\}\). For a local unitary \(U=U_A\otimes U_B\), the families \(\{U_AF_i^AU_A^\dagger\}\) and \(\{U_BF_j^BU_B^\dagger\}\) are again traceless, orthonormal and satisfy the same norm bound. Hence the set \({\cal K}_0' = \{U_AF_i^AU_A^\dagger\otimes I_B,\; I_A\otimes U_BF_j^BU_B^\dagger\}\) is another orthonormal basis of the same local operator spaces and \[\sum_{X'\in{\cal K}_0'} (X'\otimes I - I\otimes X')^2 = C .\] Consequently \(W_1(U\rho U^\dagger, U\sigma U^\dagger)=W_1(\rho,\sigma)\) for all \(\rho,\sigma\). Because \(U\mathcal{SEP}U^\dagger=\mathcal{SEP}\), we obtain \[{\cal E}(U\rho U^\dagger)=\inf_{\sigma\in\mathcal{SEP}}W_1(U\rho U^\dagger,\sigma) = \inf_{\sigma\in\mathcal{SEP}}W_1(\rho, U^\dagger\sigma U) = {\cal E}(\rho).\] (iv) Every LOCC channel \(\Lambda\) is CPTP. By the data processing inequality, \(W_1(\Lambda(\rho),\Lambda(\sigma))\le W_1(\rho,\sigma)\) for all \(\sigma\). Since \(\Lambda(\mathcal{SEP})\subseteq\mathcal{SEP}\), \[{\cal E}(\Lambda(\rho))=\inf_{\sigma'\in\mathcal{SEP}}W_1(\Lambda(\rho),\sigma') \le \inf_{\sigma\in\mathcal{SEP}}W_1(\Lambda(\rho),\Lambda(\sigma)) \le \inf_{\sigma\in\mathcal{SEP}}W_1(\rho,\sigma)={\cal E}(\rho).\] (v) On the finite‑dimensional space, the map \(\rho\mapsto W_1(\rho,\sigma)\) is Lipschitz continuous with respect to the trace norm, uniformly in \(\sigma\) [26]. The set \(\mathcal{SEP}\) is compact, therefore the infimum of a family of equi‑Lipschitz functions is continuous. \(\sqcup\) Proposition 2 demonstrates a significant advantage of our approach: all five axioms are proved in a single unified proposition, while for most existing measures (e.g., entanglement of formation, relative entropy of entanglement, negativity) one or more properties require separate, often technically involved proofs. Convexity follows from the convexity of \(W_1\), monotonicity from its universal data-processing inequality, and continuity from the compactness of separable states. This economy is possible because \(W_1\) is built from a cost operator that respects the bipartite structure and because its dual Lipschitz norm faithfully captures the non-local character of quantum states. Consequently, the Wasserstein entanglement measure provides not only a new quantifier but also a new framework in which the axiomatic properties of entanglement are transparent consequences of the underlying metric geometry.
Let \(\rho=|\psi\rangle\langle\psi|\) be pure with Schmidt decomposition \(|\psi\rangle=\sum_{i=1}^{d}\sqrt{\mu_i}\,|u_i\rangle_A\otimes|v_i\rangle_B\), where \(\mu_1\ge\cdots\ge\mu_d>0\), \(\sum_i\mu_i=1\) and \(d=\mathop{\rm min}\{d_A,d_B\}\ge 2\) (otherwise \({\cal E}\equiv0\) trivially).
Theorem 1. For every pure state, \[{\cal E}(\rho)\;\ge\; \frac{1}{d}\left(\Bigl(\sum_{i=1}^d\sqrt{\mu_i}\Bigr)^2-1\right) \;\ge\; 0.\] For the maximally entangled state (\(\mu_i=1/d\)), the right‑hand side equals \((d-1)/d\).
Proof. Define the maximally entangled state projector \[H_0 = |\Phi^+\rangle\langle\Phi^+| = \frac{1}{d}\sum_{i,j=1}^d |u_i v_i\rangle\langle u_j v_j|.\] We first verify \(\|H_0\|_{\mathrm{Lip}}\le1\). Let \(A\in\mathcal{B}(\mathcal{H}_A)\) with \(\|A\|\le1\). Since \(H_0\) is a rank‑one projection, a direct computation yields \[\|[A\otimes I_B , H_0]\|^2 = \langle\Phi^+|(A^\dagger A)\otimes I_B|\Phi^+\rangle - |\langle\Phi^+|A\otimes I_B|\Phi^+\rangle|^2 = \frac{1}{d} \operatorname{Tr}(A^\dagger A) - \frac{|\operatorname{Tr}A|^2}{d^2}.\] Because \(\|A\|\le1\), we have \(\operatorname{Tr}(A^\dagger A)\le d\,\|A\|^2\le d\), hence \[\|[A\otimes I_B , H_0]\|^2 \le \frac{d}{d} = 1,\] so \(\|[A\otimes I_B , H_0]\|\le1\). An identical bound holds for \(I_A\otimes B\), and therefore \(\|H_0\|_{\mathrm{Lip}}\le1\). Now evaluate the two terms. Because \(|\psi\rangle\) is supported on \(span\{|u_i v_i\rangle\}\), \[\operatorname{Tr}[H_0\rho] = \langle\psi|H_0|\psi\rangle = \frac{1}{d}\Bigl(\sum_{i=1}^d\sqrt{\mu_i}\Bigr)^2 .\] For any product state \(\sigma_A\otimes\sigma_B\), write \(\sigma_A = \sum_k p_k|\alpha_k\rangle\langle\alpha_k|\), \(\sigma_B = \sum_l q_l|\beta_l\rangle\langle\beta_l|\). Then \[\operatorname{Tr}[H_0(\sigma_A\otimes\sigma_B)] = \frac{1}{d} \sum_{k,l} p_k q_l \,|\langle\alpha_k|\beta_l^*\rangle|^2 \le \frac{1}{d},\] where \(|\beta_l^*\rangle\) denotes complex conjugation in the Schmidt basis. Since any separable state is a convex combination of product states, \(\alpha(H_0)=1/d\). Inserting these values into 8 yields \[{\cal E}(\rho) \ge \operatorname{Tr}[H_0\rho] - \alpha(H_0) = \frac{1}{d}\Bigl(\bigl(\sum_{i=1}^d\sqrt{\mu_i}\bigr)^2 - 1\Bigr).\] Non‑negativity follows from \(\sum_i\sqrt{\mu_i}\ge\sqrt{\sum_i\mu_i}=1\). For the maximally entangled state the bound equals \((d-1)/d\). \(\sqcup\)
We now extend the analysis to arbitrary mixed states. The Bloch expansion with respect to the local bases \(\{F_i^A\},\{F_j^B\}\) from 3 plays a central role. Any bipartite state \(\rho\) can be uniquely written as \[\label{eq:bloch} \rho= \frac{I_{AB}}{d_A d_B} + \frac{1}{d_B}\sum_{i=1}^{d_A^2-1} s_i\,F_i^A\otimes I_B + \frac{1}{d_A}\sum_{j=1}^{d_B^2-1} t_j\,I_A\otimes F_j^B + \sum_{i,j} r_{ij}\,F_i^A\otimes F_j^B,\tag{10}\] with the correlation matrix \(R=(r_{ij})\) where \[\label{eq:coeff} r_{ij} = \mathop{\rm Tr}\bigl[\rho\,(F_i^A\otimes F_j^B)\bigr].\tag{11}\]
We consider observables of the form \[\label{eq:Hform} H = \sum_{i,j} h_{ij}\,F_i^A\otimes F_j^B,\tag{12}\] with a real matrix \({\boldsymbol{H}}=(h_{ij})\).
Lemma 1. For every \(H\) of the form 12 , \[\|H\|_{\mathrm{Lip}} \le 2\sum_{i,j} |h_{ij}| = 2\|{\boldsymbol{H}}\|_{1,1}.\]
Proof. Take \(A\in\mathcal{B}(\mathcal{H}_A)\) with \(\|A\|\le1\). Then \[[A\otimes I_B, H] = \sum_{i,j} h_{ij}\,[A,F_i^A]\otimes F_j^B .\] Using the operator norm and the fact \(\|F_j^B\|\le1\), we bound \[\|[A\otimes I_B, H]\| \le \sum_{i,j} |h_{ij}|\, \|[A,F_i^A]\| \le 2\sum_{i,j} |h_{ij}|,\] because \(\|[A,F_i^A]\|\le 2\|A\|\le 2\). The same estimate holds for \(I_A\otimes B\). Hence \(\|H\|_{\mathrm{Lip}}\le 2\|{\boldsymbol{H}}\|_{1,1}\). \(\sqcup\)
The following lemma is unchanged; its proof relies only on the orthonormality of the bases.
Lemma 2. For \(H\) of the form 12 , define \[c_A = \sqrt{1-\frac{1}{d_A}},\qquad c_B = \sqrt{1-\frac{1}{d_B}}.\] Then \[\alpha(H) = \mathop{\rm max}_{\sigma\in\mathcal{SEP}}\mathop{\rm Tr}[H\sigma] = c_A\,c_B\,\|{\boldsymbol{H}}\|_{\infty}.\]
Proof. Any product state can be written as \(\sigma_A = I/d_A + \sum_i u_i F_i^A\) and \(\sigma_B = I/d_B + \sum_j v_j F_j^B\) with real vectors \(u=(u_i)\), \(v=(v_j)\). Since \(\{F_i^A\}\) and \(\{F_j^B\}\) are orthonormal and traceless, \[\mathop{\rm Tr}[(\sigma_A)^2] = \frac{1}{d_A} + \|u\|_2^2 \le 1,\qquad \mathop{\rm Tr}[(\sigma_B)^2] = \frac{1}{d_B} + \|v\|_2^2 \le 1,\] hence \(\|u\|_2 \le c_A\) and \(\|v\|_2 \le c_B\). For \(H = \sum_{i,j} h_{ij} F_i^A\otimes F_j^B\), \[\mathop{\rm Tr}[H\sigma] = \sum_{i,j} h_{ij} u_i v_j = u^{\mathsf T}{\boldsymbol{H}}v \le \|{\boldsymbol{H}}\|_{\infty}\,\|u\|_2\,\|v\|_2 \le c_A c_B \|{\boldsymbol{H}}\|_{\infty}.\] The bound is attained by choosing product states whose Bloch vectors are proportional to the left and right singular vectors of \({\boldsymbol{H}}\). \(\sqcup\)
Using the dual formulation together with Lemma 1 we obtain a first general lower bound in terms of the entrywise norm and the rank of the correlation matrix.
Theorem 2. Let \(\rho\) have correlation matrix \(R\), and denote \(r = \mathop{\rm rank}(R)\), \(n_A = d_A^2-1\), \(n_B = d_B^2-1\). Then \[{\cal E}(\rho) \ge \frac{1}{2\sqrt{r\, n_A n_B}} \bigl( \|R\|_{\mathop{\rm Tr}} - c_A c_B \bigr).\]
Proof. From 8 and Lemma 2, \[{\cal E}(\rho) \ge \mathop{\rm max}_{H:\|H\|_{\mathrm{Lip}}\le 1} \bigl( \operatorname{Tr}[H\rho] - c_A c_B \|{\boldsymbol{H}}\|_{\infty} \bigr).\] Restrict to \(H\) of the form 12 . Let \(R = U\Sigma V^{\mathsf T}\) be a singular value decomposition, \(\Sigma = \mathop{\rm diag}(\sigma_1,\dots,\sigma_r,0,\dots,0)\). Define \(\Sigma' = \mathop{\rm diag}(\operatorname{sgn}(\sigma_1),\dots,\operatorname{sgn}(\sigma_r),0,\dots,0)\) and \(S = U\Sigma' V^{\mathsf T}\). Then \(\|S\|_{\infty}=1\) and \(\operatorname{Tr}(S^{\mathsf T} R) = \|R\|_{\mathop{\rm Tr}}\).
By Lemma 1, \(\|S\|_{\mathrm{Lip}} \le 2\|S\|_{1,1}\). Moreover, \[\|S\|_{1,1} = \sum_{i,j} |S_{ij}| \le \sqrt{n_A n_B} \, \|S\|_{\mathrm{F}} = \sqrt{n_A n_B \, r}.\] Therefore \(\|S\|_{\mathrm{Lip}} \le 2\sqrt{r n_A n_B}\). Choose \(t = 1/(2\sqrt{r n_A n_B})\) and set \(H = t S\). Then \(\|H\|_{\mathrm{Lip}}\le 1\) and \(\|\mathbf{H}\|_{\infty}=t\). Substituting into the dual bound, \[{\cal E}(\rho) \ge t \,\operatorname{Tr}(S^{\mathsf T} R) - c_A c_B\,t = t\bigl( \|R\|_{\mathop{\rm Tr}} - c_A c_B \bigr).\] Insert \(t\) finishes the proof. \(\sqcup\)
In the important case of two qubits, a more refined estimate on the Lipschitz constant can be obtained, recovering a stronger constant.
Theorem 3. For a two‑qubit system (\(d_A=d_B=2\)) choose the Pauli basis \(F_i = \sigma_i/\sqrt2\), \(i=1,2,3\). Then, for any state \(\rho\), \[{\cal E}(\rho) \ge \frac{1}{3}\Bigl( \|R\|_{\mathop{\rm Tr}} - \frac{1}{2} \Bigr),\] where \(R_{ij} = \frac{1}{2}\operatorname{Tr}[\rho(\sigma_i\otimes\sigma_j)]\) is the usual correlation matrix.
Proof. Recall \(H = \frac{1}{2}\sum_{i,j} h_{ij}\sigma_i\otimes\sigma_j\). Write \({\boldsymbol{H}}= (h_{ij})\) and take its singular value decomposition \({\boldsymbol{H}}= U\Sigma V^{\mathsf T}\) with \(\Sigma = \mathop{\rm diag}(s_1,s_2,s_3)\), \(s_1\ge s_2\ge s_3\ge0\). Define \[A_k = \sum_{i=1}^3 U_{ik}\sigma_i,\qquad B_k = \sum_{j=1}^3 V_{jk}\sigma_j .\] Because \(U,V\) are orthogonal and \(\{\sigma_i\}\) satisfy \(\operatorname{Tr}(\sigma_i\sigma_j)=2\delta_{ij}\), one verifies \(\|A_k\|=\|B_k\|=1\). Moreover, \[H = \frac{1}{2}\sum_{k=1}^3 s_k \, A_k\otimes B_k .\] Now for any \(C\) with \(\|C\|\le1\), \[\| [C\otimes I, H] \| \le \frac{1}{2}\sum_{k=1}^3 s_k\,\|[C,A_k]\|\,\|B_k\| \le \frac{1}{2}\sum_{k=1}^3 s_k\cdot 2\|C\|\|A_k\| = \sum_{k=1}^3 s_k .\] The same estimate holds for \(I\otimes D\). Since \(\sum_{k=1}^3 s_k \le 3 s_1 = 3\|{\boldsymbol{H}}\|_{\infty}\), we obtain \(\|H\|_{\mathrm{Lip}} \le 3\|{\boldsymbol{H}}\|_{\infty}\). Now apply the dual formulation 8 together with Lemma 2 (note that for the Pauli basis \(c_A=c_B=1/\sqrt2\), hence \(c_Ac_B=1/2\)): \[{\cal E}(\rho) \ge \mathop{\rm max}_{\|{\boldsymbol{H}}\|_{\infty}\le 1/3} \bigl( \operatorname{Tr}({\boldsymbol{H}}^{\mathsf T} R) - \tfrac12 \|{\boldsymbol{H}}\|_{\infty} \bigr) \ge \frac{1}{3}\|R\|_{\mathop{\rm Tr}} - \frac{1}{6} = \frac{1}{3}\bigl( \|R\|_{\mathop{\rm Tr}} - \tfrac12 \bigr).\] \(\sqcup\)
We first recall the fundamental dual inequality that has been used throughout
Theorem 4. For any state \(\rho\), \[{\cal E}(\rho) \;\ge\; \sup_{H:\|H\|_{\mathrm{Lip}}\le 1} \Bigl( \mathop{\rm Tr}[H\rho] - \mathop{\rm max}_{\sigma\in\mathcal{SEP}}\mathop{\rm Tr}[H\sigma] \Bigr).\]
Proof. The original quantum Wasserstein distance satisfies (see [26]) \[W_1(\rho,\sigma) \ge \mathop{\rm max}_{H:\|H\|_{\mathrm{Lip}}\le 1} \bigl(\mathop{\rm Tr}[H\rho]-\mathop{\rm Tr}[H\sigma]\bigr).\] Taking the infimum over \(\sigma\in\mathcal{SEP}\) and exchanging \(\inf\) and \(\mathop{\rm max}\) (justified by compactness and convexity) yields \[{\cal E}(\rho) \ge \sup_{H:\|H\|_{\mathrm{Lip}}\le 1}\inf_{\sigma\in\mathcal{SEP}}\bigl(\mathop{\rm Tr}[H\rho]-\mathop{\rm Tr}[H\sigma]\bigr) = \sup_{H:\|H\|_{\mathrm{Lip}}\le 1} \Bigl( \mathop{\rm Tr}[H\rho] - \mathop{\rm max}_{\sigma\in\mathcal{SEP}}\mathop{\rm Tr}[H\sigma] \Bigr).\] The inequality is generally strict; an equality would require a stronger Lipschitz condition matching exactly the original dual of \(W_1\). \(\sqcup\)
A widely used tool for detecting entanglement in the laboratory are entanglement witnesses [20], [35]. We now show that the Wasserstein entanglement measure \(E(\rho)\) provides a direct quantitative upgrade of any witness: the amount by which a state violates a witness gives a concrete lower bound on its geometric entanglement.
Theorem 5. Let \(W\) be a self‑adjoint operator on \(\mathcal{H}_{AB}\) which is an entanglement witness, i.e. \[\operatorname{Tr}[W\sigma] \ge 0 \qquad \forall\, \sigma \in \mathrm{SEP},\] and assume that the witness is normalised so that \(\mathop{\rm min}_{\sigma\in\mathrm{SEP}} \operatorname{Tr}[W\sigma]=0\) (this can always be achieved by adding a suitable multiple of the identity). If the Lipschitz seminorm of \(W\) satisfies \(\|W\|_{\mathrm{Lip}} \le L\) for some \(L>0\), then for every bipartite state \(\rho\), \[E(\rho) \;\ge\; \frac{-\operatorname{Tr}[W\rho]}{L}. \label{eq:witnessbound}\qquad{(1)}\] In particular, if an experiment measures \(\operatorname{Tr}[W\rho] = -\varepsilon < 0\), it immediately follows that \(E(\rho) \ge \varepsilon/L\).
Proof. In the dual lower bound 8 choose the trial observable \(H = -W/L\). Because \(\|H\|_{\mathrm{Lip}} = \|W\|_{\mathrm{Lip}}/L \le 1\), it is admissible. The support functional 9 gives \[\alpha(H) = \mathop{\rm max}_{\sigma\in\mathrm{SEP}} \operatorname{Tr}[H\sigma] = -\frac{1}{L}\,\mathop{\rm min}_{\sigma\in\mathrm{SEP}} \operatorname{Tr}[W\sigma] = 0,\] where the last equality uses the witness property and the normalisation condition. Inserting \(H\) and \(\alpha(H)\) into the dual bound yields \[E(\rho) \;\ge\; \operatorname{Tr}[H\rho] - \alpha(H) \;=\; -\frac{1}{L}\,\operatorname{Tr}[W\rho],\] which is exactly ?? . \(\sqcup\) The constant \(L\) can be estimated using the techniques of Lemmas 4 and Theorem 7. For instance, in the two‑qubit setting with the Pauli basis \(\sigma_i/\sqrt{2}\), one often uses the optimal entanglement witness \(W = \frac{1}{4}\bigl(I\otimes I - \sum_{i=1}^3 \sigma_i\otimes\sigma_i\bigr)\), for which a straightforward application of Lemma 4 gives \(L \le 3/2\); thus any measured negative expectation value \(\langle W\rangle_\rho = -\varepsilon\) immediately implies \(E(\rho) \ge \frac{2}{3}\,\varepsilon\).
Proposition 3. For any bipartite state \(\rho\), define the witness-optimised bound \[\mathcal{W}(\rho) \;=\; \sup\big\{ -\operatorname{Tr}[W\rho] \;:\; W = W^\dagger,\; \operatorname{Tr}[W\sigma]\ge 0 \;\; \forall\,\sigma\in\mathrm{SEP},\; \mathop{\rm min}_{\sigma\in\mathrm{SEP}} \operatorname{Tr}[W\sigma]=0,\; \|W\|_{\mathrm{Lip}}\le 1 \big\}. \label{eq:Wdef}\qquad{(2)}\] Then \(\mathcal{W}(\rho)\) coincides exactly with the dual lower bound 8 , \[\mathcal{W}(\rho) \;=\; \mathop{\rm max}_{\|H\|_{\mathrm{Lip}}\le 1}\bigl( \operatorname{Tr}[H\rho] - \alpha(H) \bigr), \label{eq:Wdual}\qquad{(3)}\] and consequently provides a lower bound to the Wasserstein entanglement measure: \[E(\rho) \;\ge\; \mathcal{W}(\rho). \label{eq:Wbound}\qquad{(4)}\]
Proof. Let \(H\) be any self‑adjoint operator with \(\|H\|_{\mathrm{Lip}}\le 1\). Define \(W = \alpha(H) I - H\). For every separable state \(\sigma\), \[\operatorname{Tr}[W\sigma] = \alpha(H) - \operatorname{Tr}[H\sigma] \ge 0,\] by definition of \(\alpha(H) \equiv \mathop{\rm max}_{\sigma\in\mathrm{SEP}} \operatorname{Tr}[H\sigma]\), and the minimum over \(\sigma\in\mathrm{SEP}\) is zero (attained at the maximiser). Moreover, \(\|W\|_{\mathrm{Lip}} = \|H\|_{\mathrm{Lip}} \le 1\). Hence \(W\) is admissible in the supremum ?? , and we have \[-\operatorname{Tr}[W\rho] = \operatorname{Tr}[H\rho] - \alpha(H).\] Taking the supremum over all such \(H\) gives \[\mathcal{W}(\rho) \;\ge\; \mathop{\rm max}_{\|H\|_{\mathrm{Lip}}\le 1}\bigl( \operatorname{Tr}[H\rho] - \alpha(H) \bigr).\] Conversely, let \(W\) be any witness satisfying the conditions in ?? . Set \(H = -W\); clearly \(\|H\|_{\mathrm{Lip}} \le 1\). Because \(\operatorname{Tr}[W\sigma] \ge 0\) and the minimum is zero, \[\alpha(H) = \mathop{\rm max}_{\sigma\in\mathrm{SEP}} \operatorname{Tr}[-W\sigma] = - \mathop{\rm min}_{\sigma\in\mathrm{SEP}} \operatorname{Tr}[W\sigma] = 0.\] Therefore, \[\operatorname{Tr}[H\rho] - \alpha(H) = -\operatorname{Tr}[W\rho] \;\le\; \mathop{\rm max}_{\|H\|_{\mathrm{Lip}}\le 1}\bigl( \operatorname{Tr}[H\rho] - \alpha(H) \bigr).\] Since this holds for every admissible \(W\), the reverse inequality \[\mathcal{W}(\rho) \;\le\; \mathop{\rm max}_{\|H\|_{\mathrm{Lip}}\le 1}\bigl( \operatorname{Tr}[H\rho] - \alpha(H) \bigr)\] follows. The two inequalities prove ?? . The final bound ?? is now immediate from the fundamental inequality \(E(\rho) \ge \mathop{\rm max}_{\|H\|_{\mathrm{Lip}}\le 1} (\operatorname{Tr}[H\rho] - \alpha(H))\) established in 8 . \(\sqcup\)
Proposition 3 shows that the dual lower bound has a transparent operational meaning: it is the largest violation achievable by any normalised entanglement witness whose Lipschitz constant is at most \(1\). Because the Lipschitz seminorm can be bounded linearly in terms of the expansion coefficients of \(W\) in the local operator bases \(\{F_i^A\}\) and \(\{F_j^B\}\) (cf.Lemma 1), the optimisation ?? can be cast as a semi‑definite program—harnessing standard techniques for entanglement witnesses but now providing a guaranteed quantitative distance to separability.
We now turn to the behaviour of \({\cal E}\) under tensor products, a property that directly reflects the metric structure of the underlying quantum Wasserstein distance.
Theorem 6. Let \(\rho_1\in{\cal D}(\mathcal{H}_{A_1B_1})\) and \(\rho_2\in{\cal D}(\mathcal{H}_{A_2B_2})\) be bipartite states. Then \[{\cal E}(\rho_1\otimes\rho_2)\le {\cal E}(\rho_1)+{\cal E}(\rho_2).\]
Proof. For \(i=1,2\) choose separable states \(\sigma_i\) that achieve the infimum in the definition of \({\cal E}(\rho_i)\) (the infimum is attained by compactness). The tensor product \(\sigma_1\otimes\sigma_2\) is separable, therefore \[{\cal E}(\rho_1\otimes\rho_2)\le W_1(\rho_1\otimes\rho_2,\sigma_1\otimes\sigma_2).\] The cost operator for the product system is \(C_{12}=C_1\otimes I + I\otimes C_2\), because the set of local observables on \(\mathcal{H}_{A_1B_1}\otimes\mathcal{H}_{A_2B_2}\) consists exactly of operators of the form \(X_1\otimes I\) and \(I\otimes X_2\) with \(X_i\in{\cal K}_0^{(i)}\). If \(\tau_i\) is an optimal coupling between \(\rho_i\) and \(\sigma_i\), then \(\tau_1\otimes\tau_2\) is a coupling between the product states and \[\mathop{\rm Tr}[C_{12}\,(\tau_1\otimes\tau_2)]=\mathop{\rm Tr}[C_1\tau_1]+\mathop{\rm Tr}[C_2\tau_2].\] Taking the infimum over couplings gives \[W_1(\rho_1\otimes\rho_2,\sigma_1\otimes\sigma_2)^2 \le W_1(\rho_1,\sigma_1)^2+W_1(\rho_2,\sigma_2)^2.\] Using \(\sqrt{a+b}\le\sqrt a+\sqrt b\) we obtain \[W_1(\rho_1\otimes\rho_2,\sigma_1\otimes\sigma_2)\le W_1(\rho_1,\sigma_1)+W_1(\rho_2,\sigma_2) ={\cal E}(\rho_1)+{\cal E}(\rho_2),\] which completes the proof. \(\sqcup\) A further consequence of the metric nature of \(W_1\) is a simple relationship between \({\cal E}\) and the trace distance.
Theorem 7. For every bipartite state \(\rho\), \[\inf_{\sigma\in\mathcal{SEP}}\|\rho-\sigma\|_1\le 2\,{\cal E}(\rho).\]
Proof. By Lemma 3.3 of [26], for any two states \(\rho,\sigma\) one has \(\|\rho-\sigma\|_1\le 2\,W_1(\rho,\sigma)\). Taking the infimum over \(\sigma\in\mathcal{SEP}\) yields \[\inf_{\sigma\in\mathcal{SEP}}\|\rho-\sigma\|_1\le 2\inf_{\sigma\in\mathcal{SEP}}W_1(\rho,\sigma) =2\,{\cal E}(\rho).\] \(\sqcup\) More generally, the Lipschitz constraint controls the deviation of any local observable from its maximal separable value.
Theorem 8. Let \(A\in\mathcal{B}(\mathcal{H}_A),\;B\in\mathcal{B}(\mathcal{H}_B)\) be self‑adjoint with \(\|A\|\le1,\|B\|\le1\). Then \[\bigl|\mathop{\rm Tr}[(A\otimes B)\rho]-\mathop{\rm max}_{\sigma\in\mathcal{SEP}}\mathop{\rm Tr}[(A\otimes B)\sigma]\bigr| \le 2\,{\cal E}(\rho).\]
Proof. Consider \(H=A\otimes B\). For any \(X=C\otimes I\) with \(\|C\|\le1\), \[\|[X,H]\|=\|[C,A]\otimes B\|\le \|[C,A]\|\,\|B\|\le 2\|C\|\|A\|\|B\|\le2.\] The same estimate holds for \(X=I\otimes D\). Hence \(\|H\|_{\mathrm{Lip}}\le2\). Thus \(H/2\) satisfies \(\|H/2\|_{\mathrm{Lip}}\le1\) and by the dual representation \[{\cal E}(\rho)\ge \mathop{\rm Tr}[(H/2)\rho]-\alpha(H/2) =\frac{1}{2}\bigl(\mathop{\rm Tr}[H\rho]-\alpha(H)\bigr).\] Because \(\alpha(H)=\mathop{\rm max}_{\sigma\in\mathcal{SEP}}\mathop{\rm Tr}[H\sigma]\), rearranging gives \(\mathop{\rm Tr}[H\rho]-\alpha(H)\le 2{\cal E}(\rho)\). Replacing \(H\) by \(-H\) yields the absolute value. \(\sqcup\)
Theorem 9. Let \(\rho\) be a pure state drawn from the Haar measure on \(\mathbb{C}^d\otimes\mathbb{C}^d\) (\(d\ge2\)). Then \[\mathbb{E}\bigl[{\cal E}(\rho)\bigr]\;\ge\;\frac{\pi}{4}\Bigl(1-\frac{1}{d}\Bigr).\]
Proof. By Theorem 1, \[{\cal E}(\rho)\ge\frac{1}{d}\Bigl(\bigl(\sum_{i=1}^d\sqrt{\mu_i}\bigr)^2-1\Bigr),\] hence \[\mathbb{E}[{\cal E}(\rho)]\ge\frac{1}{d}\Bigl(\mathbb{E}\bigl[(\sum_i\sqrt{\mu_i})^2\bigr]-1\Bigr).\] The Schmidt coefficients \(\mu_i\) are distributed according to the Dirichlet distribution \(\operatorname{Dir}(1,\dots,1)\) on the simplex [36]. Using standard Dirichlet integrals, \[\mathbb{E}[\mu_i]=\frac{1}{d},\qquad \mathbb{E}[\sqrt{\mu_i}\sqrt{\mu_j}]=\frac{\pi}{4d}\;(i\neq j).\] Therefore \[\mathbb{E}\bigl[(\sum_i\sqrt{\mu_i})^2\bigr] = d\cdot\frac{1}{d} + d(d-1)\cdot\frac{\pi}{4d} = 1 + \frac{\pi}{4}(d-1),\] and the bound follows. \(\sqcup\)
The quantum Wasserstein distance \(W_1\) equips the state space with a canonical Lipschitz structure that can be tailored to a given set of observables. Defining an entanglement measure as the distance to the separable cone embeds entanglement theory into the mature field of optimal transport and metric geometry. This viewpoint is genuinely new: all previously known entanglement measures are based on divergences, fidelity, or norms that do not carry a built‑in Lipschitz dual. In contrast, the Lipschitz constraint in 5 naturally singles out the non‑local components of an observable, thereby providing a geometric separation between classical correlations (captured by the separable states) and genuine quantum correlations.
The dual formulation \({\cal E}(\rho)=\mathop{\rm max}_{\|H\|_{\mathrm{Lip}}\le1} (\mathop{\rm Tr}[H\rho]-\alpha(H))\) should be compared to the usual witness‑based dual \(\widetilde{C}(\rho)=\mathop{\rm max}_{L\ge0,\mathop{\rm Tr}L=1}(-\mathop{\rm Tr}[W\rho])\). Whereas the witness dual involves a positivity constraint and a specific linear witness, our Lipschitz dual involves a commutator‑type smoothness condition. This difference has two important consequences:
The Lipschitz ball is the unit ball of a norm that is often easier to characterise in finite dimensions (norm equivalence, explicit constants).
The dual is manifestly stable under local unitaries and LOCC operations directly from the metric data processing inequality, avoiding the need for optimisation over all possible witnesses.
From an operational viewpoint, Theorems 5 and 3 are particularly significant. They demonstrate that \(E\) is not only a mathematically natural geometric measure but also experimentally testable: any laboratory implementation of a Lipschitz-bounded entanglement witness immediately translates a negative expectation value into a guaranteed lower bound on the entanglement content. Moreover, the witness-based variational principle reveals that the task of finding the best lower bound on \(E\) is equivalent to searching for an optimal Lipschitz-1 witness, a problem that can be tackled with standard semidefinite programming. This bridges the abstract optimal-transport formalism with practical entanglement quantification and opens the door to numerical estimation of \(E\) for arbitrary states, substantially enhancing the applicability of the measure in realistic quantum information tasks. The subadditivity (Theorem 6) and the correlation function bound (Theorem 8) are direct offsprings of this metric structure and have no natural counterpart in the witness formalism.
The Lipschitz formulation opens bridges to several branches of mathematics
Non‑commutative geometry: The commutator condition \(\|[X,H]\|\le1\) can be interpreted as a quantum analogue of the classical \(1\)‑Lipschitz condition on a manifold; hence \({\cal E}\) is associated with a quantum metric space.
Free probability: The bound for random pure states (Theorem 9) already employs exact Dirichlet moments; one may investigate the large‑\(d\) limit in the framework of free semicircular laws and their Wasserstein transport.
Large deviations: The conjecture \(\lim_{n\to\infty}-\frac{1}{n}\log\inf_{\sigma_n\in\mathcal{SEP}_n} \mathop{\rm Tr}[\rho^{\otimes n}\sigma_n]={\cal E}(\rho)\) (Conjecture 4) remains a challenging open problem. Resolving it would likely require new insights at the interface of optimal transport and quantum hypothesis testing.
Convex geometry: The set \(\{\rho:{\cal E}(\rho)\le t\}\) is a convex neighbourhood of the separable cone; its volume ratio with the state space and its polar body (related to the Lipschitz ball) could lead to a Blaschke–Santaló inequality.
Additionally, we record a further conjecture that point towards deeper connections. A natural question is whether the quantity \(E(\rho)\) controls the asymptotic distinguishability of a tensor power \(\rho^{\otimes n}\) from the set of separable states. This would be a quantum Sanov theorem with rate function \(E\). While a full proof is currently out of reach, we can establish a partial lower bound that uses the geometry of \(W_1\) and mirrors classical large‑deviation heuristics.
Conjecture 4. For any state \(\rho\), \[\lim_{n\to\infty} -\frac{1}{n} \log \inf_{\sigma_n\in\mathcal{SEP}_n} \mathop{\rm Tr}\bigl[(\rho^{\otimes n})\sigma_n\bigr] = {\cal E}(\rho),\] where \(\mathcal{SEP}_n\) denotes the separable states on \(\mathcal{H}_{AB}^{\otimes n}\).
We include this statement as a conjecture because a complete proof would require a non‑trivial extension of the quantum Sanov theorem [32], [37] to the metric \(W_1\). The corresponding upper bound is known, while the lower bound remains an outstanding challenge.
We have introduced the Wasserstein entanglement measure, a novel distance-based quantifier rooted in quantum optimal transport. By leveraging the data-processing inequality of the order-1 quantum Wasserstein distance, the measure simultaneously fulfills faithfulness, convexity, LOCC monotonicity, and continuity in a unified geometric framework. The Lipschitz dual formulation transforms the geometric minimization into a tractable maximization, enabling explicit lower bounds for pure and mixed states. For two-qubit systems we obtain a sharp constant \(1/3\), and for Haar-random pure states the expected entanglement is shown to be at least \(\frac{\pi}{4}(1-1/d)\). The metric structure further yields subadditivity, trace-distance bounds, and limits on local observables. A central contribution of this work is the quantitative link to entanglement witnesses: any witness violation provides a certified lower bound on \(E\), and the dual variational problem is rigorously equivalent to finding an optimal Lipschitz-1 witness. This connection makes the measure experimentally accessible and amenable to semidefinite programming, thereby bridging abstract resource theories with practical entanglement detection. Our results establish the Wasserstein entanglement measure as a versatile and physically meaningful tool at the interface of quantum information, optimal transport, and non-commutative geometry. Future challenges include proving the conjectured quantum Sanov theorem with rate \(E\), extending the framework to infinite dimensions and continuous-variable systems, and exploring its implications for quantum complexity and many-body physics. The Lipschitz dual method introduced here is expected to become a standard instrument in entanglement theory and beyond.