The Sharma-Mittal Entropy is Subadditive and Supermodular on the Majorization Lattice

Roberto Bruno and Ugo Vaccaro
Department of Computer Science, University of Salerno
84084 Fisciano (SA), Italy
Email: {rbruno,uvaccaro}@unisa.it


Abstract

We prove that Sharma-Mittal entropy is a subadditive and supermodular function on the lattice of all \(n\)-dimensional probability distributions, ordered according to the partial order relation defined by majorization among vectors. Our result unifies and greatly extends analogous results presented in the literature for the Shannon entropy, the Tsallis entropy, and the Rényi entropy.

1 Introduction↩︎

The mathematical concept of majorization has a rich history, with applications across a wide range of disciplines. Majorization theory arose in mathematical economics [1], where it was employed to rigorously explain the vague notion that the components of a given vector are “more nearly equal’’ than the components of a different vector. Presently, majorization theory finds applications in many areas, ranging from pure mathematics to combinatorics [2], [3], MO?, from information and communication theory [4][7], Mu?, Sa1? to thermodynamics and quantum theory [8][10], NV?, from mathematical chemistry [11] to optimization [12], among the others.

The quantification of uncertainty and diversity in complex systems relies heavily on generalized entropy measures. While the classical Shannon and Boltzmann-Gibbs entropies assume extensivity and independent subsystems, modern applications in non-extensive statistical mechanics and complex systems demand more flexible frameworks Gell?. The Sharma-Mittal entropy [13], a versatile two-parameter generalization, might serve this purpose. By decoupling the degree of non-extensivity from the deformation of the probability distribution, it recovers both the Rényi and Tsallis entropies as limiting cases [14]. Several other entropic functionals can be seen as particular cases of the Sharma-Mittal entropy [15]. Consequently, the Sharma-Mittal entropy has found applications across diverse domains, including holographic dark energy modeling in cosmology Sa?, topic modeling evaluation in natural language processing [16], and the algorithmic quantification of generative design variety [17]. Numerous other applications of the Sharma-Mittal entropy in various fields of Physics are discussed in [18][28], Fi+?, Fra2?, Hu?, Sa?, masi2?, Maz?, Jerin?, and references quoted therein.

Simultaneously, the mathematical framework of majorization has proven useful for formally ordering probability distributions by their relative “disorder” or “concentration.” In particular, the probability simplex under the majorization pre-order forms a lattice—the majorization lattice—equipped with well-defined meet (greatest lower bound) and join (least upper bound) operations [4], [29]. The interplay between entropy measures and the majorization lattice has important physical implications, most notably in Quantum Theory. In the context of bipartite entanglement theory, state transformations under Local Operations and Classical Communication (LOCC), are strictly governed by majorization [9], NV?. Recent literature has extensively utilized the majorization lattice to model approximate bipartite entanglement transformations [30], probabilistic pure state conversions De?, and the detection of high-dimensional quantum steering [31]. Other applications of the majorization lattice to Quantum Theory are discussed in the papers [32], [33], Bo3?, Ma++?.

Despite several studies, the structural properties of generalized entropies on the majorization lattice remain incompletely charted. Cicalese and Vaccaro [4] established that the Shannon entropy is subadditive and supermodular on the majorization lattice. Harremoës Har2? greatly simplified the proofs of the results in [4]. Recently, Yadav and Shkel [34] extended the results of [4] by proving the subadditivity and supermodularity of the Rényi entropy, and Bhatia et al.[29] proved the subadditivity of the Tsallis entropy.

In this paper, we characterize the subadditive, superadditive and supermodular behavior of the Sharma-Mittal entropy \(S_{\alpha,\beta}(\mathbf{p})\) on the majorization lattice. Specifically, we prove that the Sharma-Mittal entropy \(S_{\alpha,\beta}(\mathbf{p})\) is subadditive for \(\alpha\geq 0\) and \(\beta\geq 1\), is superadditive for \(\alpha<0\) and \(\beta\leq 1\), and is supermodular for \(\alpha> 0\) and \(\beta\leq \alpha\). This provides a unified treatment and considerable extension of the results in [4], [29], [34], since the Sharma-Mittal entropy encompasses all the information measures studied therein. We also investigate the regime where \(\beta > \alpha\). We demonstrate a fundamental structural breakdown: by constructing explicit counterexamples (specifically evaluated at \(\alpha=2\), \(\beta=3\)), we prove that the Sharma-Mittal entropy generally loses both supermodularity and submodularity when \(\beta > \alpha\) and \(\alpha>0\).

2 Preliminaries↩︎

Let \({\cal P}_n = \{\mathbf{p}=(p_1,\dots,p_n): p_i\geq 0, \sum_{i=1}^n p_i = 1\}\) be the \((n-1)\)-dimensional probability simplex. Throughout this paper, we assume that all vectors \(\mathbf{p}=(p_1,\dots,p_n)\in {\cal P}_n\) are ordered in non increasing fashion, that is, \(p_1\geq \dots\geq p_n\). We recall the basics of majorization theory MO?.

Definition 1. For any \(\mathbf{p},\mathbf{q}\in {\cal P}_n\), we say that \(\mathbf{p}\) is majorized* by \(\mathbf{q}\) (equivalently, that \(\mathbf{q}\) majorizes \(\mathbf{p}\)), denoted by \(\mathbf{p} \preceq \mathbf{q}\), if it holds that \[\label{p60q} \sum_{i=1}^k p_i \leq \sum_{i=1}^k q_i, \quadfor\;k=1,\ldots,n.\tag{1}\] *

In the more general case, in which \(\mathbf{p}\) and \(\mathbf{q}\) have a different number of components, we pad the shorter vectors with an appropriate number of 0’s and apply the same definition. The majorization relation \(\preceq\) is a partial order relation on \({\cal P}_n,\) that is, \(({\cal P}_n,\preceq)\) is a poset.

Definition 2. A function \(\phi:{\cal P}_{n}\rightarrow \mathbb{R}\) is said to be Schur-convex if for every \(\mathbf{p},\mathbf{q}\) \(\in\) \({\cal P}_{n}\) satisfying \(\mathbf{p}\preceq \mathbf{q}\), it holds that \[\begin{align} \phi(\mathbf{p})\leq \phi(\mathbf{q}) \end{align}\] The function \(\phi\) is Schur-concave if \(-\phi\) is Schur-convex.

We recall [35] that a lattice is a quadruple \(⟨{\cal L},\sqsubseteq,\wedge,\vee⟩\) where \({\cal L}\) is a set, \(\sqsubseteq\) is a partial ordering on \({\cal L}\), and for all \(a,b\in {\cal L}\) there is a unique greatest lower bound (glb) \(a\wedge b\) and a unique least upper bound (lub) \(a\vee b\). More precisely, \(a\wedge b\) and \(a\vee b\) are the elements of \({\cal L}\) that satisfy the conditions \[a\wedge b\sqsubseteq a, a\wedge b\sqsubseteq b, \quad a\sqsubseteq a\vee b, b\sqsubseteq a\vee b,\] and for each \(c\) and \(d\) such that \(c\sqsubseteq a, c\sqsubseteq b, a\sqsubseteq d, b\sqsubseteq d\) one has that \[c\sqsubseteq a \wedge b \quad {and}\quad a\vee b\sqsubseteq d.\]

Bapat bapat? showed that the partially ordered set \((\mathcal{P}_n,\preceq)\) induced by majorization is a lattice. Cicalese and Vaccaro [4] gave explicit constructions for the greatest lower bound (glb) and the least upper bound (lub) of any two elements. We recall their algorithm below. Given \(\mathbf{p}\) and \(\mathbf{q}\) in \(\mathcal{P}_n\), the glb \(\mathbf{p}\wedge \mathbf{q}=\mathbf{r}=(r_1, \ldots , r_n)\) is given by \[\begin{align} \sum^{k}_{i=1} r_i = \min\left\{\sum_{i=1}^{k}p_i,\sum_{i=1}^{k}q_i \right\}.\label{eq:glb} \end{align}\tag{2}\] Define a vector \(\beta(\mathbf{p},\mathbf{q})=\mathbf{w}\in \mathbb{R}^{n}\) such that \[\begin{align} \sum^{k}_{i=1} w_i=\max\left\{\sum_{i=1}^{k}p_i,\sum_{i=1}^{k}q_i \right\}. \label{eq:lub} \end{align}\tag{3}\] If \(\mathbf{w}\in\mathcal{P}_n\), then the lub \(\mathbf{p}\vee \mathbf{q}\) is equal to \(\mathbf{w}\). Otherwise, \(\mathbf{p}\vee \mathbf{q}\) is obtained by repeatedly replacing each maximal consecutive block of coordinates of \(\mathbf{w}\) that violates the non-increasing order by its average, until the resulting vector lies in \(\mathcal{P}_n\). Equivalent methods to obtain \(\mathbf{p}\wedge \mathbf{q}\) and \(\mathbf{p}\vee \mathbf{q}\) in terms of Lorents curves are presented in [34], [36], Har2?.

2.1 Entropies↩︎

Le \(\mathbf{p}\in {\cal P}_n\). We assume that all logarithms in this paper are of base 2. The Shannon entropy is defined as \[H(\mathbf{p})=-\sum_{i=1}^np_i\log p_i.\] The Rényi entropy Re? of order \(\alpha \in [0,\infty]\), is defined as \[H_{\alpha}(\mathbf{p})= \frac{1}{1-\alpha} \log\left( \sum_{i=1}^{n} p_i^ {\alpha}\right).\] For \(\alpha \rightarrow 1\), the Rényi entropy reduces to the Shannon entropy. The Rényi entropy is Schur‑concave.

For \(\alpha\in [0,\infty)\), the Tsallis entropy Ts? is defined as \[T_\alpha(\mathbf{p})=\frac{1-\sum_{i=1}^np_i^\alpha}{\alpha-1}.\] For \(\alpha \rightarrow 1\), also the Tsallis entropy reduces to the Shannon entropy.

For \(\alpha, \beta\in [0,\infty)\), the Sharma-Mittal entropy [13], [14] is defined as \[\label{eq:Sharma-Mittal} S_{\alpha,\beta}(\mathbf{p})=\frac{1}{1-\beta}\left[\left (\sum_{i=1}^np_i^\alpha\right )^{\frac{1-\beta}{1-\alpha}}-1\right ].\tag{4}\] One can see that \(\lim_{\beta\to 1}S_{\alpha,\beta}(\mathbf{p})=\ln(2)H_\alpha(\mathbf{p})\) and for \(\alpha\neq 1\), it holds that \(\lim_{\beta\to \alpha}S_{\alpha,\beta}(\mathbf{p})=T_\alpha(\mathbf{p})\). From the definition of \(H_\alpha(\mathbf{p})\), one can see that \(\sum_{i=1}^n p_i^\alpha=2^{(1-\alpha)H_\alpha(\mathbf{p})}\). Therefore, one can rewrite the Sharma-Mittal entropy as follows \[\label{eq:sharma-mittal-renyi} S_{\alpha,\beta}(\mathbf{p}) = \frac{1}{1-\beta} \left[ \left( 2^{(1-\alpha)H_\alpha(\mathbf{p})} \right)^{\frac{1-\beta}{1-\alpha}} - 1 \right] = \frac{2^{(1-\beta)H_\alpha(\mathbf{p})} - 1}{1-\beta}.\tag{5}\] Equivalently, if one defines the function \(\phi_\beta(x)\) as \[\label{phib} \phi_\beta(x)=\begin{cases} \frac{2^{(1-\beta)x}-1}{1-\beta}, &for \beta\neq 1,\\ \ln(2)x, & for\beta=1, \end{cases}\tag{6}\] then one has \[\label{S61R} S_{\alpha,\beta}(\mathbf{p}) =\phi_\beta(H_\alpha(\mathbf{p})).\tag{7}\]

3 Subadditivity of the Sharma-Mittal entropy on the majorization lattice↩︎

We first recall the notion of subadditive (resp. superadditive) functions on an abstract lattice.

Definition 3. A real-valued function \(\phi:\mathcal{P} \rightarrow \mathbb{R}\) defined on a lattice \((\mathcal{P},\preceq,\wedge,\vee)\) is subadditive (resp. superadditive) if \(\forall\) \(\mathbf{x}\), \(\mathbf{y}\) \(\in\) \(\mathcal{P}\), it holds that \[\begin{align} \phi(\mathbf{x}\wedge \mathbf{y}) \leq(\text{resp. } \geq) \;\phi(\mathbf{x})+\phi(\mathbf{y}). \end{align}\]

The authors of [4] proved that the Shannon entropy is subadditive on the majorization lattice \(({\cal P}_n,\preceq,\wedge,\vee)\). This result was later extended to the Tsallis entropy in [29]. Recently, the authors of [34] proved the analogous result for the Rényi entropy.

In this section, we extend the above results to the more general case of Sharma-Mittal entropies \(S_{\alpha,\beta}\). To this end, we first establish some preliminary results. In the following, whenever \(\alpha < 0\), we assume that \(p_i > 0\) for all \(i\), since otherwise the Sharma-Mittal entropy \(S_{\alpha,\beta}(\mathbf{p})\) would not be well-defined.

Lemma 1. The Sharma-Mittal entropy \[S_{\alpha,\beta}(\mathbf{p})=\frac{1}{1-\beta}\left[\left (\sum_{i=1}^np_i^\alpha\right )^{\frac{1-\beta}{1-\alpha}}-1\right ]\] is Schur-concave for \(\alpha\geq 0\) and Schur-convex for \(\alpha<0\).

Proof. To determine the Schur-concavity or Schur-convexity of the Sharma-Mittal entropy \(S_{\alpha,\beta}(\mathbf{p})\), we apply the Schur-Ostrowski criterion MO?: a function \(\phi:D\subset\mathbb{R}^n\to \mathbb{R}\) is Schur-concave (resp. Schur-convex) if and only if \(\phi\) is symmetric (that is, it is invariant under any permutation of its arguments), continuously differentiable, and for all \(i\neq j\) and \(z=(z_1,\dots,z_n)\in D\) it holds that \[(z_i - z_j)\left( \frac{\partial \phi}{\partial p_i} - \frac{\partial \phi}{\partial p_j} \right) \leq 0 \quad (\text{resp. } \geq 0).\] Thus, because \(S_{\alpha,\beta}(\mathbf{p})\) is symmetric and continuously differentiable, it follows that it is Schur-concave (resp. Schur-convex) if and only if for all \(i \neq j\): \[\label{eq:nec95suff95cond} (p_i - p_j)\left( \frac{\partial S_{\alpha,\beta}}{\partial p_i} - \frac{\partial S_{\alpha,\beta}}{\partial p_j} \right) \leq 0 \quad (\text{resp. } \geq 0).\tag{8}\] To this end, let us compute the partial derivative of \(S_{\alpha,\beta}(\mathbf{p})\) with respect to a generic component \(p_i\) of \(\mathbf{p}\). For the sake of notation, let \(A=\sum_{k=1}^n p_k^\alpha\). It follows: \[\frac{\partial S_{\alpha,\beta}}{\partial p_i} = \frac{1}{1-\beta} \left[ \frac{1-\beta}{1-\alpha} A^{\frac{1-\beta}{1-\alpha} - 1} \cdot \alpha p_i^{\alpha-1} \right] = \frac{\alpha}{1-\alpha} A^{\frac{\alpha-\beta}{1-\alpha}} p_i^{\alpha-1}.\] Thus, we need to study the behaviour of the following expression: \[\label{eq:diff} (p_i-p_j)\left(\frac{\partial S_{\alpha,\beta}}{\partial p_i} - \frac{\partial S_{\alpha,\beta}}{\partial p_j}\right) =(p_i-p_j) \frac{\alpha}{1-\alpha} A^{\frac{\alpha-\beta}{1-\alpha}} \left( p_i^{\alpha-1} - p_j^{\alpha-1} \right).\tag{9}\] Since the term \(A^{\frac{\alpha-\beta}{1-\alpha}}\) is always strictly positive, it does not affect the sign of the expression 9 . Thus, we just need to study the sign of the function \[g(\alpha)=(p_i-p_j) \frac{\alpha}{1-\alpha} \left( p_i^{\alpha-1} - p_j^{\alpha-1} \right).\] We now examine how \(g(\alpha)\) behaves with respect to \(\alpha\). We assume \(p_i\neq p_j\) to avoid trivialities.

Case: \(\alpha\geq 0\). We consider two subcases: \(0<\alpha<1\) and \(\alpha>1\). For the boundary cases \(\alpha=0\) and \(\alpha=1\), we observe that \(S_{\alpha,\beta}\) is Schur-concave. In fact, for \(\alpha=0\), it holds trivially since \(g(0)=0\), while for \(\alpha=1\), it holds because \(\lim_{\alpha\to 1} g(\alpha)=-(p_i-p_j)(\ln p_i-\ln p_j)\leq0\).

Let \(0<\alpha<1\). In this case, since the function \(x\to x^{\alpha-1}\) is strictly decreasing, the factors \[(p_i-p_j)\quad\text{and}\quad (p_i^{\alpha-1} - p_j^{\alpha-1})\] always have opposite signs. Hence, \[(p_i-p_j)(p_i^{\alpha-1} - p_j^{\alpha-1})< 0.\] Therefore, since the term \(\frac{\alpha}{1-\alpha}> 0\), it follows that \(g(\alpha)< 0\).

Let \(\alpha>1\). Since the function \(x\to x^{\alpha-1}\) is strictly increasing, the factors \[(p_i-p_j)\quad\text{and}\quad (p_i^{\alpha-1} - p_j^{\alpha-1})\] always have the same sign. Hence, \[(p_i-p_j)(p_i^{\alpha-1} - p_j^{\alpha-1})> 0.\] From this, since the term \(\frac{\alpha}{1-\alpha}<0\), it follows that \(g(\alpha)< 0\).

Case: \(\alpha<0\). Since the function \(x\to x^{\alpha-1}\) is strictly decreasing, the factors \[(p_i-p_j)\quad\text{and}\quad (p_i^{\alpha-1} - p_j^{\alpha-1})\] always have opposite signs. Thus, since \[(p_i-p_j)(p_i^{\alpha-1} - p_j^{\alpha-1})< 0,\] and the term \(\frac{\alpha}{1-\alpha}< 0\), it follows that \(g(\alpha)>0\). ◻

The following technical Lemma is needed in the proof of Theorem 1, that represents the main result of this section.

Lemma 2. Let \(\mathbf{p},\mathbf{q}\in {\cal P}_n\), and let \(\mathbf{p}\otimes\mathbf{q}\) denote the tensor product of \(\mathbf{p}\) and \(\mathbf{q}\), that is, \[(\mathbf{p}\otimes\mathbf{q})_{i,j} = p_iq_j \quad\forall i,j=1,\dots,n.\] Then, for any \(\alpha,\beta\in\mathbb{R}\), it holds that \[S_{\alpha,\beta}(\mathbf{p}\otimes\mathbf{q})=S_{\alpha,\beta}(\mathbf{p})+S_{\alpha,\beta}(\mathbf{q})+(1-\beta)S_{\alpha,\beta}(\mathbf{p})S_{\alpha,\beta}(\mathbf{q}).\]

Proof. We first observe that as \(\beta\to 1\) and as \((\alpha,\beta)\to(1,1)\), the Sharma-Mittal entropy reduces to the Rényi and Shannon entropies, respectively, for which the equality holds. Thus, we can restrict our attention to \(\alpha,\beta\in\mathbb{R}\setminus\{1\}\). We observe that from 4 , we have that \[1+(1-\beta)S_{\alpha,\beta}(\mathbf{p})=\left(\sum_{i=1}^n p_i^\alpha\right)^{\frac{1-\beta}{1-\alpha}}.\] Thus, it follows that \[\begin{align} 1+(1-\beta)S_{\alpha,\beta}(\mathbf{p}\otimes\mathbf{q})&=\left(\sum_{i,j} (p_iq_j)^\alpha\right)^{\frac{1-\beta}{1-\alpha}}\nonumber\\ &=\left(\sum_{i,j} p_i^\alpha q_j^\alpha\right)^{\frac{1-\beta}{1-\alpha}}\nonumber\\ &=\left(\left(\sum_{i} p_i^\alpha\right) \left(\sum_{j}q_j^\alpha\right)\right)^{\frac{1-\beta}{1-\alpha}}\nonumber\\ &=\left(\sum_{i} p_i^\alpha\right)^{\frac{1-\beta}{1-\alpha}}\left(\sum_{j}q_j^\alpha\right)^{\frac{1-\beta}{1-\alpha}}\nonumber\\ &=\left(1+(1-\beta)S_{\alpha,\beta}(\mathbf{p})\right)\left(1+(1-\beta)S_{\alpha,\beta}(\mathbf{q})\right)\nonumber\\ &=1+(1-\beta)S_{\alpha,\beta}(\mathbf{p})+(1-\beta)S_{\alpha,\beta}(\mathbf{q})+(1-\beta)^2S_{\alpha,\beta}(\mathbf{p})S_{\alpha,\beta}(\mathbf{q}).\label{eq:last95step} \end{align}\tag{10}\] Rearranging 10 to isolate \(S_{\alpha,\beta}(\mathbf{p}\otimes\mathbf{q})\) and dividing by \((1-\beta)\), we obtain \[S_{\alpha,\beta}(\mathbf{p}\otimes\mathbf{q})=S_{\alpha,\beta}(\mathbf{p})+S_{\alpha,\beta}(\mathbf{q})+(1-\beta)S_{\alpha,\beta}(\mathbf{p})S_{\alpha,\beta}(\mathbf{q}),\] which concludes the proof. ◻

We now present the main result of the section.

Theorem 1. Let \(\beta\in \mathbb{R}\), and let \(\mathbf{p},\mathbf{q}\in {\cal P}_n\). Then, for \(\alpha\geq 0\), it holds that \[S_{\alpha,\beta}(\mathbf{p}\wedge\mathbf{q}) \leq S_{\alpha,\beta}(\mathbf{p})+S_{\alpha,\beta}(\mathbf{q})+(1-\beta)S_{\alpha,\beta}(\mathbf{p})S_{\alpha,\beta}(\mathbf{q}),\] and for \(\alpha<0\), it holds that \[S_{\alpha,\beta}(\mathbf{p}\wedge\mathbf{q}) \geq S_{\alpha,\beta}(\mathbf{p})+S_{\alpha,\beta}(\mathbf{q})+(1-\beta)S_{\alpha,\beta}(\mathbf{p})S_{\alpha,\beta}(\mathbf{q}).\]

Proof. First, we observe that \(\mathbf{p}\) and \(\mathbf{q}\) are aggregations of the tensor product \(\mathbf{p}\otimes\mathbf{q}\), that is, \(p_i=\sum_{j} p_iq_j\) for all \(i\) and \(q_j=\sum_i p_iq_j\) for all \(j\). In the context of majorization, an aggregation of a distribution always majorizes the original one. Consequently, we have \[\mathbf{p}\otimes\mathbf{q}\preceq \mathbf{p}\quad \text{and} \quad\mathbf{p}\otimes\mathbf{q}\preceq\mathbf{q}.\] Moreover, by definition of the greatest lower bound \(\mathbf{p}\wedge\mathbf{q}\), it immediately follows that \[\label{eq:prod95maj} \mathbf{p}\otimes\mathbf{q}\preceq\mathbf{p}\wedge\mathbf{q}.\tag{11}\] Now, from 11 and Lemma 1, we have that for \(\alpha\geq 0\), the Sharma-Mittal entropy \(S_{\alpha,\beta}\) is Schur-concave, yielding \[\label{eq:conc} S_{\alpha,\beta}(\mathbf{p}\wedge\mathbf{q})\leq S_{\alpha,\beta}(\mathbf{p}\otimes\mathbf{q}).\tag{12}\] Conversely, for \(\alpha<0\), it is Schur-convex, yielding \[\label{eq:convex} S_{\alpha,\beta}(\mathbf{p}\wedge\mathbf{q})\geq S_{\alpha,\beta}(\mathbf{p}\otimes\mathbf{q}).\tag{13}\] Finally, applying Lemma 2 to expand \(S_{\alpha,\beta}(\mathbf{p}\otimes\mathbf{q})\) in 12 and 13 yields the desired inequalities, concluding the proof. ◻

As a direct consequence of Theorem 1, we obtain the following subadditivity and superadditivity properties for the Sharma-Mittal entropy, depending on the values of the parameters \(\alpha\) and \(\beta\).

Corollary 1. For every \(\alpha\geq 0\) and \(\beta\geq 1\), the Sharma-Mittal entropy \(S_{\alpha,\beta}\) is subadditive on the majorization lattice, that is, for any \(\mathbf{p},\mathbf{q}\in{\cal P}_n\), it holds that \[\label{eq:rsuba} S_{\alpha,\beta}(\mathbf{p}\wedge\mathbf{q}) \le S_{\alpha,\beta}(\mathbf{p})+S_{\alpha,\beta}(\mathbf{q}).\qquad{(1)}\]

Proof. It follows from Theorem 1 since the Sharma-Mittal entropy is non-negative and for \(\beta\geq1\) the term \((1-\beta)S_{\alpha,\beta}(\mathbf{p})S_{\alpha,\beta}(\mathbf{q})\) is smaller or equal to \(0\). ◻

For \(\beta\to 1\), we recover Theorem 1 of [34], for \(\beta\to \alpha\) we obtain Theorem 3.3 of [29], and for \((\alpha,\beta)\to (1,1)\), we get Theorem 2 of [4].

Although the classical definition of the Sharma-Mittal entropy posits that \(\alpha,\beta\geq 0\), we state the following result for the record.

Corollary 2. For every \(\alpha< 0\) and \(\beta\leq 1\), the Sharma-Mittal entropy \(S_{\alpha,\beta}\) is superadditive on the majorization lattice, that is, for any \(\mathbf{p},\mathbf{q}\in{\cal P}_n\), it holds that \[\label{eq:rsupera} S_{\alpha,\beta}(\mathbf{p}\wedge\mathbf{q}) \geq S_{\alpha,\beta}(\mathbf{p})+S_{\alpha,\beta}(\mathbf{q}).\qquad{(2)}\]

Proof. It follows from Theorem 1 since the Sharma-Mittal entropy is non-negative and for \(\beta\leq1\) the term \((1-\beta)S_{\alpha,\beta}(\mathbf{p})S_{\alpha,\beta}(\mathbf{q})\) is greater or equal to \(0\). ◻

4 Supermodularity of the Sharma-Mittal entropy on the majorization lattice↩︎

We recall the definition of supermodular functions on an abstract lattice.

Definition 4. A real-valued function \(\phi:\mathcal{P} \rightarrow \mathbb{R}\) defined on a lattice \((\mathcal{P},\preceq,\wedge,\vee)\) is supermodular if \(\forall\) \(\mathbf{x}\), \(\mathbf{y}\) \(\in\) \(\mathcal{P}\), it holds that \[\begin{align} \phi(\mathbf{x})+\phi(\mathbf{y}) \leq \phi(\mathbf{x}\wedge \mathbf{y})+\phi(\mathbf{x}\vee \mathbf{y}). \end{align}\]

The paper [4] proved supermodularity for the Shannon entropy, [34] proved supermodularity for Rényi and Tsallis entropy. In this section, we extend these results to the more general case of Sharma-Mittal entropies.

Theorem 2. Let \(\alpha >0\) and \(\beta\le \alpha\). Then the Sharma–Mittal entropy is supermodular on the majorization lattice, that is, for any \(\mathbf{p},\mathbf{q}\in {\cal P}_n\), it holds that \[S_{\alpha,\beta}(\mathbf{p})+S_{\alpha,\beta}(\mathbf{q}) \le S_{\alpha,\beta}(\mathbf{p}\wedge \mathbf{q})+S_{\alpha,\beta}(\mathbf{p}\vee \mathbf{q}).\]

Proof. Let \[\mathbf{w}=\mathbf{p}\wedge \mathbf{q}, \qquad \mathbf{v}=\mathbf{p}\vee \mathbf{q},\] and define the function \(g_\alpha(\mathbf{p}):{\cal P}_n\to \mathbb{R}\) as \[g_\alpha(\mathbf{p})=\sum_{i=1}^n p_i^\alpha.\] Recalling the definition of the Tsallis entropy of order \(\alpha\in(0,\infty)\) \[T_\alpha(\mathbf{p})=\frac{1-\sum_{i=1}^np_i^\alpha}{\alpha-1},\] it follows that \[g_\alpha(\mathbf{p})=(1-\alpha)T_\alpha(\mathbf{p})+1.\] Moreover, define the function \(h_{\alpha,\beta}:\mathbb{R^+}\to \mathbb{R}\) as \[\label{unzidvgp} h_{\alpha,\beta}(x)=\frac{1}{1-\beta}\left(x^{\frac{1-\beta}{1-\alpha}}-1\right).\tag{14}\] We can express the Sharma-Mittal entropy as a function of \(g_\alpha(\mathbf{p})\) as follows: \[\label{rqzcxygw} S_{\alpha,\beta}(\mathbf{p})=h_{\alpha,\beta}(g_\alpha(\mathbf{p})).\tag{15}\]

We now proceed by considering three distinct cases depending on the value of \(\alpha\).

Case: \(0<\alpha<1\). In this case, since \(1-\alpha> 0\), and since the Tsallis entropy is Schur-concave and subadditive on the majorization lattice [29], and supermodular for \(\alpha\in[0,\infty)\) [34], we have that the function \(g_\alpha(\mathbf{p})\) is also Schur-concave, subadditive and supermodular.

Moreover, since \(\alpha\in(0,1)\) and \(\beta\leq \alpha\), the exponent \(\frac{1-\beta}{1-\alpha}\) in ([eq:h95def]) is greater or equal to \(1\), and both the first derivative \(h'_{\alpha,\beta}(x)=\frac{1}{1-\alpha}x^{\frac{1-\beta}{1-\alpha}-1}\) and the second derivative \(h''_{\alpha,\beta}(x)=\frac{\frac{1-\beta}{1-\alpha}-1}{1-\alpha}x^{\frac{1-\beta}{1-\alpha}-2}\) are non negative. Thus, both the function \(h_{\alpha,\beta}(x)\) and its derivative \(h'_{\alpha,\beta}(x)\) are increasing in \(x\).

Let \[a=g_\alpha(\mathbf{p}), \qquad b=g_\alpha(\mathbf{q}), \qquad c=g_\alpha(\mathbf{w}), \qquad d=g_\alpha(\mathbf{v}).\] Because \(g_\alpha\) is Schur-concave and \[\mathbf{w}\preceq \mathbf{p},\; \mathbf{q}\preceq \mathbf{v},\] it follows \[c\ge a, \qquad c\ge b, \qquad a\ge d, \qquad b\ge d.\] Moreover, without loss of generality, assuming \(a\geq b\), yields \[\label{tykuxwgl} c\ge a\ge b\ge d.\tag{16}\] Now, since \(g_\alpha\) is supermodular on the majorization lattice, we have that \[\label{yrazebph} a+b\leq c+d.\tag{17}\] Let us define \[\delta=(c-a)\ge0.\] From ([eq:g95supermod]), it holds that \[d-b\ge -\delta.\] Hence, there exists \(\varepsilon\ge0\) such that \[\label{kfsrtpua} d=b-\delta+\varepsilon.\tag{18}\] Thus, we have that \(c=a+\delta\) and \(d=b-\delta+\varepsilon\).

Define the function \(G(t)\) as \[G(t)=h_{\alpha,\beta}(a+t)+h_{\alpha,\beta}(b-t), \qquad 0\le t\le \delta,\] where \(h_{\alpha,\beta}\) is defined in ([eq:h95def]). First, we observe that the argument \(b-t\) is always non negative for \(t\in[0,\delta]\). This holds because \(g_\alpha\) is subadditive and, therefore, since \(g_\alpha(\mathbf{p}\wedge\mathbf{q})=c\leq a+b=g_\alpha(\mathbf{p})+g_\alpha(\mathbf{q})\), we have \(b-t\geq b-\delta=a+b-c\geq0\) for \(t\in[0,\delta]\).

Let \[G'(t)=h'_{\alpha,\beta}(a+t)-h'_{\alpha,\beta}(b-t)\] be the derivative of \(G(t)\). Since, from the chain of inequalities in ([eq:order]), we have that \[a+t\ge b-t,\] and since \(h'_{\alpha,\beta}\) is increasing, it follows that \[G'(t)\ge0.\] Thus, the function \(G\) is increasing and, evaluating \(G(t)\) at \(t=\delta\) and \(t=0\) we obtain

\[\label{xkyaedgr} G(\delta)=h_{\alpha,\beta}(a+\delta)+h_{\alpha,\beta}(b-\delta)\geq h_{\alpha,\beta}(a)+h_{\alpha,\beta}(b)=G(0).\tag{19}\] Moreover, since \(h_{\alpha,\beta}\) is increasing and \(\varepsilon\geq 0\), we have that \[\label{bmxhlrcd} h_{\alpha,\beta}(b-\delta+\varepsilon)\geq h_{\alpha,\beta}(b-\delta).\tag{20}\] Combining [eq:b43eps] with [eq:G], we finally get \[h_{\alpha,\beta}(c)+h_{\alpha,\beta}(d)=h_{\alpha,\beta}(a+\delta)+h_{\alpha,\beta}(b-\delta+\varepsilon)\geq h_{\alpha,\beta}(a)+h_{\alpha,\beta}(b).\] Recalling the definitions of \(a,b,c\) and \(d\), and [eq:sharma95m95as95tsa] we obtain \[S_{\alpha,\beta}(\mathbf{w})+S_{\alpha,\beta}(\mathbf{v}) \ge S_{\alpha,\beta}(\mathbf{p})+S_{\alpha,\beta}(\mathbf{q}),\] which demonstrates that \(S_{\alpha,\beta}\) is supermodular in the range \(0<\alpha < 1\) and \(\beta\le \alpha\).

Case: \(\alpha>1\). In this case, since \(1-\alpha< 0\), we have that the function \(g_\alpha(\mathbf{p})\) is Schur-convex and submodular on the majorization lattice. Moreover, since \(\alpha>1\) and \(\beta\leq \alpha\), we have that the first derivative \(h'_{\alpha,\beta}(x)=\frac{1}{1-\alpha}x^{\frac{1-\beta}{1-\alpha}-1}\) is negative and the second derivative \(h''_{\alpha,\beta}(x)=\frac{\frac{1-\beta}{1-\alpha}-1}{1-\alpha}x^{\frac{1-\beta}{1-\alpha}-2}\) is positive. Thus, the function \(h(x)\) is convex and decreasing in \(x\), and its derivative \(h'(x)\) is increasing in \(x\). Let \[a=g_\alpha(\mathbf{p}), \qquad b=g_\alpha(\mathbf{q}), \qquad c=g_\alpha(\mathbf{w}), \qquad d=g_\alpha(\mathbf{v}).\] Because \(g_\alpha\) is Schur-convex and \[\mathbf{w}\preceq \mathbf{p},\; \mathbf{q}\preceq \mathbf{v},\] it follows \[c\leq a, \qquad c\leq b, \qquad a\leq d, \qquad b\leq d.\] Moreover, without loss of generality, assuming \(a\geq b\), yields \[\label{eq:order95convex} c\leq b\leq a\leq d.\tag{21}\] Since \(g_\alpha\) is submodular on the majorization lattice, we have that \[\label{eq:g95submodular} a+b\geq c+d.\tag{22}\] Let us define \[\delta=(a-c)\ge0.\] From (22 ), it holds that \[d-b\leq \delta.\] Hence, there exists \(\varepsilon\ge0\) such that \[\label{eq:d95convex} d=b+\delta-\varepsilon.\tag{23}\] Thus, we have that \(c=a-\delta\) and \(d=b+\delta-\varepsilon\). Define the function \(G(t)\) as \[G(t)=h_{\alpha,\beta}(a-t)+h_{\alpha,\beta}(b+t), \qquad 0\le t\le \delta,\] where \(h_{\alpha,\beta}\) is defined in ([eq:h95def]). First, we observe that the argument \(a-t\geq a-\delta=c\) is always non negative for \(t\in[0,\delta]\). Let \[G'(t)=-h'_{\alpha,\beta}(a-t)+h'_{\alpha,\beta}(b+t)\] be the derivative of \(G(t)\). We observe that \(G'(t)\geq 0\) for all \(t\) which satisfy \[h'_{\alpha,\beta}(b+t)\geq h'_{\alpha,\beta}(a-t).\] But since \(h'\) is increasing, this implies that \(G'(t)\geq 0\) for all \(t\) such that \[t\geq \frac{a-b}{2}.\] Thus, the function \(G\) is increasing for \(t\in[\frac{a-b}{2},\delta]\). Moreover, evaluating \(G(t)\) at \(t=a-b\leq a-c=\delta\), we have that \[\label{eq:g40a-b41} G(a-b)=h_{\alpha,\beta}(a-(a-b))+h_{\alpha,\beta}(b+(a-b))=h_{\alpha,\beta}(b)+h_{\alpha,\beta}(a)=G(0).\tag{24}\] Therefore, from 24 , we have that \(G(0)=G(a-b)\), and since the function \(G\) is increasing for \(t\in[\frac{a-b}{2},\delta]\), it follows that \[\label{eq:G95convex} G(\delta)=h_{\alpha,\beta}(a-\delta)+h_{\alpha,\beta}(b+\delta)\geq h_{\alpha,\beta}(a)+h_{\alpha,\beta}(b)=G(0).\tag{25}\] Moreover, since \(h_{\alpha,\beta}\) is decreasing and \(\varepsilon\geq 0\), we have that \[\label{eq:b-eps} h_{\alpha,\beta}(b+\delta-\varepsilon)\geq h_{\alpha,\beta}(b+\delta).\tag{26}\] Combining 26 with 25 , we finally get \[h_{\alpha,\beta}(c)+h_{\alpha,\beta}(d)=h_{\alpha,\beta}(a-\delta)+h_{\alpha,\beta}(b+\delta-\varepsilon)\geq h_{\alpha,\beta}(a)+h_{\alpha,\beta}(b).\] Recalling the definitions of \(a,b,c\) and \(d\), and [eq:sharma95m95as95tsa] we obtain \[S_{\alpha,\beta}(\mathbf{w})+S_{\alpha,\beta}(\mathbf{v}) \ge S_{\alpha,\beta}(\mathbf{p})+S_{\alpha,\beta}(\mathbf{q}),\] which demonstrates that \(S_{\alpha,\beta}\) is supermodular in the range \(\alpha>1\) and \(\beta\le \alpha\).

Case: \(\alpha=1\). We recall that the Sharma-Mittal entropy \(S_{\alpha,\beta}(\mathbf{p})\) can be expressed as a function of the Rényi entropy \(H_\alpha(\mathbf{p})\) as shown in 7 . Thus, as \(\alpha\to1\), the Rényi entropy converges to the Shannon entropy \(H(p)\), yielding the following representation: \[\label{eq:S951} S_{1,\beta}(\mathbf{p}) =\phi_\beta(H(\mathbf{p})),\tag{27}\] where the function \(\phi_\beta(x)\) is defined as: \[\label{eq:phi} \phi_\beta(x)=\begin{cases} \frac{2^{(1-\beta)x}-1}{1-\beta}, &for \beta\neq 1,\\ \ln(2)x, & for\beta=1. \end{cases}\tag{28}\] Given that the Shannon entropy \(H(\mathbf{p})\) is Schur-concave, subadditive and supermodular on the majorization lattice [4], and that, for \(\beta\leq\alpha=1\), the first derivative \(\phi'_\beta(x)\) and second derivative \(\phi''_\beta\) of the function \(\phi_\beta(x)\) are non negative, we can proceed following the same idea of the Case \(\alpha\in(0,1)\). Let \[a=H(\mathbf{p}), \qquad b=H(\mathbf{q}), \qquad c=H(\mathbf{w}), \qquad d=H(\mathbf{v}).\] Because Shannon entropy \(H(\mathbf{p})\) is Schur-concave and \[\mathbf{w}\preceq \mathbf{p},\; \mathbf{q}\preceq \mathbf{v},\] it follows \[c\ge a, \qquad c\ge b, \qquad a\ge d, \qquad b\ge d.\] Moreover, without loss of generality, assuming \(a\geq b\), yields \[\label{exjvpbag} c\ge a\ge b\ge d.\tag{29}\] Since the Shannon entropy satisfies the supermodularity property, we have \[\label{uteqzobv} c+d\ge a+b.\tag{30}\] Let us define \[\delta=(c-a)\ge0.\] From ([eq:2]), it holds that \[d-b\ge -\delta.\] Hence there exists \(\varepsilon\ge0\) such that \[\label{ctqoriev} d=b-\delta+\varepsilon.\tag{31}\] Therefore \[(c,d) = (a+\delta,\;b-\delta+\varepsilon).\] Define the function \(F(t)\) as \[F(t)=\phi_\beta(a+t)+\phi_\beta(b-t), \qquad 0\le t\le \delta.\] where \(\phi_\beta\) is defined in (28 ). We first check that the argument \(b-t\geq b-\delta\) remains non negative, therefore it stays within the domain of \(\phi_\beta.\) Since the Shannon entropy is subadditive on the majorization lattice, we get that \(c=H(\mathbf{p}\wedge \mathbf{q})\leq H(\mathbf{p})+H(\mathbf{q})=a+b.\) Therefore \(c\leq a+b\) ensures \(b-t\geq b-\delta=a+b-c\geq 0\) for \(t\in[0,\delta].\)

Because \(\beta\le\alpha=1\), we have that \(\phi_\beta''(x)=(\ln 2)^2(1-\beta)2^{(1-\beta)x}\ge0\). Therefore, the function \(\phi_\beta\) is convex, and its derivative \(\phi_\beta'\) is increasing. Differentiating \(F(t)\) yields \[F'(t) = \phi_\beta'(a+t)-\phi_\beta'(b-t).\] From the chain of inequalities in ([eq:1]), it follows that \[a+t\ge b-t,\] and since \(\phi_\beta'\) is increasing, we obtain that \[F'(t)\ge0.\] Thus, the function \(F(t)\) is increasing in \(t\). Evaluating \(F(t)\) at \(t=\delta\) and \(t=0\) gives \[\label{iprsvbyw} F(\delta)=\phi_\beta(a+\delta)+\phi_\beta(b-\delta) \ge \phi_\beta(a)+\phi_\beta(b)=F(0).\tag{32}\] Finally, since \(\phi_\beta\) is increasing and \(\varepsilon\ge0\), we obtain the inequality \[\label{xfnpyqlg} \phi_\beta(b-\delta+\varepsilon) \ge \phi_\beta(b-\delta).\tag{33}\] Combining ([ineq]) with ([eq:4]),we get \[\phi_\beta(c)+\phi_\beta(d) = \phi_\beta(a+\delta)+\phi_\beta(b-\delta+\varepsilon) \ge \phi_\beta(a)+\phi_\beta(b).\] Recalling the definitions of \(a,b,c, d\), and (27 ) we obtain \[S_{1,\beta}(\mathbf{w})+S_{1,\beta}(\mathbf{v}) \ge S_{1,\beta}(\mathbf{p})+S_{1,\beta}(\mathbf{q}),\] which demonstrates that \(S_{\alpha,\beta}\) is also supermodular for \(\alpha=1\) and \(\beta\le \alpha\). ◻

5 The behavior of \(S_{\alpha,\beta}\) for \(\beta>\alpha\)↩︎

In this section, we show that the Sharma-Mittal entropy \(S_{\alpha,\beta}(\mathbf{p})\) is generally neither supermodular nor submodular on the majorization lattice for \(\beta > \alpha\), and \(\alpha>0\). To this purpose, we construct two explicit counterexamples by finding one pair of probability distributions that breaks supermodularity and another pair that breaks submodularity. Recall the definition of the Sharma-Mittal entropy: \[S_{\alpha, \beta}(\mathbf{p}) = \frac{1}{1-\beta} \left[ \left( \sum_{i=1}^n p_i^\alpha \right)^{\frac{1-\beta}{1-\alpha}} - 1 \right].\]

For our counterexamples, let us fix \(\alpha = 2\) and \(\beta = 3\). This choice simplifies the entropy formula to: \[S_{2,3}(\mathbf{p}) = \frac{1}{-2} \left[ \left( \sum_{i=1}^n p_i^2 \right)^{\frac{-2}{-1}} - 1 \right] = \frac{1}{2} \left( 1 - \left( \sum_{i=1}^n p_i^2 \right)^2 \right)\]

We also recall that in the majorization lattice, given two probability distributions \(\mathbf{p}\) and \(\mathbf{q}\) ordered in a non-increasing fashion, their least upper bound \(\mathbf{p} \vee \mathbf{q}\) and greatest lower bound \(\mathbf{p} \wedge \mathbf{q}\) are constructed by taking the unique distributions whose cumulative sums correspond, respectively, to the component-wise maximum and minimum of the cumulative sums of \(\mathbf{p}\) and \(\mathbf{q}\) (see 3 and 2 for further details).

We can now present the two counterexamples in the following subsections.

5.1 Counterexample 1: \(S_{\alpha, \beta}(\mathbf{p})\) is not supermodular for \(\beta>\alpha\)↩︎

In order to disprove supermodularity, we must show an instance, i.e., a pair of probability distributions \(\mathbf{p}\) and \(\mathbf{q}\), where the property fails, that is, \[S_{2,3}(\mathbf{p} \vee \mathbf{q}) + S_{2,3}(\mathbf{p} \wedge \mathbf{q}) < S_{2,3}(\mathbf{p}) + S_{2,3}(\mathbf{q}).\]

For this purpose, let \(n=4\) and consider the following distributions: \[\mathbf{p} = (0.5, 0.3, 0.1, 0.1)\quad\text{and }\quad\mathbf{q} = (0.4, 0.4, 0.2, 0.0).\] One can verify that in the majorization lattice their least upper bound and greatest lower bound are, respectively: \[\mathbf{p} \vee \mathbf{q} = (0.5, 0.3, 0.2, 0.0)\quad\text{and }\quad\mathbf{p} \wedge \mathbf{q} = (0.4, 0.4, 0.1, 0.1).\] Let us evaluate their Sharma-Mittal entropy \(S_{\alpha,\beta}\) for \(\alpha=2\) and \(\beta=3>1\): \[\begin{align} S_{2,3}(\mathbf{p})&=\frac{1}{2}-\frac{1}{2}\left(0.5^2+0.3^2+0.1^2+0.1^2\right)^2={0.4352}\\ S_{2,3}(\mathbf{q})&=\frac{1}{2}-\frac{1}{2}\left(0.4^2+0.4^2+0.2^2+0.0^2\right)^2={0.4352}\\ S_{2,3}(\mathbf{p} \vee \mathbf{q})&=\frac{1}{2}-\frac{1}{2}\left(0.5^2+0.3^2+0.2^2+0.0^2\right)^2={0.4278}\\ S_{2,3}(\mathbf{p} \wedge \mathbf{q})&=\frac{1}{2}-\frac{1}{2}\left(0.4^2+0.4^2+0.1^2+0.1^2\right)^2={0.4422} \end{align}\] Then, one can see that for the chosen pair of probability distributions, the following inequality holds \[S_{2,3}(\mathbf{p})+S_{2,3}(\mathbf{q})={0.8704}>{0.8700}=S_{2,3}(\mathbf{p} \vee \mathbf{q})+S_{2,3}(\mathbf{p} \wedge \mathbf{q}),\] which shows that \(S_{2,3}\) is not supermodular.

5.2 Counterexample 2: \(S_{\alpha, \beta}(\mathbf{p})\) is not submodular for \(\beta>\alpha\)↩︎

In a similar fashion to the previous section, to disprove submodularity, we must exhibit an instance where the submodular property fails, that is, \[S_{2,3}(\mathbf{p} \vee \mathbf{q}) + S_{2,3}(\mathbf{p} \wedge \mathbf{q}) > S_{2,3}(\mathbf{p}) + S_{2,3}(\mathbf{q}).\]

Let \(n=4\) and consider the following pair of distributions: \[\boldsymbol{p}=(0.5,0.2,0.2,0.1)\quad\text{and }\quad\boldsymbol{q}=(0.4,0.4,0.15,0.05).\] Their respective least upper bound and greatest lower bound are: \[\mathbf{p} \vee \mathbf{q}=(0.5,0.3,0.15,0.05)\quad\text{and }\quad\mathbf{p} \wedge \mathbf{q}=(0.4,0.3,0.2,0.1).\] Evaluating their Sharma-Mittal entropy \(S_{\alpha,\beta}\) for \(\alpha=2\) and \(\beta=3\) yields: \[\begin{align} S_{2,3}(\mathbf{p})&=\frac{1}{2}-\frac{1}{2}\left(0.5^2+0.2^2+0.2^2+0.1^2\right)^2={0.4422}\\ S_{2,3}(\mathbf{q})&=\frac{1}{2}-\frac{1}{2}\left(0.4^2+0.4^2+0.15^2+0.05^2\right)^2={0.4404875}\\ S_{2,3}(\mathbf{p} \vee \mathbf{q})&=\frac{1}{2}-\frac{1}{2}\left(0.5^2+0.3^2+0.15^2+0.05^2\right)^2={0.4333875}\\ S_{2,3}(\mathbf{p} \wedge \mathbf{q})&=\frac{1}{2}-\frac{1}{2}\left(0.4^2+0.3^2+0.2^2+0.1^2\right)^2={0.455} \end{align}\] Consequently, for this choice of probability distributions, the following inequality holds \[S_{2,3}(\mathbf{p} \vee \mathbf{q})+S_{2,3}(\mathbf{p} \wedge \mathbf{q})=0.8883875>0.8826875= S_{2,3}(\mathbf{p})+ S_{2,3}(\mathbf{q}),\] which demonstrates that \(S_{2,3}\) is not submodular.

Taken together, these two counterexamples demonstrate that for \(\beta > \alpha\), and \(\alpha>0\), the Sharma-Mittal entropy is neither supermodular nor submodular on the majorization lattice.

References↩︎

[1]
B. C. Arnold and J. M. Sarabia, Majorization and the Lorenz order with applications in applied mathematics and economics, Springer, vol. 7, 2018.
[2]
G. H. Hardy, J. E. Littlewood, G. Pólya, Inequalities, Cambridge University Press, 1934.
[3]
M. Madiman, L. Wang, and J. O. Woo, Majorization and Rényi entropy inequalities via Sperner theory, Discrete Mathematics, vol. 342(10), pp. 2911–2932, 2019.
[4]
F. Cicalese and U. Vaccaro, Supermodularity and subadditivity properties of the entropy on the majorization lattice, IEEE Transactions in Information Theory, Vol. 48, Issue 4, pp. 933–938, 2002.
[5]
F. Cicalese, L. Gargano, and U. Vaccaro, Minimum-entropy couplings and their applications, IEEE Transactions on Information Theory, vol. 65(6), pp. 3436–-3451, 2019.
[6]
M. A. Khan, S. I. Bradanovic, N. Latif, . Pecaric, and J. Pecaric, Majorization Inequality and Information Theory, Element, Zagreb, 2019.
[7]
H. Witsenhausen, Some aspects of convexity useful in information theory, IEEE Transaction on Information Theory, vol. 26(3), pp. 265–271, 1980.
[8]
G. Bellomo, G. M. Bosyk, Majorization, across the (quantum) universe, Cambridge University Press, 2019.
[9]
M.A. Nielsen, Conditions for a class of entanglement transformations, Phys. Rev. Lett. 83, 436–439, 1999.
[10]
T. Sagawa, Entropy, divergence, and majorization in classical and quantum thermodynamics, Springer Nature, 2022.
[11]
M. Bianchi, G. P. Clemente, A. Cornaro, J. L. Palacios, and A. Torreiro, New trends in majorization techniques for bounding topological indices, Bounds in Chemical Graph Theory-Basics, pp. 3–66, 2017.
[12]
G. Dahl, Principal majorization ideals and optimization, Linear Algebra and its Applications, vol. 331(1–3), pp. 113–130, 2001.
[13]
B. D. Sharma and D. P. Mittal, New non-additive measures of entropy for discrete probability distributions, Journal of Mathematical Sciences, vol. 10(75), pp. 28–40, 1975.
[14]
M. Masi, A step beyond Tsallis and Rényi entropies, Physics Letters A, vol. 338(3–5), pp. 217–224, 2005.
[15]
I. J. Taneja, On generalized information measures and their applications, Advances in Electronics and Electron Physics, vol. 76, pp. 327–413, 1989.
[16]
S. Koltcov, V. Ignatenko, and O. Koltsova, Estimating topic modeling performance with Sharma–Mittal entropy, Entropy, MDPI, vol. 21(7), pp. 660, 2019.
[17]
F. Ahmed, S. K. Ramachandran, M. Fuge, S. Hunter, and S. Miller, Design variety measurement using Sharma–Mittal entropy, Journal of Mechanical Design, American Society of Mechanical Engineers, vol. 143(6), 2020.
[18]
C. Beck, Generalised information and entropy measures in physics, Contemporary Physics, vol. 50(4), pp. 495-–510, 2009.
[19]
M. Carannante and A. Mazzoccoli, Wavelet energy entropy for predictability and cross-market similarity in crude oil benchmarks, Axioms, MDPI, vol. 15(4), pp. 253, 2026.
[20]
T. D. Frank and A. Daffertshofer, Exact time-dependent solutions of the Renyi Fokker–-Planck equation and the Fokker-–Planck equations related to the entropies proposed by Sharma and Mittal, Physica A: Statistical Mechanics and its Applications, vol. 285(3–4), pp. 351–366, 2000.
[21]
S. N. Gashti, A. Anand, M. A. S. Afshar, M. R. Alipour, Y. Sekhmani, B. Pourhassan, İ. Sakallı, and J. Sadeghi, AdS-Schwarzschild-like black hole thermodynamics: Loop quantum gravity impact on topology and universality, Nuclear Physics B, vol. 1021, pp. 117188, 2025.
[22]
S. Ghaffari, A. H. Ziaie, H. Moradpour, F. Asghariyan, F. Feleppa, and M. Tavayef, Black hole thermodynamics in Sharma–-Mittal generalized entropy formalism, General Relativity and Gravitation, vol. 51(7), pp. 93, 2019.
[23]
L. Kaur, B. Singh Khehra and A. Singh, Sharma-Mittal entropy and whale optimization algorithm based multilevel thresholding approach for image segmentation, Artificial Intelligence and Sustainable Computing: Proceedings of ICSISCET 2021, pp. 451–467, 2022.
[24]
R. P. Mondaini and S. C. A. Neto, A comprehensive review of Sharma-Mittal entropy measures and their usefulness in the study of discrete probability distributions in mathematical biology, Trends in Biomathematics: Exploring Epidemics, Eco-Epidemiological Systems, and Optimal Control Strategies: Selected Works from the BIOMAT Consortium Lecture, pp. 321–355, 2024.
[25]
F. Nielsen and R. Nock, A closed-form expression for the Sharma–Mittal entropy of exponential families, Journal of Physics A: Mathematical and Theoretical, vol. 45(3), pp. 32003, 2012.
[26]
R. Rudamenko, D. Abulkhanov, K. Semenov, M. Diskin, and A. Savchenko, Learning when to be sparse: adaptive activations via two–parameter entropy, Workshop on Scientific Methods for Understanding Deep Learning, 2026.
[27]
A. M. Scarfone, Legendre structure of the thermostatistics theory based on the Sharma–-Taneja-–Mittal entropy, Physica A: Statistical Mechanics and its Applications, vol. 365(1), pp. 63–70, 2006.
[28]
D.-P. Xuan, Z.-X. Wang, and S.-M. Fei, Quantum speed limits based on the Sharma–Mittal entropy, Annalen der Physik, vol. 538(1), pp. e00383, 2026.
[29]
P. K. Bhatia, S. Singh, and V. Kumar, On some properties of Tsallis entropy on majorization lattice, Acta Mathematica Academiae Paedagogicae Nyı́regyháziensis, vol. 31(2), pp. 331–340, 2015.
[30]
G. M. Bosyk, G. Sergioli, H. Freytes, F. Holik, and G. Bellomo, Approximate transformations of bipartite pure-state entanglement from the majorization lattice, Physica A: Statistical Mechanics and its Applications, vol. 473, pp. 403–411, 2017.
[31]
M.-C. Yang and C.-F. Qiao, Witness high-dimensional quantum steering via majorization lattice, npj Quantum Information, vol. 12(55), 2026.
[32]
G.M. Bosyk, H. Freytes, G. Bellomo, and G. Sergioli, The lattice of trumping majorization for 4D probability vectors and 2D catalysts. Nat. Sci Rep 8, 3671 (2018).
[33]
J.-L. Li and C.-F. Qiao, The optimal uncertainty relation, Annalen der Physik, vol. 531(10), pp. 1900143, 2019.
[34]
A. K. Yadav and Y. Y. Shkel, Geometry of Rényi entropy on the majorization lattice, arXiv preprint arXiv:2605.09655, 2026.
[35]
B. A. Davey and H. A. Priestly, Introduction to lattices and order, Cambridge University Press, 2002.
[36]
P. Cuff, T. Cover, G. Kumar, and Lei Zhao, A lattice of gambles, IEEE International Symposium on Information Theory, pp. 1762–1766, 2011.