May 18, 2026
We prove that Sharma-Mittal entropy is a subadditive and supermodular function on the lattice of all \(n\)-dimensional probability distributions, ordered according to the partial order relation defined by majorization among vectors. Our result unifies and greatly extends analogous results presented in the literature for the Shannon entropy, the Tsallis entropy, and the Rényi entropy.
The mathematical concept of majorization has a rich history, with applications across a wide range of disciplines. Majorization theory arose in mathematical economics [1], where it was employed to rigorously explain the vague notion that the components of a given vector are “more nearly equal’’ than the components of a different vector. Presently, majorization theory finds applications in many areas, ranging from pure mathematics to combinatorics [2], [3], MO?, from information and communication theory [4]–[7], Mu?, Sa1? to thermodynamics and quantum theory [8]–[10], NV?, from mathematical chemistry [11] to optimization [12], among the others.
The quantification of uncertainty and diversity in complex systems relies heavily on generalized entropy measures. While the classical Shannon and Boltzmann-Gibbs entropies assume extensivity and independent subsystems, modern applications in non-extensive statistical mechanics and complex systems demand more flexible frameworks Gell?. The Sharma-Mittal entropy [13], a versatile two-parameter generalization, might serve this purpose. By decoupling the degree of non-extensivity from the deformation of the probability distribution, it recovers both the Rényi and Tsallis entropies as limiting cases [14]. Several other entropic functionals can be seen as particular cases of the Sharma-Mittal entropy [15]. Consequently, the Sharma-Mittal entropy has found applications across diverse domains, including holographic dark energy modeling in cosmology Sa?, topic modeling evaluation in natural language processing [16], and the algorithmic quantification of generative design variety [17]. Numerous other applications of the Sharma-Mittal entropy in various fields of Physics are discussed in [18]–[28], Fi+?, Fra2?, Hu?, Sa?, masi2?, Maz?, Jerin?, and references quoted therein.
Simultaneously, the mathematical framework of majorization has proven useful for formally ordering probability distributions by their relative “disorder” or “concentration.” In particular, the probability simplex under the majorization pre-order forms a lattice—the majorization lattice—equipped with well-defined meet (greatest lower bound) and join (least upper bound) operations [4], [29]. The interplay between entropy measures and the majorization lattice has important physical implications, most notably in Quantum Theory. In the context of bipartite entanglement theory, state transformations under Local Operations and Classical Communication (LOCC), are strictly governed by majorization [9], NV?. Recent literature has extensively utilized the majorization lattice to model approximate bipartite entanglement transformations [30], probabilistic pure state conversions De?, and the detection of high-dimensional quantum steering [31]. Other applications of the majorization lattice to Quantum Theory are discussed in the papers [32], [33], Bo3?, Ma++?.
Despite several studies, the structural properties of generalized entropies on the majorization lattice remain incompletely charted. Cicalese and Vaccaro [4] established that the Shannon entropy is subadditive and supermodular on the majorization lattice. Harremoës Har2? greatly simplified the proofs of the results in [4]. Recently, Yadav and Shkel [34] extended the results of [4] by proving the subadditivity and supermodularity of the Rényi entropy, and Bhatia et al.[29] proved the subadditivity of the Tsallis entropy.
In this paper, we characterize the subadditive, superadditive and supermodular behavior of the Sharma-Mittal entropy \(S_{\alpha,\beta}(\mathbf{p})\) on the majorization lattice. Specifically, we prove that the Sharma-Mittal entropy \(S_{\alpha,\beta}(\mathbf{p})\) is subadditive for \(\alpha\geq 0\) and \(\beta\geq 1\), is superadditive for \(\alpha<0\) and \(\beta\leq 1\), and is supermodular for \(\alpha> 0\) and \(\beta\leq \alpha\). This provides a unified treatment and considerable extension of the results in [4], [29], [34], since the Sharma-Mittal entropy encompasses all the information measures studied therein. We also investigate the regime where \(\beta > \alpha\). We demonstrate a fundamental structural breakdown: by constructing explicit counterexamples (specifically evaluated at \(\alpha=2\), \(\beta=3\)), we prove that the Sharma-Mittal entropy generally loses both supermodularity and submodularity when \(\beta > \alpha\) and \(\alpha>0\).
Let \({\cal P}_n = \{\mathbf{p}=(p_1,\dots,p_n): p_i\geq 0, \sum_{i=1}^n p_i = 1\}\) be the \((n-1)\)-dimensional probability simplex. Throughout this paper, we assume that all vectors \(\mathbf{p}=(p_1,\dots,p_n)\in {\cal P}_n\) are ordered in non increasing fashion, that is, \(p_1\geq \dots\geq p_n\). We recall the basics of majorization theory MO?.
Definition 1. For any \(\mathbf{p},\mathbf{q}\in {\cal P}_n\), we say that \(\mathbf{p}\) is majorized* by \(\mathbf{q}\) (equivalently, that \(\mathbf{q}\) majorizes \(\mathbf{p}\)), denoted by \(\mathbf{p} \preceq \mathbf{q}\), if it holds that \[\label{p60q} \sum_{i=1}^k p_i \leq \sum_{i=1}^k q_i, \quadfor\;k=1,\ldots,n.\tag{1}\] *
In the more general case, in which \(\mathbf{p}\) and \(\mathbf{q}\) have a different number of components, we pad the shorter vectors with an appropriate number of 0’s and apply the same definition. The majorization relation \(\preceq\) is a partial order relation on \({\cal P}_n,\) that is, \(({\cal P}_n,\preceq)\) is a poset.
Definition 2. A function \(\phi:{\cal P}_{n}\rightarrow \mathbb{R}\) is said to be Schur-convex if for every \(\mathbf{p},\mathbf{q}\) \(\in\) \({\cal P}_{n}\) satisfying \(\mathbf{p}\preceq \mathbf{q}\), it holds that \[\begin{align} \phi(\mathbf{p})\leq \phi(\mathbf{q}) \end{align}\] The function \(\phi\) is Schur-concave if \(-\phi\) is Schur-convex.
We recall [35] that a lattice is a quadruple \(⟨{\cal L},\sqsubseteq,\wedge,\vee⟩\) where \({\cal L}\) is a set, \(\sqsubseteq\) is a partial ordering on \({\cal L}\), and for all \(a,b\in {\cal L}\) there is a unique greatest lower bound (glb) \(a\wedge b\) and a unique least upper bound (lub) \(a\vee b\). More precisely, \(a\wedge b\) and \(a\vee b\) are the elements of \({\cal L}\) that satisfy the conditions \[a\wedge b\sqsubseteq a, a\wedge b\sqsubseteq b, \quad a\sqsubseteq a\vee b, b\sqsubseteq a\vee b,\] and for each \(c\) and \(d\) such that \(c\sqsubseteq a, c\sqsubseteq b, a\sqsubseteq d, b\sqsubseteq d\) one has that \[c\sqsubseteq a \wedge b \quad {and}\quad a\vee b\sqsubseteq d.\]
Bapat bapat? showed that the partially ordered set \((\mathcal{P}_n,\preceq)\) induced by majorization is a lattice. Cicalese and Vaccaro [4] gave explicit constructions for the greatest lower bound (glb) and the least upper bound (lub) of any two elements. We recall their algorithm below. Given \(\mathbf{p}\) and \(\mathbf{q}\) in \(\mathcal{P}_n\), the glb \(\mathbf{p}\wedge \mathbf{q}=\mathbf{r}=(r_1, \ldots , r_n)\) is given by \[\begin{align} \sum^{k}_{i=1} r_i = \min\left\{\sum_{i=1}^{k}p_i,\sum_{i=1}^{k}q_i \right\}.\label{eq:glb} \end{align}\tag{2}\] Define a vector \(\beta(\mathbf{p},\mathbf{q})=\mathbf{w}\in \mathbb{R}^{n}\) such that \[\begin{align} \sum^{k}_{i=1} w_i=\max\left\{\sum_{i=1}^{k}p_i,\sum_{i=1}^{k}q_i \right\}. \label{eq:lub} \end{align}\tag{3}\] If \(\mathbf{w}\in\mathcal{P}_n\), then the lub \(\mathbf{p}\vee \mathbf{q}\) is equal to \(\mathbf{w}\). Otherwise, \(\mathbf{p}\vee \mathbf{q}\) is obtained by repeatedly replacing each maximal consecutive block of coordinates of \(\mathbf{w}\) that violates the non-increasing order by its average, until the resulting vector lies in \(\mathcal{P}_n\). Equivalent methods to obtain \(\mathbf{p}\wedge \mathbf{q}\) and \(\mathbf{p}\vee \mathbf{q}\) in terms of Lorents curves are presented in [34], [36], Har2?.
Le \(\mathbf{p}\in {\cal P}_n\). We assume that all logarithms in this paper are of base 2. The Shannon entropy is defined as \[H(\mathbf{p})=-\sum_{i=1}^np_i\log p_i.\] The Rényi entropy Re? of order \(\alpha \in [0,\infty]\), is defined as \[H_{\alpha}(\mathbf{p})= \frac{1}{1-\alpha} \log\left( \sum_{i=1}^{n} p_i^ {\alpha}\right).\] For \(\alpha \rightarrow 1\), the Rényi entropy reduces to the Shannon entropy. The Rényi entropy is Schur‑concave.
For \(\alpha\in [0,\infty)\), the Tsallis entropy Ts? is defined as \[T_\alpha(\mathbf{p})=\frac{1-\sum_{i=1}^np_i^\alpha}{\alpha-1}.\] For \(\alpha \rightarrow 1\), also the Tsallis entropy reduces to the Shannon entropy.
For \(\alpha, \beta\in [0,\infty)\), the Sharma-Mittal entropy [13], [14] is defined as \[\label{eq:Sharma-Mittal} S_{\alpha,\beta}(\mathbf{p})=\frac{1}{1-\beta}\left[\left (\sum_{i=1}^np_i^\alpha\right )^{\frac{1-\beta}{1-\alpha}}-1\right ].\tag{4}\] One can see that \(\lim_{\beta\to 1}S_{\alpha,\beta}(\mathbf{p})=\ln(2)H_\alpha(\mathbf{p})\) and for \(\alpha\neq 1\), it holds that \(\lim_{\beta\to \alpha}S_{\alpha,\beta}(\mathbf{p})=T_\alpha(\mathbf{p})\). From the definition of \(H_\alpha(\mathbf{p})\), one can see that \(\sum_{i=1}^n p_i^\alpha=2^{(1-\alpha)H_\alpha(\mathbf{p})}\). Therefore, one can rewrite the Sharma-Mittal entropy as follows \[\label{eq:sharma-mittal-renyi} S_{\alpha,\beta}(\mathbf{p}) = \frac{1}{1-\beta} \left[ \left( 2^{(1-\alpha)H_\alpha(\mathbf{p})} \right)^{\frac{1-\beta}{1-\alpha}} - 1 \right] = \frac{2^{(1-\beta)H_\alpha(\mathbf{p})} - 1}{1-\beta}.\tag{5}\] Equivalently, if one defines the function \(\phi_\beta(x)\) as \[\label{phib} \phi_\beta(x)=\begin{cases} \frac{2^{(1-\beta)x}-1}{1-\beta}, &for \beta\neq 1,\\ \ln(2)x, & for\beta=1, \end{cases}\tag{6}\] then one has \[\label{S61R} S_{\alpha,\beta}(\mathbf{p}) =\phi_\beta(H_\alpha(\mathbf{p})).\tag{7}\]
We first recall the notion of subadditive (resp. superadditive) functions on an abstract lattice.
Definition 3. A real-valued function \(\phi:\mathcal{P} \rightarrow \mathbb{R}\) defined on a lattice \((\mathcal{P},\preceq,\wedge,\vee)\) is subadditive (resp. superadditive) if \(\forall\) \(\mathbf{x}\), \(\mathbf{y}\) \(\in\) \(\mathcal{P}\), it holds that \[\begin{align} \phi(\mathbf{x}\wedge \mathbf{y}) \leq(\text{resp. } \geq) \;\phi(\mathbf{x})+\phi(\mathbf{y}). \end{align}\]
The authors of [4] proved that the Shannon entropy is subadditive on the majorization lattice \(({\cal P}_n,\preceq,\wedge,\vee)\). This result was later extended to the Tsallis entropy in [29]. Recently, the authors of [34] proved the analogous result for the Rényi entropy.
In this section, we extend the above results to the more general case of Sharma-Mittal entropies \(S_{\alpha,\beta}\). To this end, we first establish some preliminary results. In the following, whenever \(\alpha < 0\), we assume that \(p_i > 0\) for all \(i\), since otherwise the Sharma-Mittal entropy \(S_{\alpha,\beta}(\mathbf{p})\) would not be well-defined.
Lemma 1. The Sharma-Mittal entropy \[S_{\alpha,\beta}(\mathbf{p})=\frac{1}{1-\beta}\left[\left (\sum_{i=1}^np_i^\alpha\right )^{\frac{1-\beta}{1-\alpha}}-1\right ]\] is Schur-concave for \(\alpha\geq 0\) and Schur-convex for \(\alpha<0\).
Proof. To determine the Schur-concavity or Schur-convexity of the Sharma-Mittal entropy \(S_{\alpha,\beta}(\mathbf{p})\), we apply the Schur-Ostrowski criterion MO?: a function \(\phi:D\subset\mathbb{R}^n\to \mathbb{R}\) is Schur-concave (resp. Schur-convex) if and only if \(\phi\) is symmetric (that is, it is invariant under any permutation of its arguments), continuously differentiable, and for all \(i\neq j\) and \(z=(z_1,\dots,z_n)\in D\) it holds that \[(z_i - z_j)\left( \frac{\partial \phi}{\partial p_i} - \frac{\partial \phi}{\partial p_j} \right) \leq 0 \quad (\text{resp. } \geq 0).\] Thus, because \(S_{\alpha,\beta}(\mathbf{p})\) is symmetric and continuously differentiable, it follows that it is Schur-concave (resp. Schur-convex) if and only if for all \(i \neq j\): \[\label{eq:nec95suff95cond} (p_i - p_j)\left( \frac{\partial S_{\alpha,\beta}}{\partial p_i} - \frac{\partial S_{\alpha,\beta}}{\partial p_j} \right) \leq 0 \quad (\text{resp. } \geq 0).\tag{8}\] To this end, let us compute the partial derivative of \(S_{\alpha,\beta}(\mathbf{p})\) with respect to a generic component \(p_i\) of \(\mathbf{p}\). For the sake of notation, let \(A=\sum_{k=1}^n p_k^\alpha\). It follows: \[\frac{\partial S_{\alpha,\beta}}{\partial p_i} = \frac{1}{1-\beta} \left[ \frac{1-\beta}{1-\alpha} A^{\frac{1-\beta}{1-\alpha} - 1} \cdot \alpha p_i^{\alpha-1} \right] = \frac{\alpha}{1-\alpha} A^{\frac{\alpha-\beta}{1-\alpha}} p_i^{\alpha-1}.\] Thus, we need to study the behaviour of the following expression: \[\label{eq:diff} (p_i-p_j)\left(\frac{\partial S_{\alpha,\beta}}{\partial p_i} - \frac{\partial S_{\alpha,\beta}}{\partial p_j}\right) =(p_i-p_j) \frac{\alpha}{1-\alpha} A^{\frac{\alpha-\beta}{1-\alpha}} \left( p_i^{\alpha-1} - p_j^{\alpha-1} \right).\tag{9}\] Since the term \(A^{\frac{\alpha-\beta}{1-\alpha}}\) is always strictly positive, it does not affect the sign of the expression 9 . Thus, we just need to study the sign of the function \[g(\alpha)=(p_i-p_j) \frac{\alpha}{1-\alpha} \left( p_i^{\alpha-1} - p_j^{\alpha-1} \right).\] We now examine how \(g(\alpha)\) behaves with respect to \(\alpha\). We assume \(p_i\neq p_j\) to avoid trivialities.
Case: \(\alpha\geq 0\). We consider two subcases: \(0<\alpha<1\) and \(\alpha>1\). For the boundary cases \(\alpha=0\) and \(\alpha=1\), we observe that \(S_{\alpha,\beta}\) is Schur-concave. In fact, for \(\alpha=0\), it holds trivially since \(g(0)=0\), while for \(\alpha=1\), it holds because \(\lim_{\alpha\to 1} g(\alpha)=-(p_i-p_j)(\ln p_i-\ln p_j)\leq0\).
Let \(0<\alpha<1\). In this case, since the function \(x\to x^{\alpha-1}\) is strictly decreasing, the factors \[(p_i-p_j)\quad\text{and}\quad (p_i^{\alpha-1} - p_j^{\alpha-1})\] always have opposite signs. Hence, \[(p_i-p_j)(p_i^{\alpha-1} - p_j^{\alpha-1})< 0.\] Therefore, since the term \(\frac{\alpha}{1-\alpha}> 0\), it follows that \(g(\alpha)< 0\).
Let \(\alpha>1\). Since the function \(x\to x^{\alpha-1}\) is strictly increasing, the factors \[(p_i-p_j)\quad\text{and}\quad (p_i^{\alpha-1} - p_j^{\alpha-1})\] always have the same sign. Hence, \[(p_i-p_j)(p_i^{\alpha-1} - p_j^{\alpha-1})> 0.\] From this, since the term \(\frac{\alpha}{1-\alpha}<0\), it follows that \(g(\alpha)< 0\).
Case: \(\alpha<0\). Since the function \(x\to x^{\alpha-1}\) is strictly decreasing, the factors \[(p_i-p_j)\quad\text{and}\quad (p_i^{\alpha-1} - p_j^{\alpha-1})\] always have opposite signs. Thus, since \[(p_i-p_j)(p_i^{\alpha-1} - p_j^{\alpha-1})< 0,\] and the term \(\frac{\alpha}{1-\alpha}< 0\), it follows that \(g(\alpha)>0\). ◻
The following technical Lemma is needed in the proof of Theorem 1, that represents the main result of this section.
Lemma 2. Let \(\mathbf{p},\mathbf{q}\in {\cal P}_n\), and let \(\mathbf{p}\otimes\mathbf{q}\) denote the tensor product of \(\mathbf{p}\) and \(\mathbf{q}\), that is, \[(\mathbf{p}\otimes\mathbf{q})_{i,j} = p_iq_j \quad\forall i,j=1,\dots,n.\] Then, for any \(\alpha,\beta\in\mathbb{R}\), it holds that \[S_{\alpha,\beta}(\mathbf{p}\otimes\mathbf{q})=S_{\alpha,\beta}(\mathbf{p})+S_{\alpha,\beta}(\mathbf{q})+(1-\beta)S_{\alpha,\beta}(\mathbf{p})S_{\alpha,\beta}(\mathbf{q}).\]
Proof. We first observe that as \(\beta\to 1\) and as \((\alpha,\beta)\to(1,1)\), the Sharma-Mittal entropy reduces to the Rényi and Shannon entropies, respectively, for which the equality holds. Thus, we can restrict our attention to \(\alpha,\beta\in\mathbb{R}\setminus\{1\}\). We observe that from 4 , we have that \[1+(1-\beta)S_{\alpha,\beta}(\mathbf{p})=\left(\sum_{i=1}^n p_i^\alpha\right)^{\frac{1-\beta}{1-\alpha}}.\] Thus, it follows that \[\begin{align} 1+(1-\beta)S_{\alpha,\beta}(\mathbf{p}\otimes\mathbf{q})&=\left(\sum_{i,j} (p_iq_j)^\alpha\right)^{\frac{1-\beta}{1-\alpha}}\nonumber\\ &=\left(\sum_{i,j} p_i^\alpha q_j^\alpha\right)^{\frac{1-\beta}{1-\alpha}}\nonumber\\ &=\left(\left(\sum_{i} p_i^\alpha\right) \left(\sum_{j}q_j^\alpha\right)\right)^{\frac{1-\beta}{1-\alpha}}\nonumber\\ &=\left(\sum_{i} p_i^\alpha\right)^{\frac{1-\beta}{1-\alpha}}\left(\sum_{j}q_j^\alpha\right)^{\frac{1-\beta}{1-\alpha}}\nonumber\\ &=\left(1+(1-\beta)S_{\alpha,\beta}(\mathbf{p})\right)\left(1+(1-\beta)S_{\alpha,\beta}(\mathbf{q})\right)\nonumber\\ &=1+(1-\beta)S_{\alpha,\beta}(\mathbf{p})+(1-\beta)S_{\alpha,\beta}(\mathbf{q})+(1-\beta)^2S_{\alpha,\beta}(\mathbf{p})S_{\alpha,\beta}(\mathbf{q}).\label{eq:last95step} \end{align}\tag{10}\] Rearranging 10 to isolate \(S_{\alpha,\beta}(\mathbf{p}\otimes\mathbf{q})\) and dividing by \((1-\beta)\), we obtain \[S_{\alpha,\beta}(\mathbf{p}\otimes\mathbf{q})=S_{\alpha,\beta}(\mathbf{p})+S_{\alpha,\beta}(\mathbf{q})+(1-\beta)S_{\alpha,\beta}(\mathbf{p})S_{\alpha,\beta}(\mathbf{q}),\] which concludes the proof. ◻
We now present the main result of the section.
Theorem 1. Let \(\beta\in \mathbb{R}\), and let \(\mathbf{p},\mathbf{q}\in {\cal P}_n\). Then, for \(\alpha\geq 0\), it holds that \[S_{\alpha,\beta}(\mathbf{p}\wedge\mathbf{q}) \leq S_{\alpha,\beta}(\mathbf{p})+S_{\alpha,\beta}(\mathbf{q})+(1-\beta)S_{\alpha,\beta}(\mathbf{p})S_{\alpha,\beta}(\mathbf{q}),\] and for \(\alpha<0\), it holds that \[S_{\alpha,\beta}(\mathbf{p}\wedge\mathbf{q}) \geq S_{\alpha,\beta}(\mathbf{p})+S_{\alpha,\beta}(\mathbf{q})+(1-\beta)S_{\alpha,\beta}(\mathbf{p})S_{\alpha,\beta}(\mathbf{q}).\]
Proof. First, we observe that \(\mathbf{p}\) and \(\mathbf{q}\) are aggregations of the tensor product \(\mathbf{p}\otimes\mathbf{q}\), that is, \(p_i=\sum_{j} p_iq_j\) for all \(i\) and \(q_j=\sum_i p_iq_j\) for all \(j\). In the context of majorization, an aggregation of a distribution always majorizes the original one. Consequently, we have \[\mathbf{p}\otimes\mathbf{q}\preceq \mathbf{p}\quad \text{and} \quad\mathbf{p}\otimes\mathbf{q}\preceq\mathbf{q}.\] Moreover, by definition of the greatest lower bound \(\mathbf{p}\wedge\mathbf{q}\), it immediately follows that \[\label{eq:prod95maj} \mathbf{p}\otimes\mathbf{q}\preceq\mathbf{p}\wedge\mathbf{q}.\tag{11}\] Now, from 11 and Lemma 1, we have that for \(\alpha\geq 0\), the Sharma-Mittal entropy \(S_{\alpha,\beta}\) is Schur-concave, yielding \[\label{eq:conc} S_{\alpha,\beta}(\mathbf{p}\wedge\mathbf{q})\leq S_{\alpha,\beta}(\mathbf{p}\otimes\mathbf{q}).\tag{12}\] Conversely, for \(\alpha<0\), it is Schur-convex, yielding \[\label{eq:convex} S_{\alpha,\beta}(\mathbf{p}\wedge\mathbf{q})\geq S_{\alpha,\beta}(\mathbf{p}\otimes\mathbf{q}).\tag{13}\] Finally, applying Lemma 2 to expand \(S_{\alpha,\beta}(\mathbf{p}\otimes\mathbf{q})\) in 12 and 13 yields the desired inequalities, concluding the proof. ◻
As a direct consequence of Theorem 1, we obtain the following subadditivity and superadditivity properties for the Sharma-Mittal entropy, depending on the values of the parameters \(\alpha\) and \(\beta\).
Corollary 1. For every \(\alpha\geq 0\) and \(\beta\geq 1\), the Sharma-Mittal entropy \(S_{\alpha,\beta}\) is subadditive on the majorization lattice, that is, for any \(\mathbf{p},\mathbf{q}\in{\cal P}_n\), it holds that \[\label{eq:rsuba} S_{\alpha,\beta}(\mathbf{p}\wedge\mathbf{q}) \le S_{\alpha,\beta}(\mathbf{p})+S_{\alpha,\beta}(\mathbf{q}).\qquad{(1)}\]
Proof. It follows from Theorem 1 since the Sharma-Mittal entropy is non-negative and for \(\beta\geq1\) the term \((1-\beta)S_{\alpha,\beta}(\mathbf{p})S_{\alpha,\beta}(\mathbf{q})\) is smaller or equal to \(0\). ◻
For \(\beta\to 1\), we recover Theorem 1 of [34], for \(\beta\to \alpha\) we obtain Theorem 3.3 of [29], and for \((\alpha,\beta)\to (1,1)\), we get Theorem 2 of [4].
Although the classical definition of the Sharma-Mittal entropy posits that \(\alpha,\beta\geq 0\), we state the following result for the record.
Corollary 2. For every \(\alpha< 0\) and \(\beta\leq 1\), the Sharma-Mittal entropy \(S_{\alpha,\beta}\) is superadditive on the majorization lattice, that is, for any \(\mathbf{p},\mathbf{q}\in{\cal P}_n\), it holds that \[\label{eq:rsupera} S_{\alpha,\beta}(\mathbf{p}\wedge\mathbf{q}) \geq S_{\alpha,\beta}(\mathbf{p})+S_{\alpha,\beta}(\mathbf{q}).\qquad{(2)}\]
Proof. It follows from Theorem 1 since the Sharma-Mittal entropy is non-negative and for \(\beta\leq1\) the term \((1-\beta)S_{\alpha,\beta}(\mathbf{p})S_{\alpha,\beta}(\mathbf{q})\) is greater or equal to \(0\). ◻
We recall the definition of supermodular functions on an abstract lattice.
Definition 4. A real-valued function \(\phi:\mathcal{P} \rightarrow \mathbb{R}\) defined on a lattice \((\mathcal{P},\preceq,\wedge,\vee)\) is supermodular if \(\forall\) \(\mathbf{x}\), \(\mathbf{y}\) \(\in\) \(\mathcal{P}\), it holds that \[\begin{align} \phi(\mathbf{x})+\phi(\mathbf{y}) \leq \phi(\mathbf{x}\wedge \mathbf{y})+\phi(\mathbf{x}\vee \mathbf{y}). \end{align}\]
The paper [4] proved supermodularity for the Shannon entropy, [34] proved supermodularity for Rényi and Tsallis entropy. In this section, we extend these results to the more general case of Sharma-Mittal entropies.
Theorem 2. Let \(\alpha >0\) and \(\beta\le \alpha\). Then the Sharma–Mittal entropy is supermodular on the majorization lattice, that is, for any \(\mathbf{p},\mathbf{q}\in {\cal P}_n\), it holds that \[S_{\alpha,\beta}(\mathbf{p})+S_{\alpha,\beta}(\mathbf{q}) \le S_{\alpha,\beta}(\mathbf{p}\wedge \mathbf{q})+S_{\alpha,\beta}(\mathbf{p}\vee \mathbf{q}).\]
Proof. Let \[\mathbf{w}=\mathbf{p}\wedge \mathbf{q}, \qquad \mathbf{v}=\mathbf{p}\vee \mathbf{q},\] and define the function \(g_\alpha(\mathbf{p}):{\cal P}_n\to \mathbb{R}\) as \[g_\alpha(\mathbf{p})=\sum_{i=1}^n p_i^\alpha.\] Recalling the definition of the Tsallis entropy of order \(\alpha\in(0,\infty)\) \[T_\alpha(\mathbf{p})=\frac{1-\sum_{i=1}^np_i^\alpha}{\alpha-1},\] it follows that \[g_\alpha(\mathbf{p})=(1-\alpha)T_\alpha(\mathbf{p})+1.\] Moreover, define the function \(h_{\alpha,\beta}:\mathbb{R^+}\to \mathbb{R}\) as \[\label{unzidvgp} h_{\alpha,\beta}(x)=\frac{1}{1-\beta}\left(x^{\frac{1-\beta}{1-\alpha}}-1\right).\tag{14}\] We can express the Sharma-Mittal entropy as a function of \(g_\alpha(\mathbf{p})\) as follows: \[\label{rqzcxygw} S_{\alpha,\beta}(\mathbf{p})=h_{\alpha,\beta}(g_\alpha(\mathbf{p})).\tag{15}\]
We now proceed by considering three distinct cases depending on the value of \(\alpha\).
Case: \(0<\alpha<1\). In this case, since \(1-\alpha> 0\), and since the Tsallis entropy is Schur-concave and subadditive on the majorization lattice [29], and supermodular for \(\alpha\in[0,\infty)\) [34], we have that the function \(g_\alpha(\mathbf{p})\) is also Schur-concave, subadditive and supermodular.
Moreover, since \(\alpha\in(0,1)\) and \(\beta\leq \alpha\), the exponent \(\frac{1-\beta}{1-\alpha}\) in ([eq:h95def]) is greater or equal to \(1\), and both the first derivative \(h'_{\alpha,\beta}(x)=\frac{1}{1-\alpha}x^{\frac{1-\beta}{1-\alpha}-1}\) and the second derivative \(h''_{\alpha,\beta}(x)=\frac{\frac{1-\beta}{1-\alpha}-1}{1-\alpha}x^{\frac{1-\beta}{1-\alpha}-2}\) are non negative. Thus, both the function \(h_{\alpha,\beta}(x)\) and its derivative \(h'_{\alpha,\beta}(x)\) are increasing in \(x\).
Let \[a=g_\alpha(\mathbf{p}), \qquad b=g_\alpha(\mathbf{q}), \qquad c=g_\alpha(\mathbf{w}), \qquad d=g_\alpha(\mathbf{v}).\] Because \(g_\alpha\) is Schur-concave and \[\mathbf{w}\preceq \mathbf{p},\; \mathbf{q}\preceq \mathbf{v},\] it follows \[c\ge a, \qquad c\ge b, \qquad a\ge d, \qquad b\ge d.\] Moreover, without loss of generality, assuming \(a\geq b\), yields \[\label{tykuxwgl} c\ge a\ge b\ge d.\tag{16}\] Now, since \(g_\alpha\) is supermodular on the majorization lattice, we have that \[\label{yrazebph} a+b\leq c+d.\tag{17}\] Let us define \[\delta=(c-a)\ge0.\] From ([eq:g95supermod]), it holds that \[d-b\ge -\delta.\] Hence, there exists \(\varepsilon\ge0\) such that \[\label{kfsrtpua} d=b-\delta+\varepsilon.\tag{18}\] Thus, we have that \(c=a+\delta\) and \(d=b-\delta+\varepsilon\).
Define the function \(G(t)\) as \[G(t)=h_{\alpha,\beta}(a+t)+h_{\alpha,\beta}(b-t), \qquad 0\le t\le \delta,\] where \(h_{\alpha,\beta}\) is defined in ([eq:h95def]). First, we observe that the argument \(b-t\) is always non negative for \(t\in[0,\delta]\). This holds because \(g_\alpha\) is subadditive and, therefore, since \(g_\alpha(\mathbf{p}\wedge\mathbf{q})=c\leq a+b=g_\alpha(\mathbf{p})+g_\alpha(\mathbf{q})\), we have \(b-t\geq b-\delta=a+b-c\geq0\) for \(t\in[0,\delta]\).
Let \[G'(t)=h'_{\alpha,\beta}(a+t)-h'_{\alpha,\beta}(b-t)\] be the derivative of \(G(t)\). Since, from the chain of inequalities in ([eq:order]), we have that \[a+t\ge b-t,\] and since \(h'_{\alpha,\beta}\) is increasing, it follows that \[G'(t)\ge0.\] Thus, the function \(G\) is increasing and, evaluating \(G(t)\) at \(t=\delta\) and \(t=0\) we obtain
\[\label{xkyaedgr} G(\delta)=h_{\alpha,\beta}(a+\delta)+h_{\alpha,\beta}(b-\delta)\geq h_{\alpha,\beta}(a)+h_{\alpha,\beta}(b)=G(0).\tag{19}\] Moreover, since \(h_{\alpha,\beta}\) is increasing and \(\varepsilon\geq 0\), we have that \[\label{bmxhlrcd} h_{\alpha,\beta}(b-\delta+\varepsilon)\geq h_{\alpha,\beta}(b-\delta).\tag{20}\] Combining [eq:b43eps] with [eq:G], we finally get \[h_{\alpha,\beta}(c)+h_{\alpha,\beta}(d)=h_{\alpha,\beta}(a+\delta)+h_{\alpha,\beta}(b-\delta+\varepsilon)\geq h_{\alpha,\beta}(a)+h_{\alpha,\beta}(b).\] Recalling the definitions of \(a,b,c\) and \(d\), and [eq:sharma95m95as95tsa] we obtain \[S_{\alpha,\beta}(\mathbf{w})+S_{\alpha,\beta}(\mathbf{v}) \ge S_{\alpha,\beta}(\mathbf{p})+S_{\alpha,\beta}(\mathbf{q}),\] which demonstrates that \(S_{\alpha,\beta}\) is supermodular in the range \(0<\alpha < 1\) and \(\beta\le \alpha\).
Case: \(\alpha>1\). In this case, since \(1-\alpha< 0\), we have that the function \(g_\alpha(\mathbf{p})\) is Schur-convex and submodular on the majorization lattice. Moreover, since \(\alpha>1\) and \(\beta\leq \alpha\), we have that the first derivative \(h'_{\alpha,\beta}(x)=\frac{1}{1-\alpha}x^{\frac{1-\beta}{1-\alpha}-1}\) is negative and the second derivative \(h''_{\alpha,\beta}(x)=\frac{\frac{1-\beta}{1-\alpha}-1}{1-\alpha}x^{\frac{1-\beta}{1-\alpha}-2}\) is positive. Thus, the function \(h(x)\) is convex and decreasing in \(x\), and its derivative \(h'(x)\) is increasing in \(x\). Let \[a=g_\alpha(\mathbf{p}), \qquad b=g_\alpha(\mathbf{q}), \qquad c=g_\alpha(\mathbf{w}), \qquad d=g_\alpha(\mathbf{v}).\] Because \(g_\alpha\) is Schur-convex and \[\mathbf{w}\preceq \mathbf{p},\; \mathbf{q}\preceq \mathbf{v},\] it follows \[c\leq a, \qquad c\leq b, \qquad a\leq d, \qquad b\leq d.\] Moreover, without loss of generality, assuming \(a\geq b\), yields \[\label{eq:order95convex} c\leq b\leq a\leq d.\tag{21}\] Since \(g_\alpha\) is submodular on the majorization lattice, we have that \[\label{eq:g95submodular} a+b\geq c+d.\tag{22}\] Let us define \[\delta=(a-c)\ge0.\] From (22 ), it holds that \[d-b\leq \delta.\] Hence, there exists \(\varepsilon\ge0\) such that \[\label{eq:d95convex} d=b+\delta-\varepsilon.\tag{23}\] Thus, we have that \(c=a-\delta\) and \(d=b+\delta-\varepsilon\). Define the function \(G(t)\) as \[G(t)=h_{\alpha,\beta}(a-t)+h_{\alpha,\beta}(b+t), \qquad 0\le t\le \delta,\] where \(h_{\alpha,\beta}\) is defined in ([eq:h95def]). First, we observe that the argument \(a-t\geq a-\delta=c\) is always non negative for \(t\in[0,\delta]\). Let \[G'(t)=-h'_{\alpha,\beta}(a-t)+h'_{\alpha,\beta}(b+t)\] be the derivative of \(G(t)\). We observe that \(G'(t)\geq 0\) for all \(t\) which satisfy \[h'_{\alpha,\beta}(b+t)\geq h'_{\alpha,\beta}(a-t).\] But since \(h'\) is increasing, this implies that \(G'(t)\geq 0\) for all \(t\) such that \[t\geq \frac{a-b}{2}.\] Thus, the function \(G\) is increasing for \(t\in[\frac{a-b}{2},\delta]\). Moreover, evaluating \(G(t)\) at \(t=a-b\leq a-c=\delta\), we have that \[\label{eq:g40a-b41} G(a-b)=h_{\alpha,\beta}(a-(a-b))+h_{\alpha,\beta}(b+(a-b))=h_{\alpha,\beta}(b)+h_{\alpha,\beta}(a)=G(0).\tag{24}\] Therefore, from 24 , we have that \(G(0)=G(a-b)\), and since the function \(G\) is increasing for \(t\in[\frac{a-b}{2},\delta]\), it follows that \[\label{eq:G95convex} G(\delta)=h_{\alpha,\beta}(a-\delta)+h_{\alpha,\beta}(b+\delta)\geq h_{\alpha,\beta}(a)+h_{\alpha,\beta}(b)=G(0).\tag{25}\] Moreover, since \(h_{\alpha,\beta}\) is decreasing and \(\varepsilon\geq 0\), we have that \[\label{eq:b-eps} h_{\alpha,\beta}(b+\delta-\varepsilon)\geq h_{\alpha,\beta}(b+\delta).\tag{26}\] Combining 26 with 25 , we finally get \[h_{\alpha,\beta}(c)+h_{\alpha,\beta}(d)=h_{\alpha,\beta}(a-\delta)+h_{\alpha,\beta}(b+\delta-\varepsilon)\geq h_{\alpha,\beta}(a)+h_{\alpha,\beta}(b).\] Recalling the definitions of \(a,b,c\) and \(d\), and [eq:sharma95m95as95tsa] we obtain \[S_{\alpha,\beta}(\mathbf{w})+S_{\alpha,\beta}(\mathbf{v}) \ge S_{\alpha,\beta}(\mathbf{p})+S_{\alpha,\beta}(\mathbf{q}),\] which demonstrates that \(S_{\alpha,\beta}\) is supermodular in the range \(\alpha>1\) and \(\beta\le \alpha\).
Case: \(\alpha=1\). We recall that the Sharma-Mittal entropy \(S_{\alpha,\beta}(\mathbf{p})\) can be expressed as a function of the Rényi entropy \(H_\alpha(\mathbf{p})\) as shown in 7 . Thus, as \(\alpha\to1\), the Rényi entropy converges to the Shannon entropy \(H(p)\), yielding the following representation: \[\label{eq:S951} S_{1,\beta}(\mathbf{p}) =\phi_\beta(H(\mathbf{p})),\tag{27}\] where the function \(\phi_\beta(x)\) is defined as: \[\label{eq:phi} \phi_\beta(x)=\begin{cases} \frac{2^{(1-\beta)x}-1}{1-\beta}, &for \beta\neq 1,\\ \ln(2)x, & for\beta=1. \end{cases}\tag{28}\] Given that the Shannon entropy \(H(\mathbf{p})\) is Schur-concave, subadditive and supermodular on the majorization lattice [4], and that, for \(\beta\leq\alpha=1\), the first derivative \(\phi'_\beta(x)\) and second derivative \(\phi''_\beta\) of the function \(\phi_\beta(x)\) are non negative, we can proceed following the same idea of the Case \(\alpha\in(0,1)\). Let \[a=H(\mathbf{p}), \qquad b=H(\mathbf{q}), \qquad c=H(\mathbf{w}), \qquad d=H(\mathbf{v}).\] Because Shannon entropy \(H(\mathbf{p})\) is Schur-concave and \[\mathbf{w}\preceq \mathbf{p},\; \mathbf{q}\preceq \mathbf{v},\] it follows \[c\ge a, \qquad c\ge b, \qquad a\ge d, \qquad b\ge d.\] Moreover, without loss of generality, assuming \(a\geq b\), yields \[\label{exjvpbag} c\ge a\ge b\ge d.\tag{29}\] Since the Shannon entropy satisfies the supermodularity property, we have \[\label{uteqzobv} c+d\ge a+b.\tag{30}\] Let us define \[\delta=(c-a)\ge0.\] From ([eq:2]), it holds that \[d-b\ge -\delta.\] Hence there exists \(\varepsilon\ge0\) such that \[\label{ctqoriev} d=b-\delta+\varepsilon.\tag{31}\] Therefore \[(c,d) = (a+\delta,\;b-\delta+\varepsilon).\] Define the function \(F(t)\) as \[F(t)=\phi_\beta(a+t)+\phi_\beta(b-t), \qquad 0\le t\le \delta.\] where \(\phi_\beta\) is defined in (28 ). We first check that the argument \(b-t\geq b-\delta\) remains non negative, therefore it stays within the domain of \(\phi_\beta.\) Since the Shannon entropy is subadditive on the majorization lattice, we get that \(c=H(\mathbf{p}\wedge \mathbf{q})\leq H(\mathbf{p})+H(\mathbf{q})=a+b.\) Therefore \(c\leq a+b\) ensures \(b-t\geq b-\delta=a+b-c\geq 0\) for \(t\in[0,\delta].\)
Because \(\beta\le\alpha=1\), we have that \(\phi_\beta''(x)=(\ln 2)^2(1-\beta)2^{(1-\beta)x}\ge0\). Therefore, the function \(\phi_\beta\) is convex, and its derivative \(\phi_\beta'\) is increasing. Differentiating \(F(t)\) yields \[F'(t) = \phi_\beta'(a+t)-\phi_\beta'(b-t).\] From the chain of inequalities in ([eq:1]), it follows that \[a+t\ge b-t,\] and since \(\phi_\beta'\) is increasing, we obtain that \[F'(t)\ge0.\] Thus, the function \(F(t)\) is increasing in \(t\). Evaluating \(F(t)\) at \(t=\delta\) and \(t=0\) gives \[\label{iprsvbyw} F(\delta)=\phi_\beta(a+\delta)+\phi_\beta(b-\delta) \ge \phi_\beta(a)+\phi_\beta(b)=F(0).\tag{32}\] Finally, since \(\phi_\beta\) is increasing and \(\varepsilon\ge0\), we obtain the inequality \[\label{xfnpyqlg} \phi_\beta(b-\delta+\varepsilon) \ge \phi_\beta(b-\delta).\tag{33}\] Combining ([ineq]) with ([eq:4]),we get \[\phi_\beta(c)+\phi_\beta(d) = \phi_\beta(a+\delta)+\phi_\beta(b-\delta+\varepsilon) \ge \phi_\beta(a)+\phi_\beta(b).\] Recalling the definitions of \(a,b,c, d\), and (27 ) we obtain \[S_{1,\beta}(\mathbf{w})+S_{1,\beta}(\mathbf{v}) \ge S_{1,\beta}(\mathbf{p})+S_{1,\beta}(\mathbf{q}),\] which demonstrates that \(S_{\alpha,\beta}\) is also supermodular for \(\alpha=1\) and \(\beta\le \alpha\). ◻
In this section, we show that the Sharma-Mittal entropy \(S_{\alpha,\beta}(\mathbf{p})\) is generally neither supermodular nor submodular on the majorization lattice for \(\beta > \alpha\), and \(\alpha>0\). To this purpose, we construct two explicit counterexamples by finding one pair of probability distributions that breaks supermodularity and another pair that breaks submodularity. Recall the definition of the Sharma-Mittal entropy: \[S_{\alpha, \beta}(\mathbf{p}) = \frac{1}{1-\beta} \left[ \left( \sum_{i=1}^n p_i^\alpha \right)^{\frac{1-\beta}{1-\alpha}} - 1 \right].\]
For our counterexamples, let us fix \(\alpha = 2\) and \(\beta = 3\). This choice simplifies the entropy formula to: \[S_{2,3}(\mathbf{p}) = \frac{1}{-2} \left[ \left( \sum_{i=1}^n p_i^2 \right)^{\frac{-2}{-1}} - 1 \right] = \frac{1}{2} \left( 1 - \left( \sum_{i=1}^n p_i^2 \right)^2 \right)\]
We also recall that in the majorization lattice, given two probability distributions \(\mathbf{p}\) and \(\mathbf{q}\) ordered in a non-increasing fashion, their least upper bound \(\mathbf{p} \vee \mathbf{q}\) and greatest lower bound \(\mathbf{p} \wedge \mathbf{q}\) are constructed by taking the unique distributions whose cumulative sums correspond, respectively, to the component-wise maximum and minimum of the cumulative sums of \(\mathbf{p}\) and \(\mathbf{q}\) (see 3 and 2 for further details).
We can now present the two counterexamples in the following subsections.
In order to disprove supermodularity, we must show an instance, i.e., a pair of probability distributions \(\mathbf{p}\) and \(\mathbf{q}\), where the property fails, that is, \[S_{2,3}(\mathbf{p} \vee \mathbf{q}) + S_{2,3}(\mathbf{p} \wedge \mathbf{q}) < S_{2,3}(\mathbf{p}) + S_{2,3}(\mathbf{q}).\]
For this purpose, let \(n=4\) and consider the following distributions: \[\mathbf{p} = (0.5, 0.3, 0.1, 0.1)\quad\text{and }\quad\mathbf{q} = (0.4, 0.4, 0.2, 0.0).\] One can verify that in the majorization lattice their least upper bound and greatest lower bound are, respectively: \[\mathbf{p} \vee \mathbf{q} = (0.5, 0.3, 0.2, 0.0)\quad\text{and }\quad\mathbf{p} \wedge \mathbf{q} = (0.4, 0.4, 0.1, 0.1).\] Let us evaluate their Sharma-Mittal entropy \(S_{\alpha,\beta}\) for \(\alpha=2\) and \(\beta=3>1\): \[\begin{align} S_{2,3}(\mathbf{p})&=\frac{1}{2}-\frac{1}{2}\left(0.5^2+0.3^2+0.1^2+0.1^2\right)^2={0.4352}\\ S_{2,3}(\mathbf{q})&=\frac{1}{2}-\frac{1}{2}\left(0.4^2+0.4^2+0.2^2+0.0^2\right)^2={0.4352}\\ S_{2,3}(\mathbf{p} \vee \mathbf{q})&=\frac{1}{2}-\frac{1}{2}\left(0.5^2+0.3^2+0.2^2+0.0^2\right)^2={0.4278}\\ S_{2,3}(\mathbf{p} \wedge \mathbf{q})&=\frac{1}{2}-\frac{1}{2}\left(0.4^2+0.4^2+0.1^2+0.1^2\right)^2={0.4422} \end{align}\] Then, one can see that for the chosen pair of probability distributions, the following inequality holds \[S_{2,3}(\mathbf{p})+S_{2,3}(\mathbf{q})={0.8704}>{0.8700}=S_{2,3}(\mathbf{p} \vee \mathbf{q})+S_{2,3}(\mathbf{p} \wedge \mathbf{q}),\] which shows that \(S_{2,3}\) is not supermodular.
In a similar fashion to the previous section, to disprove submodularity, we must exhibit an instance where the submodular property fails, that is, \[S_{2,3}(\mathbf{p} \vee \mathbf{q}) + S_{2,3}(\mathbf{p} \wedge \mathbf{q}) > S_{2,3}(\mathbf{p}) + S_{2,3}(\mathbf{q}).\]
Let \(n=4\) and consider the following pair of distributions: \[\boldsymbol{p}=(0.5,0.2,0.2,0.1)\quad\text{and }\quad\boldsymbol{q}=(0.4,0.4,0.15,0.05).\] Their respective least upper bound and greatest lower bound are: \[\mathbf{p} \vee \mathbf{q}=(0.5,0.3,0.15,0.05)\quad\text{and }\quad\mathbf{p} \wedge \mathbf{q}=(0.4,0.3,0.2,0.1).\] Evaluating their Sharma-Mittal entropy \(S_{\alpha,\beta}\) for \(\alpha=2\) and \(\beta=3\) yields: \[\begin{align} S_{2,3}(\mathbf{p})&=\frac{1}{2}-\frac{1}{2}\left(0.5^2+0.2^2+0.2^2+0.1^2\right)^2={0.4422}\\ S_{2,3}(\mathbf{q})&=\frac{1}{2}-\frac{1}{2}\left(0.4^2+0.4^2+0.15^2+0.05^2\right)^2={0.4404875}\\ S_{2,3}(\mathbf{p} \vee \mathbf{q})&=\frac{1}{2}-\frac{1}{2}\left(0.5^2+0.3^2+0.15^2+0.05^2\right)^2={0.4333875}\\ S_{2,3}(\mathbf{p} \wedge \mathbf{q})&=\frac{1}{2}-\frac{1}{2}\left(0.4^2+0.3^2+0.2^2+0.1^2\right)^2={0.455} \end{align}\] Consequently, for this choice of probability distributions, the following inequality holds \[S_{2,3}(\mathbf{p} \vee \mathbf{q})+S_{2,3}(\mathbf{p} \wedge \mathbf{q})=0.8883875>0.8826875= S_{2,3}(\mathbf{p})+ S_{2,3}(\mathbf{q}),\] which demonstrates that \(S_{2,3}\) is not submodular.
Taken together, these two counterexamples demonstrate that for \(\beta > \alpha\), and \(\alpha>0\), the Sharma-Mittal entropy is neither supermodular nor submodular on the majorization lattice.
arXiv preprint arXiv:2605.09655, 2026.