January 01, 1970
We extend the entropy formula of Menon and Yu for the real Deep Linear Network (DLN) to its complex and quaternionic analogues, obtaining a unified formula for DLNs over \(\mathbb{R}\), \(\mathbb{C}\), and \(\mathbb{H}\).
In this paper, we compute the Boltzmann entropy of functions represented by deep linear networks (DLNs) over \(\mathbb{R}\), \(\mathbb{C}\), and \(\mathbb{H}\). The DLN is a basic model for the geometry of overparametrization in deep learning, with a precise geometric and thermodynamic description of its training dynamics [1].
Given depth \(N\ge 2\) and width \(d\ge 1\), a DLN over \(\mathbb{F}\) is described by the multiplication map \[\phi:{M}_{d}(\mathbb{F})^N\longrightarrow {M}_{d}(\mathbb{F}),\qquad \phi(W_N,\ldots,W_1)=W_N\cdots W_1 .\] The space of parameters is \({M}_{d}(\mathbb{F})^N\) while the space of observables is \({M}_{d}(\mathbb{F})\). For a parameter \(\mathbf{W}\), we write \[\mathbf{W}=(W_N,\ldots,W_1)\in {M}_{d}(\mathbb{F})^N,\qquad X=\phi(\mathbf{W})=W_N\cdots W_1,\] and call \(X\) the end-to-end matrix of \(\mathbf{W}\). Here \(\mathbb{F}\in\{\mathbb{R},\mathbb{C},\mathbb{H}\}\), where \(\mathbb{R}\), \(\mathbb{C}\), and \(\mathbb{H}\) denote the real numbers, complex numbers, and quaternions, respectively. We write \({M}_{d}(\mathbb{F})\) for the space of \(d\times d\) matrices over \(\mathbb{F}\), with \(GL_d(\mathbb{F})\subset {M}_{d}(\mathbb{F})\) denoting the invertible matrices. We also set \(\beta:=\dim_{\mathbb{R}}\mathbb{F}\), so \(\beta=1,2,4\) for \(\mathbb{F}=\mathbb{R},\mathbb{C},\mathbb{H}\), respectively.
Given \(X\in GL_d(\mathbb{F})\), we consider the set of balanced factorizations of \(X\), given by \[\mathcal{O}_X^{\beta} := \left\{ \mathbf{W} \in GL_d(\mathbb{F})^N : W_N\cdots W_1=X,\quad W_{p+1}^*W_{p+1}=W_pW_p^* \;\text{for } 1\le p\le N-1 \right\}.\] The Boltzmann entropy of \(X\) is the logarithm of the volume, \[S^{\beta}(X):=\log\operatorname{vol}(\mathcal{O}_X^{\beta}),\] where the volume is taken with respect to the metric induced from the real trace pairing on \({M}_{d}(\mathbb{F})^N\). Our main result (1) is an explicit formula for the entropy, generalizing Menon and Yu’s result [2] from the case \(\beta=1\) to the cases \(\beta=1, 2, 4\).
Let \(O_d\), \(U_d\), and \(Sp_d\) denote the orthogonal, unitary, and compact symplectic groups, respectively. Collectively, we write \(K_1=O_d\), \(K_2=U_d\), and \(K_4=Sp_d\), and let \(c_{\beta}:=\operatorname{vol}(K_{\beta})\), where the volume is taken with respect to the standard bi-invariant metric on \(K_{\beta}\).
We denote by \(\operatorname{van}(D)\) the Vandermonde determinant of \(D\). For a diagonal matrix \(D=\operatorname{diag}(\lambda_1,\ldots,\lambda_d),\) with real diagonal entries \(\lambda_1\ge \lambda_2\ge \cdots \ge \lambda_d,\) set \[\operatorname{van}(D):=\prod_{1\le s<l\le d}(\lambda_s-\lambda_l).\]
Theorem 1 (Entropy formula). Let \(X\in GL_d(\mathbb{F})\) have distinct singular values \(\sigma_1>\cdots>\sigma_d>0,\) and write \(\Sigma=\operatorname{diag}(\sigma_1,\ldots,\sigma_d)\). Then \[\label{eq:entropy-formula-vandermonde} S^{\beta}(X) = (N-1)\log c_{\beta} +\frac{\beta}{2}\log\frac{\operatorname{van}(\Sigma^2)}{\operatorname{van}(\Sigma^{2/N})} +(\beta-1) \left( \frac{d}{2}\log N+ \log\frac{\det\Sigma}{\det(\Sigma^{1/N})} \right).\tag{1}\] Equivalently, \[\label{eq:entropy-formula-entrywise} \begin{align} S^{\beta}(X) ={}& (N-1)\log c_{\beta} +\frac{\beta}{2}\sum_{1\le s<l\le d} \log\left( \frac{\sigma_s^2-\sigma_l^2}{\sigma_s^{2/N}-\sigma_l^{2/N}} \right)\\ &+(\beta-1) \left( \frac{d}{2}\log N+ \left(1-\frac{1}{N}\right)\sum_{s=1}^d\log\sigma_s \right). \end{align}\tag{2}\]
The assumption of distinct singular values is used to identify the orbit \(\mathcal{O}_X^\beta\) with \(K_\beta^{N-1}\) via 3. The right-hand side of 2 , however, extends real-analytically through collisions of singular values [3]. Indeed, for \(1\le s<l\le d\), if \(a=\sigma_s^{2/N}\) and \(b=\sigma_l^{2/N}\), then \[\label{eq:intro-finite-sum-extension} \frac{\sigma_s^2-\sigma_l^2}{\sigma_s^{2/N}-\sigma_l^{2/N}} = \frac{a^N-b^N}{a-b} = \sum_{m=0}^{N-1}a^{N-1-m}b^m,\tag{3}\] a positive polynomial for \(a,b>0\). Throughout the paper, quotients of this form are understood by 3 when singular values collide.
Remark 1 (Analogy with RMT). The parameter \(\beta=\dim_{\mathbb{R}}\mathbb{F}\) plays a role analogous to the Dyson index in random matrix theory (RMT). In the classical Gaussian ensembles [4], the volumes of isospectral orbits carry Vandermonde factors with exponents \(\beta=1,2,4\) [5]. The entropy formula 2 contains an analogous ratio of Vandermondes, \[\prod_{1\le s<l\le d} \left( \frac{\sigma_s^2-\sigma_l^2}{\sigma_s^{2/N}-\sigma_l^{2/N}} \right)^{\beta/2}.\] An important difference is worth emphasizing. By 3 , the entropy term exhibits no singular-value repulsion, in contrast with the usual logarithmic repulsion in Dyson Brownian motion. Unlike the \(\beta\)-ensembles of Dumitriu and Edelman [6], we do not presently have a tridiagonal matrix model for the entropy formula at arbitrary \(\beta>0\).
For \(\mathbb{F}=\mathbb{R}\), 1 recovers the entropy formula for the real DLN [2]. Menon and Yu’s proof proceeds by diagonalizing certain tridiagonal Jacobi matrices using the theory of Chebyshev polynomials. The same ideas also yield an orthonormal basis for the tangent space of the balanced manifold. The resulting basis allows them to prove that the multiplication map \(\phi\) restricts to a Riemannian submersion from the balanced manifold in parameter space onto the space of end-to-end matrices [2]. In this paper, we do not compute the full orthonormal basis or pursue the corresponding Riemannian submersion for the complex and quaternionic DLNs. For the Kähler reduction of the complex DLN, see [7].
We recall the DLN metric \(g_N\) on \(GL_d(\mathbb{F})\), the space of invertible end-to-end matrices, following [2], [8]. For \(X\in GL_d(\mathbb{F})\), define the real-linear operator \[\mathcal{A}_{N,X}(Z) = \sum_{p=1}^N (XX^*)^{(N-p)/N}\,Z\,(X^*X)^{(p-1)/N}, \qquad Z\in {M}_{d}(\mathbb{F}).\] The powers are defined by functional calculus for positive Hermitian matrices. We regard \({M}_{d}(\mathbb{F})\) as a real inner product space with the pairing \(\langle Y,Z\rangle:=\operatorname{Re}\operatorname{Tr}(Y^*Z)\). With respect to this pairing, \(\mathcal{A}_{N,X}\) is positive. Under the canonical identification \(T_XGL_d(\mathbb{F})={M}_{d}(\mathbb{F})\), we define the DLN metric by \[\label{eq:intro-dln-metric} g_{N,X}(Z_1,Z_2) = \operatorname{Re}\operatorname{Tr}\left(Z_1^*\,\mathcal{A}_{N,X}^{-1}(Z_2)\right), \qquad Z_1,Z_2\in {M}_{d}(\mathbb{F}).\tag{4}\] All tangent spaces and determinants below are understood over the underlying real vector spaces. In the quaternionic case, the same formulas are justified by the cyclicity of \(\operatorname{Re}\operatorname{Tr}\).
Proposition 1. Let \(X\in GL_d(\mathbb{F})\) have singular value decomposition \(X=U_N\Sigma U_0^*\), and assume that the singular values of \(X\) are distinct. Let \(\beta=\dim_{\mathbb{R}}\mathbb{F}\). Then \[\label{eq:division-volume-determinant} \operatorname{vol}(\mathcal{O}_X^{\beta})\, \det_{\mathbb{R}}(\mathcal{A}_{N,X}^{-1})^{1/4} = c_{\beta}^{N-1} N^{(\beta-2)d/4} \left( \frac{\det\Sigma}{\det(\Sigma^{1/N})} \right)^{(\beta-2)/2}.\qquad{(1)}\] Equivalently, \[\label{eq:division-entropy-determinant} S^{\beta}(X) = (N-1)\log c_{\beta} -\frac{1}{4}\log\det_{\mathbb{R}}(\mathcal{A}_{N,X}^{-1}) +\frac{\beta-2}{4}d\log N +\frac{\beta-2}{2}\log\frac{\det\Sigma}{\det(\Sigma^{1/N})}.\qquad{(2)}\]
When \(\mathbb{F}=\mathbb{C}\), we have \(\beta=2\), and ?? reduces to \[\label{eq:complex-volume-determinant} \operatorname{vol}(\mathcal{O}_X^{2})\, \det_{\mathbb{R}}(\mathcal{A}_{N,X}^{-1})^{1/4} = c_{2}^{N-1}.\tag{5}\]
Remark 2 (Infinite depth and renormalized entropy). Consider the renormalized operator \[\frac{1}{N}\mathcal{A}_{N,X}(Z) \longrightarrow \mathcal{A}_{\infty,X}(Z) := \int_0^1 (XX^*)^{1-t} Z (X^*X)^t\,{d}t,\] which defines the metric \(g_{\infty,X}(Z_1,Z_2) := \operatorname{Re}\operatorname{Tr}\left( Z_1^*\mathcal{A}_{\infty,X}^{-1}(Z_2) \right)\) for the DLN in the infinite-depth limit [9].
The limiting entropy can be obtained from 1 by subtracting terms that diverge as \(N\to\infty\). The same renormalization applied to 1 provides another expression for the entropy formula at infinite depth. Passing to the limit gives \[S^\beta_{\infty}(X) = -\frac{1}{4}\log\det_{\mathbb{R}}(\mathcal{A}_{\infty,X}^{-1}) +\frac{\beta-2}{2}\log\det\Sigma .\] Evaluating the determinant in terms of the singular values \(\sigma_1,\ldots,\sigma_d\) of \(X\) gives \[S^\beta_{\infty}(X) = (\beta-1)\sum_{s=1}^d \log\sigma_s + \frac{\beta}{2} \sum_{1\le s<l\le d} \log\!\left( \frac{\sigma_s^2-\sigma_l^2}{\log\sigma_s^2-\log\sigma_l^2} \right).\] As in the formula at finite depth, this expression is understood by continuity at repeated singular values.
In 2, we set notation for matrices over \(\mathbb{R}\), \(\mathbb{C}\), and \(\mathbb{H}\), and introduce the balanced manifold and the orbit \(\mathcal{O}_X^{\beta}\). In 3, we prove 3, identifying \(\mathcal{O}_X^{\beta}\) with \(K_{\beta}^{N-1}\) for matrices with distinct singular values. In 4, we compute the metric induced on this orbit by the ambient real trace pairing. In 5, we evaluate the resulting block determinants and prove 1 and 1.
For \(A=(a_{sl})\in {M}_{d}(\mathbb{F})\), set \(A^*=(\overline{a_{ls}}).\) Thus \(A^*\) is the transpose over \(\mathbb{R}\), the conjugate transpose over \(\mathbb{C}\), and the quaternionic adjoint over \(\mathbb{H}\).
All inner products are real. We use the pairing \(\left\langle A,B\right\rangle:=\operatorname{Re}\operatorname{Tr}(A^*B),\) where \(A,B\in{M}_{d}(\mathbb{F})\). For quaternionic matrices, we have the cyclic property \(\operatorname{Re}\operatorname{Tr}(AB)=\operatorname{Re}\operatorname{Tr}(BA)\) whenever both products are defined.
We use the singular value decomposition (SVD) over \(\mathbb{R},\mathbb{C},\mathbb{H}\). Thus every \(X\in{M}_{d}(\mathbb{F})\) admits a decomposition \[X=U\Sigma V^*, \qquad \Sigma=\operatorname{diag}(\sigma_1,\dots,\sigma_d), \qquad \sigma_1\ge\cdots\ge\sigma_d\ge 0,\] with \(U,V\in O_d\), \(U_d\), or \(Sp_d\), respectively, when \(\mathbb{F}=\mathbb{R},\mathbb{C}\) or \(\mathbb{H}\). In the quaternionic case this follows from the spectral theorem for quaternionic Hermitian matrices [10].
For \(\mathbf{A}, \mathbf{B} \in {M}_{d}(\mathbb{F})^N\) where \(\mathbf{A}=(A_N,\dots,A_1), \mathbf{B}=(B_N,\dots,B_1)\) set \[\label{eq:ambient-inner-product} \left\langle \mathbf{A},\mathbf{B}\right\rangle := \sum_{p=1}^N \operatorname{Re}\operatorname{Tr}(A_p^*B_p).\tag{6}\] This inner product defines the Euclidean metric on \({M}_{d}(\mathbb{F})^N\).
Recall the multiplication map \[\phi:{M}_{d}(\mathbb{F})^N\to{M}_{d}(\mathbb{F}), \qquad \phi(\mathbf{W})=W_N\cdots W_1 .\] For \(X\in GL_d(\mathbb{F})\), the fibre over \(X\) is \[\mathcal{F}_X^\beta:=\phi^{-1}(X).\] Since \(X\) is invertible, every element of \(\mathcal{F}_X^\beta\) has invertible layers.
The equations for balancedness are \[W_{p+1}^*W_{p+1}=W_pW_p^*, \qquad 1\le p\le N-1.\] Together, they define the balanced variety \[\mathcal{M}_{\mathbf{0}}^\beta := \left\{ \mathbf{W}\in{M}_{d}(\mathbb{F})^N: W_{p+1}^*W_{p+1}=W_pW_p^* \text{ for }1\le p\le N-1 \right\}.\] On \(GL_d(\mathbb{F})^N\), the same equations define the balanced manifold \[\mathcal{M}^\beta:=\mathcal{M}_{\mathbf{0}}^\beta\cap GL_d(\mathbb{F})^N .\] Hence, the set of balanced factorizations of \(X\) is \[\mathcal{O}_X^\beta:=\mathcal{F}_X^\beta\cap\mathcal{M}^\beta .\] When \(X\) has distinct singular values, by 3, we have that \(\mathcal{O}_X^\beta\) is a compact smooth submanifold of \({M}_{d}(\mathbb{F})^N\). Its volume is always taken with respect to the metric induced by 6 .
Recall that \(K_\beta\) denotes \(O_d\), \(U_d\), or \(Sp_d\) according as \(\beta=1,2,4\), respectively, and write its Lie algebra as \[\mathfrak{k}_\beta := \left\{A\in{M}_{d}(\mathbb{F}):A^*=-A\right\}.\] We equip \(\mathfrak{k}_\beta\) with the real inner product \[\label{eq:lie-inner-product} \left\langle A,B\right\rangle_{\mathfrak{k}} := \operatorname{Re}\operatorname{Tr}(A^*B).\tag{7}\]
The group \(K_\beta^{N-1}\) acts on \({M}_{d}(\mathbb{F})^N\) by \[\label{eq:compact-action} (U_{N-1},\dots,U_1)\cdot(W_N,\dots,W_1) := \bigl( W_NU_{N-1}^*, U_{N-1}W_{N-1}U_{N-2}^*, \dots, U_1W_1 \bigr).\tag{8}\] The action is isometric, preserves the end-to-end product, and preserves the equations for balancedness. Indeed, the intermediate factors cancel in the product, and for \(1\le p\le N-1\), \[\widetilde{W}_{p+1}^*\widetilde{W}_{p+1} = U_pW_{p+1}^*W_{p+1}U_p^*, \qquad \widetilde{W}_p\widetilde{W}_p^* = U_pW_pW_p^*U_p^* .\] Hence each \(\mathcal{O}_X^\beta\) is invariant under 8 .
We also use the action of \(K_\beta\times K_\beta\) on \({M}_{d}(\mathbb{F})^N\) given by \[\label{eq:endpoint-action} (Q,R)\cdot(W_N,\dots,W_1) := (QW_N,W_{N-1},\dots,W_2,W_1R^*).\tag{9}\] This action is also isometric.
Proposition 2. Let \(X\in GL_d(\mathbb{F})\) have distinct singular values, and let \(Q,R\in K_\beta\). The endpoint action 9 restricts to an isometry \[\Psi_{Q,R}:\mathcal{O}_X^\beta\longrightarrow\mathcal{O}_{QXR^*}^\beta, \qquad \Psi_{Q,R}(\mathbf{W})=(Q,R)\cdot\mathbf{W} .\] In particular, if \(X=U_N\Sigma U_0^*\) is a singular value decomposition, then \[S^\beta(X)=S^\beta(\Sigma).\]
Proof. Let \(\mathbf{W}=(W_N,\dots,W_1)\in\mathcal{F}_X^{\beta}\). Then \[\phi\bigl((Q,R)\cdot\mathbf{W}\bigr) = QW_N\cdots W_1R^* = QXR^*.\] Thus the endpoint action sends \(\mathcal{F}_X^{\beta}\) to \(\mathcal{F}_{QXR^*}^{\beta}\).
The only equations that can change are those with \(p=1\) or \(p=N-1\). For these equations, \[(QW_N)^*(QW_N)=W_N^*W_N, \qquad (W_1R^*)(W_1R^*)^*=W_1W_1^*.\] Hence the equations for balancedness are preserved. The action also preserves invertibility, so it sends \(\mathcal{M}^{\beta}\) to itself. Therefore it restricts to a map \[\mathcal{O}_X^{\beta}\longrightarrow\mathcal{O}_{QXR^*}^{\beta}.\] Its inverse is the restriction of the endpoint action by \((Q^*,R^*)\). Since the endpoint action preserves 6 , the restricted map is an isometry.
Taking \(Q=U_N^*\) and \(R=U_0^*\) gives an isometry \[\mathcal{O}_X^{\beta}\longrightarrow\mathcal{O}_{\Sigma}^{\beta}.\] The volumes are equal, and the entropy identity follows. ◻
Choose and fix a singular value decomposition \(X=U_N\Sigma U_0^*,\) and set \(\Lambda=\Sigma^{1/N}\), where the positive \(N\)-th root is taken. The center associated with this chosen SVD is the balanced factorization obtained by distributing the singular values evenly across the layers \[\mathbf{C}_X=(U_N\Lambda,\Lambda,\ldots,\Lambda U_0^*).\]
Proposition 3. Let \(X\in GL_d(\mathbb{F})\) have distinct singular values. Then the map \[\label{eq:orbit-map} \mathfrak{y}_X:K_{\beta}^{N-1}\to \mathcal{O}_X^{\beta}\qquad{(3)}\] defined by \[\mathfrak{y}_X(U_{N-1},\ldots,U_1) = \bigl( U_N\Lambda U_{N-1}^*,\, U_{N-1}\Lambda U_{N-2}^*,\, \ldots,\, U_1\Lambda U_0^* \bigr)\] is a diffeomorphism. In particular, \(\mathcal{O}_X^{\beta}=K_{\beta}^{N-1}\cdot \mathbf{C}_X.\)
The product of the displayed factors telescopes to \(X\), and the equations for balancedness are immediate. Hence \(\mathfrak{y}_X\) is well defined as a map into \(\mathcal{O}_X^\beta\). The proof that \(\mathfrak{y}_X\) is a diffeomorphism is presented in 3.
The map \(\mathfrak{y}_X\) parametrizes the orbit \(\mathcal{O}_X^\beta\) for a fixed end-to-end matrix \(X\). One could similarly parametrize the full balanced manifold \(\mathcal{M}^\beta\), allowing \(X\) to vary, but here we restrict attention to the orbit map \(\mathfrak{y}_X\). In the real case, the Riemannian geometry of \(\mathcal{M}^\beta\) is developed in [2].
We use the ordered SVD fixed in the setup for 3, \[X=U_N\Sigma U_0^*, \qquad \Sigma=\operatorname{diag}(\sigma_1,\dots,\sigma_d), \qquad \sigma_1>\cdots>\sigma_d>0,\] and set \(\Lambda=\Sigma^{1/N}\). Since the singular values are distinct and ordered, the diagonal factor \(\Sigma\) is determined by \(X\). The only remaining ambiguity in the singular-vector factors is simultaneous multiplication on the right by the same diagonal unit. We therefore introduce the diagonal stabilizer \[\Delta_{\beta}:=Z_{K_{\beta}}(\Sigma) = \{t\in K_{\beta}:t\Sigma t^*=\Sigma\}.\] Since the singular values are distinct, this stabilizer is explicitly \[\Delta_{\beta} = \begin{cases} \left\{\operatorname{diag}(\varepsilon_1,\dots,\varepsilon_d):\varepsilon_s=\pm 1\right\}, & \mathbb{F}=\mathbb{R},\\ \left\{\operatorname{diag}(z_1,\dots,z_d):|z_s|=1\right\}, & \mathbb{F}=\mathbb{C},\\ \left\{\operatorname{diag}(q_1,\dots,q_d):q_s\in Sp_1\right\}, & \mathbb{F}=\mathbb{H}. \end{cases}\] Indeed, every \(t\in\Delta_{\beta}\) gives the same singular value decomposition \[X=(U_Nt)\Sigma(U_0t)^*,\] since \(t\Sigma t^*=\Sigma\). Conversely, as shown in 2 below, every ordered SVD of \(X\) with diagonal factor \(\Sigma\) arises uniquely in this way.
Lemma 1. Let \(\mathbf{W}=(W_N,\dots,W_1)\in \mathcal{M}^{\beta}\). Then there exist matrices \(Q_0,\dots,Q_N\in K_{\beta}\) and a positive diagonal matrix \(\Lambda_{\mathbf{W}}\), with diagonal entries in nonincreasing order, such that \[W_p=Q_p\Lambda_{\mathbf{W}}Q_{p-1}^*, \qquad 1\le p\le N.\]
Proof. Choose a singular value decomposition of the first layer, \(W_1=Q_1\Lambda_{\mathbf{W}}Q_0^*,\) where \(\Lambda_{\mathbf{W}}\) is positive diagonal with diagonal entries in nonincreasing order. Suppose inductively that \(W_p=Q_p\Lambda_{\mathbf{W}}Q_{p-1}^*\) has been constructed for some \(1\le p\le N-1\). Then \[W_pW_p^*=Q_p\Lambda_{\mathbf{W}}^2Q_p^*.\] Since \(\mathbf{W}\) is balanced, \[W_{p+1}^*W_{p+1}=W_pW_p^* =Q_p\Lambda_{\mathbf{W}}^2Q_p^*.\] Define \[Q_{p+1}:=W_{p+1}Q_p\Lambda_{\mathbf{W}}^{-1}.\] Because \(\mathbf{W}\in\mathcal{M}^{\beta}\), all layers are invertible, so this is well defined. Moreover, \[Q_{p+1}^*Q_{p+1} = \Lambda_{\mathbf{W}}^{-1}Q_p^*W_{p+1}^*W_{p+1}Q_p \Lambda_{\mathbf{W}}^{-1} = I.\] Thus \(Q_{p+1}\in K_{\beta}\), and \(W_{p+1}=Q_{p+1}\Lambda_{\mathbf{W}}Q_p^*.\) The claim follows by induction. ◻
The next lemma is the only point in the proof where distinct singular values are used.
Lemma 2. Assume that \(\Sigma=\operatorname{diag}(\sigma_1,\dots,\sigma_d)\) has distinct positive entries. If \[Q\Sigma R^*=\widetilde{Q}\Sigma \widetilde{R}^*, \qquad Q,R,\widetilde{Q},\widetilde{R}\in K_{\beta},\] then there exists a unique \(t\in \Delta_{\beta}\) such that \[\widetilde{Q}=Qt, \qquad \widetilde{R}=Rt.\]
Proof. Set \[A:=Q^*\widetilde{Q}, \qquad B:=R^*\widetilde{R}.\] Then \(A,B\in K_{\beta}\) and \(\Sigma=A\Sigma B^*.\) Multiplying by the adjoint gives \[\Sigma^2=A\Sigma^2A^*, \qquad \Sigma^2=B\Sigma^2B^*.\] Hence \(A\) and \(B\) commute with \(\Sigma^2\). Since \(\Sigma^2\) is diagonal with distinct real diagonal entries, both \(A\) and \(B\) are diagonal. Indeed, for \(s\ne l\), \[(\sigma_s^2-\sigma_l^2)A_{sl}=0, \qquad (\sigma_s^2-\sigma_l^2)B_{sl}=0.\] Thus \[A=\operatorname{diag}(a_1,\dots,a_d), \qquad B=\operatorname{diag}(b_1,\dots,b_d),\] with \(\left|a_s\right|=\left|b_s\right|=1\) for each \(s\). Returning to \(\Sigma=A\Sigma B^*\) and comparing diagonal entries gives \(a_s\sigma_s\overline{b_s}=\sigma_s,\) for \(1\le s\le d.\) Since \(\sigma_s>0\), we obtain \(a_s=b_s\) for every \(s\). Therefore \(A=B=t\) for some \(t\in\Delta_{\beta}\), and the stated formulas follow. Uniqueness follows from \(t=Q^*\widetilde{Q}\). ◻
Proof of 3. Viewed as a map to \({M}_{d}(\mathbb{F})^N\), \(\mathfrak{y}_X\) is smooth by construction. Its image lies in \(\mathcal{O}_X^{\beta}\) because the product of the factors telescopes to \(X\), and the equations for balancedness are immediate from the definition.
To prove surjectivity, let \(\mathbf{W}=(W_N,\dots,W_1)\in \mathcal{O}_X^{\beta}.\) By 1, there exist \(Q_0,\dots,Q_N\in K_{\beta}\) and a positive diagonal matrix \(\Lambda_{\mathbf{W}}\), with diagonal entries in nonincreasing order, such that \(W_p=Q_p\Lambda_{\mathbf{W}}Q_{p-1}^*,\) for \(1\le p\le N.\) Multiplying the factors gives \(X=Q_N\Lambda_{\mathbf{W}}^NQ_0^*.\) Since \(\Lambda_{\mathbf{W}}^N\) is positive diagonal with entries in nonincreasing order, this is a singular value decomposition of \(X\) with ordered singular values. Comparing with \(X=U_N\Sigma U_0^*\), whose singular values are strictly decreasing, we obtain \(\Lambda_{\mathbf{W}}^N=\Sigma,\) and \(\Lambda_{\mathbf{W}}=\Lambda.\) Now compare the two singular value decompositions \[X=U_N\Sigma U_0^*=Q_N\Sigma Q_0^*.\] By 2, there exists \(t\in\Delta_{\beta}\) such that \[Q_N=U_Nt, \qquad Q_0=U_0t.\] Set \[V_p:=Q_pt^*, \qquad 1\le p\le N-1.\] Since \(t\) commutes with \(\Lambda\), we have \[W_p=Q_p\Lambda Q_{p-1}^* =V_p\Lambda V_{p-1}^*, \qquad 2\le p\le N-1,\] and also \[W_N=U_N\Lambda V_{N-1}^*, \qquad W_1=V_1\Lambda U_0^*.\] Thus \(\mathbf{W}=\mathfrak{y}_X(V_{N-1},\dots,V_1),\) so \(\mathfrak{y}_X\) is surjective.
To prove injectivity, suppose \[\mathfrak{y}_X(A_{N-1},\dots,A_1)=\mathfrak{y}_X(B_{N-1},\dots,B_1).\] Comparing the first layer gives \(A_1\Lambda U_0^*=B_1\Lambda U_0^*,\) hence \(A_1=B_1\). If \(A_{p-1}=B_{p-1}\) for some \(2\le p\le N-1\), then comparison of the \(p\)-th layer gives \[A_p\Lambda A_{p-1}^* = B_p\Lambda B_{p-1}^* = B_p\Lambda A_{p-1}^*,\] so \(A_p=B_p\). Induction proves injectivity.
It remains to show that \(\mathfrak{y}_X\) is an immersion. For \(\mathbf{V},\mathbf{U}\in K_{\beta}^{N-1}\), the orbit map satisfies \[\mathfrak{y}_X(\mathbf{V}\mathbf{U})=\mathbf{V}\cdot \mathfrak{y}_X(\mathbf{U}),\] where multiplication on the left is componentwise and the dot denotes the \(K_{\beta}^{N-1}\)-action on \({M}_{d}(\mathbb{F})^N\). Hence it is enough to check the differential at the identity.
Let \(\mathbf{a}=(a_{N-1},\dots,a_1)\in \mathfrak k_{\beta}^{N-1}.\) Differentiating \[\tau\longmapsto \mathfrak{y}_X(\exp(\tau a_{N-1}),\dots,\exp(\tau a_1))\] at \(\tau=0\) gives the tangent vector with layer components \[\label{eq:general-tangent-formula} \begin{align} \bigl(({d}\mathfrak{y}_X)_{\mathbf{e}}(\mathbf{a})\bigr)_N &=-U_N\Lambda a_{N-1},\\ \bigl(({d}\mathfrak{y}_X)_{\mathbf{e}}(\mathbf{a})\bigr)_p &=a_p\Lambda-\Lambda a_{p-1}, \qquad 2\le p\le N-1,\\ \bigl(({d}\mathfrak{y}_X)_{\mathbf{e}}(\mathbf{a})\bigr)_1 &=a_1\Lambda U_0^*. \end{align}\tag{10}\] If \(({d}\mathfrak{y}_X)_{\mathbf{e}}(\mathbf{a})=0\), then the first layer gives \(a_1\Lambda U_0^*=0,\) so \(a_1=0\). If \(a_{p-1}=0\) for some \(2\le p\le N-1\), then the \(p\)-th layer gives \(a_p\Lambda=0,\) so \(a_p=0\). Hence \(a_1=\cdots=a_{N-1}=0\). Therefore \(({d}\mathfrak{y}_X)_{\mathbf{e}}\) is injective, and \(\mathfrak{y}_X\) is an immersion.
Since \(K_{\beta}^{N-1}\) is compact, the injective immersion \(\mathfrak{y}_X\) is proper as a map to \({M}_{d}(\mathbb{F})^N\). A proper injective immersion is an embedding. By surjectivity, its image is \(\mathcal{O}_X^{\beta}\). Thus \(\mathcal{O}_X^{\beta}\) is an embedded submanifold of \({M}_{d}(\mathbb{F})^N\), and \(\mathfrak{y}_X\) is a diffeomorphism onto \(\mathcal{O}_X^{\beta}\). Since \(K_{\beta}^{N-1}\) is compact, \(\mathcal{O}_X^{\beta}\) is compact.
Finally, \(\mathfrak{y}_X(\mathbf{e})= \mathbf{C}_X\), and the equivariance above gives \[\mathcal{O}_X^{\beta}=\mathfrak{y}_X(K_{\beta}^{N-1})=K_{\beta}^{N-1}\cdot \mathbf{C}_X.\] ◻
Corollary 1. At the diagonal point \(\mathbf{C}=(\Lambda,\dots,\Lambda)\in \mathcal{O}_{\Sigma}^{\beta}\), the tangent space is \[T_{\mathbf{C}}\mathcal{O}_{\Sigma}^{\beta} = \left\{\mathbf{c}(\mathbf{a}):\mathbf{a}\in \mathfrak k_{\beta}^{N-1}\right\},\] where \(\mathbf{c}(\mathbf{a})\) has layer components \[\label{eq:diagonal-tangent-formula} \begin{align} \mathbf{c}(\mathbf{a})_N &=-\Lambda a_{N-1},\\ \mathbf{c}(\mathbf{a})_p &=a_p\Lambda-\Lambda a_{p-1}, \qquad 2\le p\le N-1,\\ \mathbf{c}(\mathbf{a})_1 &=a_1\Lambda. \end{align}\tag{11}\]
Proof. Apply 3 to \(X=\Sigma\), so that \(U_N=U_0=I\) in 10 . Since \(\mathfrak{y}_{\Sigma}\) is a diffeomorphism, its differential at the identity is an isomorphism from \(\mathfrak k_{\beta}^{N-1}\) onto \(T_{\mathbf{C}}\mathcal{O}_{\Sigma}^{\beta}\). ◻
The map \(\mathfrak{y}_X\) parametrizes the compact orbit \(\mathcal{O}_X^{\beta}\) inside the fixed fibre \(\mathcal{F}_X^{\beta}\). Accordingly, 11 describes \(T_{\mathbf{C}}\mathcal{O}_{\Sigma}^{\beta}\), and not \(T_{\mathbf{C}}\mathcal{M}^{\beta}\). Tangent directions along the full balanced manifold also include directions in which the end-to-end matrix varies, equivalently variations of the singular values and the endpoint singular-vector factors. For the analogous separation between orbit directions in one fibre and directions in which the end-to-end matrix varies in the complex DLN, see [7].
By 2, the entropy depends only on the singular values of \(X\). Hence, from now on we work at the diagonal point \[\mathbf{C}=(\Lambda,\dots,\Lambda)\in \mathcal{O}_{\Sigma}^{\beta}.\] Write \(\Lambda=\operatorname{diag}(\lambda_1,\ldots,\lambda_d)\), so that \(\lambda_s=\sigma_s^{1/N}\). The tangent space at \(\mathbf{C}\) is described by 1. The next step is to choose an explicit real orthonormal basis of \(\mathfrak{k}_{\beta}\) and compute the induced metric on the corresponding orbit directions.
Set \[B_\beta= \begin{cases} \left\{1\right\}, & \mathbb{F}=\mathbb{R},\\ \left\{1,i\right\}, & \mathbb{F}=\mathbb{C},\\ \left\{1,i,j,k\right\}, & \mathbb{F}=\mathbb{H}, \end{cases} \qquad I_\beta=B_\beta\setminus\left\{1\right\}.\] Then \(|B_\beta|=\beta\) and \(|I_\beta|=\beta-1\). The set \(B_\beta\) is an orthonormal basis of \(\mathbb{F}\) over \(\mathbb{R}\), so for \(\varepsilon,\eta\in B_\beta\) one has \[\label{eq:basis-units-orthonormal} \operatorname{Re}(\overline{\varepsilon}\,\eta)=\delta_{\varepsilon\eta}.\tag{12}\] For \(1\le s<l\le d\) and \(\varepsilon\in B_\beta\), define \[\label{eq:offdiag-basis} T_{sl}^{\varepsilon}:=\frac{1}{\sqrt 2}\bigl(\varepsilon E_{sl}-\overline{\varepsilon}\,E_{ls}\bigr),\tag{13}\] and for \(1\le s\le d\) and \(\eta\in I_\beta\), define \[\label{eq:diag-basis} D_s^{\eta}:=\eta E_{ss}.\tag{14}\] Each of these matrices is skew-Hermitian, hence lies in \(\mathfrak{k}_{\beta}\). For \(\mathbb{F}=\mathbb{R}\), only the matrices \(T_{sl}^{1}\) occur. The matrices \(D_s^{\eta}\) come from the imaginary diagonal part of \(\mathfrak{k}_{\beta}\) and appear only in the complex and quaternionic cases.
Lemma 3. The set \[\left\{T_{sl}^{\varepsilon}:1\le s<l\le d,\;\varepsilon\in B_\beta\right\} \cup \left\{D_s^{\eta}:1\le s\le d,\;\eta\in I_\beta\right\}\] is a real orthonormal basis of \(\mathfrak{k}_{\beta}\) for the inner product 7 .
Proof. Orthogonality is a direct computation with matrix units. For \(q,r\in\mathbb{F}\), \[(qE_{ab})(rE_{cd})=\delta_{bc}\,qr\,E_{ad}.\] Together with 12 , this gives \[\operatorname{Re}\operatorname{Tr}\bigl((T_{sl}^{\varepsilon})^*T_{mn}^{\eta}\bigr) =\delta_{sm}\delta_{ln}\delta_{\varepsilon\eta}, \qquad \operatorname{Re}\operatorname{Tr}\bigl((D_s^{\eta})^*D_m^{\theta}\bigr) =\delta_{sm}\delta_{\eta\theta},\] while \[\operatorname{Re}\operatorname{Tr}\bigl((T_{sl}^{\varepsilon})^*D_m^{\eta}\bigr)=0.\] Thus the displayed family is orthonormal.
It remains to check that the family spans \(\mathfrak{k}_{\beta}\). A matrix \(A\in \mathfrak{k}_{\beta}\) has arbitrary purely imaginary diagonal entries and off-diagonal entries satisfying \(A_{ls}=-\overline{A_{sl}}\) for \(s<l\). Expanding each off-diagonal entry \(A_{sl}\in \mathbb{F}\) in the real basis \(B_\beta\) and each diagonal entry in the real basis \(I_\beta\) expresses \(A\) uniquely as a real linear combination of the matrices 13 and 14 . Therefore they form a real orthonormal basis. ◻
For \(1\le p\le N-1\), let \(\mathbf{e}_{p,sl}^{\varepsilon}\in \mathfrak{k}_{\beta}^{N-1}\) denote the element whose \(p\)th component is \(T_{sl}^{\varepsilon}\) and whose other components are zero. Likewise, let \(\mathbf{f}_{p,s}^{\eta}\in \mathfrak{k}_{\beta}^{N-1}\) denote the element whose \(p\)th component is \(D_s^{\eta}\) and whose other components are zero. By 3, the family \[\begin{align} &\left\{\mathbf{e}_{p,sl}^{\varepsilon}:1\le p\le N-1,\;1\le s<l\le d,\;\varepsilon\in B_\beta\right\}\\ &\cup\left\{\mathbf{f}_{p,s}^{\eta}:1\le p\le N-1,\;1\le s\le d,\;\eta\in I_\beta\right\} \end{align}\] is a real orthonormal basis of \(\mathfrak{k}_{\beta}^{N-1}\).
We now push this basis to the orbit through 11 . Set \[\label{eq:vertical-basis-vectors} \mathbf{v}_{p,sl}^{\varepsilon}:=\mathbf{c}(\mathbf{e}_{p,sl}^{\varepsilon}), \qquad \mathbf{w}_{p,s}^{\eta}:=\mathbf{c}(\mathbf{f}_{p,s}^{\eta}).\tag{15}\] Since \(\mathbf{c}:\mathfrak{k}_{\beta}^{N-1}\to T_{\mathbf{C}}\mathcal{O}_{\Sigma}^{\beta}\) is an isomorphism by 1, these vectors form a real basis of \(T_{\mathbf{C}}\mathcal{O}_{\Sigma}^{\beta}\).
Explicitly, for \(1\le r\le N\), where \(r\) denotes the layer index, \[(\mathbf{v}_{p,sl}^{\varepsilon})_r= \begin{cases} -\Lambda T_{sl}^{\varepsilon}, & r=p+1,\\ T_{sl}^{\varepsilon}\Lambda, & r=p,\\ 0, & \text{otherwise}, \end{cases} \qquad (\mathbf{w}_{p,s}^{\eta})_r= \begin{cases} -\Lambda D_s^{\eta}, & r=p+1,\\ D_s^{\eta}\Lambda, & r=p,\\ 0, & \text{otherwise}. \end{cases}\]
Let \(\iota_\Sigma\) denote the Riemannian metric on \(\mathcal{O}_{\Sigma}^{\beta}\) induced from the ambient inner product 6 .
Lemma 4. Among the vectors 15 , the following inner products, together with those obtained by interchanging the two arguments, are the only nonzero ones: \[\begin{align} \left\langle \mathbf{v}_{p,sl}^{\varepsilon},\mathbf{v}_{p,sl}^{\varepsilon}\right\rangle_{\iota_\Sigma} &=\lambda_s^2+\lambda_l^2 && (1\le p\le N-1), & \left\langle \mathbf{v}_{p,sl}^{\varepsilon},\mathbf{v}_{p+1,sl}^{\varepsilon}\right\rangle_{\iota_\Sigma} &=-\lambda_s\lambda_l && (1\le p\le N-2), \tag{16}\\ \left\langle \mathbf{w}_{p,s}^{\eta},\mathbf{w}_{p,s}^{\eta}\right\rangle_{\iota_\Sigma} &=2\lambda_s^2 && (1\le p\le N-1), & \left\langle \mathbf{w}_{p,s}^{\eta},\mathbf{w}_{p+1,s}^{\eta}\right\rangle_{\iota_\Sigma} &=-\lambda_s^2 && (1\le p\le N-2). \tag{17} \end{align}\] Any inner product that is neither listed in 16 –17 nor obtained from one of these by interchanging the two arguments vanishes.
Proof. Let \(p\) and \(q\) denote the depth labels of the two pushed-forward basis vectors. The layer index itself will be denoted by \(r\). If \(|p-q|>1\), then the two pushed-forward vectors have disjoint supports as functions of the layer index \(r\), so their inner product is zero. It is therefore enough to consider the cases \(q=p\) and \(q=p+1\); the case \(q=p-1\) follows by symmetry.
For the off-diagonal vectors, a direct computation from 13 gives \[\Lambda T_{sl}^{\varepsilon}=\frac{1}{\sqrt 2}\bigl(\lambda_s\varepsilon E_{sl}-\lambda_l\overline{\varepsilon}\,E_{ls}\bigr), \qquad T_{sl}^{\varepsilon}\Lambda=\frac{1}{\sqrt 2}\bigl(\lambda_l\varepsilon E_{sl}-\lambda_s\overline{\varepsilon}\,E_{ls}\bigr).\] Using the matrix-unit identities and 12 , we obtain \[\begin{align} \operatorname{Re}\operatorname{Tr}\bigl((\Lambda T_{sl}^{\varepsilon})^*(\Lambda T_{mn}^{\eta})\bigr) &=\frac{1}{2}(\lambda_s^2+\lambda_l^2)\,\delta_{sm}\delta_{ln}\delta_{\varepsilon\eta},\\ \operatorname{Re}\operatorname{Tr}\bigl((T_{sl}^{\varepsilon}\Lambda)^*(T_{mn}^{\eta}\Lambda)\bigr) &=\frac{1}{2}(\lambda_s^2+\lambda_l^2)\,\delta_{sm}\delta_{ln}\delta_{\varepsilon\eta},\\ \operatorname{Re}\operatorname{Tr}\bigl((\Lambda T_{sl}^{\varepsilon})^*(T_{mn}^{\eta}\Lambda)\bigr) &=\lambda_s\lambda_l\,\delta_{sm}\delta_{ln}\delta_{\varepsilon\eta}. \end{align}\] Therefore, for \(1\le p\le N-1\), \[\left\langle \mathbf{v}_{p,sl}^{\varepsilon},\mathbf{v}_{p,mn}^{\eta}\right\rangle_{\iota_\Sigma} =(\lambda_s^2+\lambda_l^2)\,\delta_{sm}\delta_{ln}\delta_{\varepsilon\eta}.\] For adjacent layers, with \(1\le p\le N-2\), the only common layer is \(p+1\), and the sign comes from the component \(-\Lambda T_{sl}^{\varepsilon}\) in \(\mathbf{v}_{p,sl}^{\varepsilon}\). Hence \[\left\langle \mathbf{v}_{p,sl}^{\varepsilon},\mathbf{v}_{p+1,mn}^{\eta}\right\rangle_{\iota_\Sigma} = -\operatorname{Re}\operatorname{Tr}\bigl((\Lambda T_{sl}^{\varepsilon})^*(T_{mn}^{\eta}\Lambda)\bigr) = -\lambda_s\lambda_l\,\delta_{sm}\delta_{ln}\delta_{\varepsilon\eta}.\] This proves 16 and shows that distinct pairs \((s,l)\) or distinct basis units are orthogonal.
For the imaginary diagonal vectors, \[\Lambda D_s^{\eta}=D_s^{\eta}\Lambda=\lambda_s\eta E_{ss}.\] Thus \[\begin{align} \operatorname{Re}\operatorname{Tr}\bigl((\Lambda D_s^{\eta})^*(\Lambda D_m^{\theta})\bigr) &=\lambda_s^2\delta_{sm}\delta_{\eta\theta},\\ \operatorname{Re}\operatorname{Tr}\bigl((D_s^{\eta}\Lambda)^*(D_m^{\theta}\Lambda)\bigr) &=\lambda_s^2\delta_{sm}\delta_{\eta\theta},\\ \operatorname{Re}\operatorname{Tr}\bigl((\Lambda D_s^{\eta})^*(D_m^{\theta}\Lambda)\bigr) &=\lambda_s^2\delta_{sm}\delta_{\eta\theta}. \end{align}\] Therefore, for \(1\le p\le N-1\), \[\left\langle \mathbf{w}_{p,s}^{\eta},\mathbf{w}_{p,m}^{\theta}\right\rangle_{\iota_\Sigma} =2\lambda_s^2\delta_{sm}\delta_{\eta\theta}.\] For adjacent layers, with \(1\le p\le N-2\), the only common layer is \(p+1\), and the sign again comes from the component \(-\Lambda D_s^{\eta}\) in \(\mathbf{w}_{p,s}^{\eta}\). Hence \[\left\langle \mathbf{w}_{p,s}^{\eta},\mathbf{w}_{p+1,m}^{\theta}\right\rangle_{\iota_\Sigma} = -\operatorname{Re}\operatorname{Tr}\bigl((\Lambda D_s^{\eta})^*(D_m^{\theta}\Lambda)\bigr) = -\lambda_s^2\delta_{sm}\delta_{\eta\theta}.\] This proves 17 . Every pairing between an off-diagonal vector and a diagonal vector vanishes because the relevant products of matrix units have zero trace. ◻
Introduce the standard tridiagonal matrix \[\label{eq:tridiagonal-L} L_{N-1}:= \begin{pmatrix} 2&-1&&&0\\ -1&2&-1&&\\ &\ddots&\ddots&\ddots&\\ &&-1&2&-1\\ 0&&&-1&2 \end{pmatrix}\in {M}_{N-1}(\mathbb{R}).\tag{18}\] For \(1\le s<l\le d\), define \[\label{eq:offdiag-block} H_{sl}:=(\lambda_s-\lambda_l)^2I_{N-1}+\lambda_s\lambda_lL_{N-1}.\tag{19}\] Then \(H_{sl}\) is the \((N-1)\times(N-1)\) tridiagonal matrix with diagonal entries \(\lambda_s^2+\lambda_l^2\) and nearest off-diagonal entries \(-\lambda_s\lambda_l\).
Let \(G_{\Sigma,\beta}\) be the matrix of inner products of \(\iota_\Sigma\) on \(T_{\mathbf{C}}\mathcal{O}_{\Sigma}^{\beta}\) in the ordered pushed-forward basis obtained as follows: group together the vectors \(\mathbf{v}_{p,sl}^{\varepsilon}\) with fixed \((s,l,\varepsilon)\) and order them by depth \(p=1,\dots,N-1\), and group together the vectors \(\mathbf{w}_{p,s}^{\eta}\) with fixed \((s,\eta)\) and again order them by depth.
Proposition 4. The matrix \(G_{\Sigma,\beta}\) is block diagonal. More precisely,
for each pair \(1\le s<l\le d\) and each \(\varepsilon\in B_\beta\), the corresponding block is \(H_{sl}\);
for each \(1\le s\le d\) and each \(\eta\in I_\beta\), the corresponding block is \(\lambda_s^2L_{N-1}\).
In particular, \[\label{eq:gram-determinant-factorization} \det G_{\Sigma,\beta} =\prod_{1\le s<l\le d}\det(H_{sl})^{\beta} \prod_{s=1}^d\det(\lambda_s^2L_{N-1})^{\beta-1}.\qquad{(4)}\]
Proof. By 4, vectors with different labels are orthogonal. For fixed \((s,l,\varepsilon)\), the vectors \[\mathbf{v}_{1,sl}^{\varepsilon},\dots,\mathbf{v}_{N-1,sl}^{\varepsilon}\] have the tridiagonal inner products described in 16 , so their block is exactly \(H_{sl}\). For fixed \((s,\eta)\), the vectors \[\mathbf{w}_{1,s}^{\eta},\dots,\mathbf{w}_{N-1,s}^{\eta}\] have the tridiagonal inner products described in 17 , so their block is \(\lambda_s^2L_{N-1}\). This proves the block decomposition.
The multiplicities are immediate: there are \(|B_\beta|=\beta\) off-diagonal copies for each pair \(1\le s<l\le d\) and \(|I_\beta|=\beta-1\) diagonal copies for each \(1\le s\le d\). Taking determinants of the blocks gives ?? . ◻
For \(\mathbb{F}=\mathbb{R}\), only one copy of each off-diagonal block \(H_{sl}\) appears. For \(\mathbb{F}=\mathbb{C}\) or \(\mathbb{H}\), the off-diagonal blocks occur with multiplicity \(\beta\), and the imaginary diagonal directions contribute the additional \(\beta-1\) copies of the blocks \(\lambda_s^2L_{N-1}\).
4 reduces the volume computation to the determinants of the tridiagonal blocks \(L_{N-1}\) and \(H_{sl}\). We now evaluate these determinants.
Lemma 5. For every \(n\ge 1\), \[\det(L_n)=n+1,\] where \(L_n\) is the \(n\times n\) tridiagonal matrix with diagonal entries \(2\) and nearest off-diagonal entries \(-1\).
Proof. Set \(\ell_n=\det(L_n)\). Expanding along the first row gives the recurrence \[\ell_n=2\ell_{n-1}-\ell_{n-2}, \qquad n\ge 3,\] with initial values \(\ell_1=2\) and \(\ell_2=3\). The sequence \(\ell_n=n+1\) satisfies the same recurrence and initial conditions, so \(\ell_n=n+1\) for all \(n\). ◻
Proposition 5. For \(1\le s<l\le d\), \[\label{eq:offdiag-determinant-sum} \det(H_{sl}) = \sum_{m=0}^{N-1} \lambda_s^{2(N-1-m)}\lambda_l^{2m}.\qquad{(5)}\] In particular, if \(\lambda_s\ne \lambda_l\), then \[\label{eq:offdiag-determinant} \det(H_{sl}) = \frac{\lambda_s^{2N}-\lambda_l^{2N}}{\lambda_s^2-\lambda_l^2} = \frac{\sigma_s^2-\sigma_l^2}{\sigma_s^{2/N}-\sigma_l^{2/N}}.\qquad{(6)}\] Moreover, \[\label{eq:diag-determinant} \det(\lambda_s^2L_{N-1})=N\lambda_s^{2(N-1)}=N\sigma_s^{2-2/N}.\qquad{(7)}\]
Proof. Equation ?? follows immediately from 5: \[\det(\lambda_s^2L_{N-1}) = \lambda_s^{2(N-1)}\det(L_{N-1}) = N\lambda_s^{2(N-1)}.\] Since \(\lambda_s^N=\sigma_s\), this is exactly ?? . For the block \(H_{sl}\), let \(D_n\) be the determinant of the \(n\times n\) tridiagonal Toeplitz matrix with diagonal entry \(\lambda_s^2+\lambda_l^2\) and nearest off-diagonal entry \(-\lambda_s\lambda_l\). Then \(D_{N-1}=\det(H_{sl})\). Expanding along the first row gives \[D_n = (\lambda_s^2+\lambda_l^2)D_{n-1} -\lambda_s^2\lambda_l^2D_{n-2}, \qquad n\ge 2,\] with initial conditions \(D_0=1\) and \(D_1=\lambda_s^2+\lambda_l^2\).
Set \[S_n=\sum_{m=0}^{n}\lambda_s^{2(n-m)}\lambda_l^{2m}.\] The sequence \(S_n\) has the same recurrence and the same initial values as \(D_n\). Hence \(D_n=S_n\) for all \(n\), and setting \(n=N-1\) gives ?? . If \(\lambda_s\ne\lambda_l\), the finite geometric sum gives \[\sum_{m=0}^{N-1} \lambda_s^{2(N-1-m)}\lambda_l^{2m} = \frac{\lambda_s^{2N}-\lambda_l^{2N}}{\lambda_s^2-\lambda_l^2}.\] Using \(\lambda_r^N=\sigma_r\) gives the second quotient in ?? . ◻
Corollary 2. Let \(G_{\Sigma,\beta}\) be the matrix of inner products of the induced metric on \(T_{\mathbf{C}}\mathcal{O}_{\Sigma}^{\beta}\) in the pushed-forward basis ordered as in 4. Assume \(\sigma_1>\cdots>\sigma_d>0\). Then \[\label{eq:total-gram-determinant-vandermonde} \det G_{\Sigma,\beta} = N^{(\beta-1)d} (\det\Sigma)^{2(\beta-1)(1-1/N)} \left( \frac{\operatorname{van}(\Sigma^2)}{\operatorname{van}(\Sigma^{2/N})} \right)^{\beta}.\tag{20}\]
Proof. Insert ?? and ?? into the block product formula ?? . The diagonal blocks contribute \[\prod_{s=1}^d \det(\lambda_s^2L_{N-1})^{\beta-1} = \prod_{s=1}^d \left(N\lambda_s^{2(N-1)}\right)^{\beta-1} = N^{(\beta-1)d}(\det\Lambda)^{2(\beta-1)(N-1)}.\] The off-diagonal blocks contribute \[\prod_{1\le s<l\le d}\det(H_{sl})^{\beta} = \prod_{1\le s<l\le d} \left( \frac{\sigma_s^2-\sigma_l^2}{\sigma_s^{2/N}-\sigma_l^{2/N}} \right)^{\beta}.\] Multiplying the two contributions gives \[\det G_{\Sigma,\beta} = N^{(\beta-1)d} (\det\Lambda)^{2(\beta-1)(N-1)} \prod_{1\le s<l\le d} \left( \frac{\sigma_s^2-\sigma_l^2}{\sigma_s^{2/N}-\sigma_l^{2/N}} \right)^{\beta}.\] Finally, \[(\det\Lambda)^{2(\beta-1)(N-1)} = (\det\Sigma)^{2(\beta-1)(1-1/N)}\] and \[\prod_{1\le s<l\le d} \left( \frac{\sigma_s^2-\sigma_l^2}{\sigma_s^{2/N}-\sigma_l^{2/N}} \right)^{\beta} = \left( \frac{\operatorname{van}(\Sigma^2)}{\operatorname{van}(\Sigma^{2/N})} \right)^{\beta}.\] Thus 20 follows. ◻
We now pull the induced metric on the orbit back to \(K_\beta^{N-1}\) by the orbit map.
Proposition 6. Assume \(X\in \mathrm{GL}_d(\mathbb{F})\) has distinct singular values. Let \(\iota_X\) denote the metric on \(\mathcal{O}_X^\beta\) induced by the ambient metric, and set \[\gamma_X:=\mathfrak{y}_X^*(\iota_X),\] where \(\mathfrak{y}_X\) is the orbit map of 3. Then \(\gamma_X\) is left-invariant on \(K_\beta^{N-1}\). Consequently, \[\label{eq:orbit-volume-from-gram} \operatorname{vol}(\mathcal{O}_X^\beta) = c_\beta^{N-1}\det(G_{X,\beta})^{1/2},\qquad{(8)}\] where \(G_{X,\beta}\) is the matrix of \(\gamma_X\) at the identity in any orthonormal basis of \(\mathfrak{k}_\beta^{N-1}\).
Proof. For \(\mathbf{V},\mathbf{U}\in K_\beta^{N-1}\), the orbit map satisfies \(\mathfrak{y}_X(\mathbf{V}\mathbf{U})=\mathbf{V}\cdot \mathfrak{y}_X(\mathbf{U}).\) The action of \(K_\beta^{N-1}\) on the orbit is by isometries for the induced metric. If \(\ell_{\mathbf{V}}\) denotes left translation by \(\mathbf{V}\), then \[(\ell_{\mathbf{V}})^*\gamma_X = (\mathfrak{y}_X\circ \ell_{\mathbf{V}})^*\iota_X = (\mathbf{V}\cdot \mathfrak{y}_X)^*\iota_X = \mathfrak{y}_X^*(\mathbf{V}\cdot)^*\iota_X = \mathfrak{y}_X^*\iota_X = \gamma_X.\] Thus \(\gamma_X\) is left-invariant.
The reference volume on \(K_\beta^{N-1}\) is the product volume induced by \(\langle A,B\rangle_{\mathfrak{k}} = \operatorname{Re}\operatorname{Tr}(A^*B)\) on each factor \(\mathfrak{k}_\beta\). Therefore \(\operatorname{vol}(K_\beta^{N-1})=c_\beta^{N-1}.\) When \(\mathbb{F}=\mathbb{R}\), this is the volume of all of \(O_d\). Choose an orthonormal basis of \(\mathfrak{k}_\beta^{N-1}\) for this product metric and extend it by left translation. In this frame, the reference metric has matrix \(I\), while \(\gamma_X\) has the constant matrix \(G_{X,\beta}\). Hence the Riemannian density of \(\gamma_X\) is \(\det(G_{X,\beta})^{1/2}\) times the reference density. Integrating over \(K_\beta^{N-1}\) gives ?? . ◻
Proof of 1. By the endpoint isometry, the entropy depends only on the singular values of \(X\). Thus \(S^\beta(X)=S^\beta(\Sigma).\) By 6 and 2, \[\operatorname{vol}(\mathcal{O}_\Sigma^\beta) = c_\beta^{N-1} N^{(\beta-1)d/2} (\det\Sigma)^{(\beta-1)(1-1/N)} \left( \frac{\operatorname{van}(\Sigma^2)}{\operatorname{van}(\Sigma^{2/N})} \right)^{\beta/2}.\] Taking logarithms gives \[S^\beta(X) = (N-1)\log c_\beta + \frac{\beta}{2} \log\left( \frac{\operatorname{van}(\Sigma^2)}{\operatorname{van}(\Sigma^{2/N})} \right) + (\beta-1) \left( \frac{d}{2}\log N + \log\left( \frac{\det\Sigma}{\det(\Sigma^{1/N})} \right) \right).\] Since \[\log\left( \frac{\det\Sigma}{\det(\Sigma^{1/N})} \right) = \left(1-\frac{1}{N}\right) \sum_{s=1}^d \log\sigma_s,\] this is equivalent to the sum formula. ◻
Lemma 6. Let \(X=U_N\Sigma U_0^*\) be a singular value decomposition over \(\mathbb{F}\). Assume \(\sigma_1>\cdots>\sigma_d>0\). Then \[\label{eq:det-real-AN-inverse} \det_{\mathbb{R}}(\mathcal{A}_{N,X}^{-1}) = N^{-\beta d} (\det\Sigma)^{-2\beta(1-1/N)} \left( \frac{\operatorname{van}(\Sigma^{2/N})}{\operatorname{van}(\Sigma^2)} \right)^{2\beta}.\tag{21}\]
Proof. Let \(u_s\) and \(v_l\) be the columns of \(U_N\) and \(U_0\). The vectors \[u_s\varepsilon v_l^*, \qquad 1\le s,l\le d,\quad \varepsilon\in B_\beta,\] form a real orthonormal basis of \(M_d(\mathbb{F})\). Indeed, \[\left\langle u_s\varepsilon v_l^*, u_m\eta v_n^* \right\rangle = \delta_{sm}\delta_{ln}\operatorname{Re}(\overline{\varepsilon}\eta) = \delta_{sm}\delta_{ln}\delta_{\varepsilon\eta}.\] For this basis, \[\mathcal{A}_{N,X}(u_s\varepsilon v_l^*) = \alpha_{sl}u_s\varepsilon v_l^*,\] where \[\alpha_{sl} := \sum_{p=1}^N \sigma_s^{2(N-p)/N} \sigma_l^{2(p-1)/N}.\] The eigenvalue does not depend on \(\varepsilon\), so each \(\alpha_{sl}\) has real multiplicity \(\beta\). For \(s=l\), \(\alpha_{ss}=N\sigma_s^{2-2/N}.\) For \(s\ne l\), the identity for a geometric series gives \[\alpha_{sl} = \frac{\sigma_s^2-\sigma_l^2}{\sigma_s^{2/N}-\sigma_l^{2/N}}.\] Hence \[\det_{\mathbb{R}}(\mathcal{A}_{N,X}^{-1}) = \prod_{s=1}^d \alpha_{ss}^{-\beta} \prod_{1\le s<l\le d}\alpha_{sl}^{-2\beta}.\] Substituting the formulas for \(\alpha_{ss}\) and \(\alpha_{sl}\) gives \[\det_{\mathbb{R}}(\mathcal{A}_{N,X}^{-1}) = N^{-\beta d} (\det\Sigma)^{-2\beta(1-1/N)} \prod_{1\le s<l\le d} \left( \frac{\sigma_s^{2/N}-\sigma_l^{2/N}}{\sigma_s^2-\sigma_l^2} \right)^{2\beta},\] which is 21 . ◻
Proof of 1. By the endpoint isometry, it suffices to use the diagonal representative \(\Sigma\). Combining ?? with 20 gives \[\operatorname{vol}(\mathcal{O}_X^\beta) = c_\beta^{N-1} N^{(\beta-1)d/2} (\det\Sigma)^{(\beta-1)(1-1/N)} \left( \frac{\operatorname{van}(\Sigma^2)}{\operatorname{van}(\Sigma^{2/N})} \right)^{\beta/2}.\] Taking the fourth root of 21 gives \[\det_{\mathbb{R}}(\mathcal{A}_{N,X}^{-1})^{1/4} = N^{-\beta d/4} (\det\Sigma)^{-\beta(1-1/N)/2} \left( \frac{\operatorname{van}(\Sigma^{2/N})}{\operatorname{van}(\Sigma^2)} \right)^{\beta/2}.\] The Vandermonde factors cancel. The powers of \(N\) combine as \[\frac{(\beta-1)d}{2}-\frac{\beta d}{4} = \frac{(\beta-2)d}{4},\] and the powers of \(\det\Sigma\) combine as \[(\beta-1)\left(1-\frac{1}{N}\right) - \frac{\beta}{2}\left(1-\frac{1}{N}\right) = \frac{\beta-2}{2}\left(1-\frac{1}{N}\right).\] Since \[(\det\Sigma)^{1-1/N} = \frac{\det\Sigma}{\det(\Sigma^{1/N})},\] we obtain \[\operatorname{vol}(\mathcal{O}_X^\beta) \det_{\mathbb{R}}(\mathcal{A}_{N,X}^{-1})^{1/4} = c_\beta^{N-1} N^{(\beta-2)d/4} \left( \frac{\det\Sigma}{\det(\Sigma^{1/N})} \right)^{(\beta-2)/2}.\] Taking logarithms gives the equivalent entropy identity, because \(\mathcal{A}_{N,X}\) is positive. ◻
The authors are grateful to Govind Menon and Tianmin Yu for helpful discussions and insightful remarks that have improved this work. This work was partially supported by NSF grant DMS 2407055.