April 28, 2026
We develop a spectral theory for the continuous- and discrete-time Blahut–Arimoto (BA) dynamics centred on a single operator, the relaxation kernel \(\mathcal{G}= \mathbb{E}_p[K^*_X \otimes K^*_X]\). Five main results are established. (i) Along the continuous-time BA flow, the free energy satisfies the exact \(\chi^2\)-dissipation identity \(\dot{F}_\beta = -\mathcal{D}(q)\), where \(\mathcal{D}(q) = \chi^2(\mathcal{T}q \| q)\) is the Pearson \(\chi^2\)-divergence. (ii) The operator \(\mathcal{G}\) admits a threefold identity: it is simultaneously the Gram matrix of the equilibrium Gibbs kernels, the linearised generator of the BA vector field, and the Fisher–Rao Hessian of the free energy at the fixed point. (iii) For the discrete BA iteration, the one-step Lyapunov dissipation decomposes spectrally as \(\Delta\mathcal{L}^{(2)} = \sum_i c_i^2\, d(\lambda_i)\), where \(\lambda_i\) are eigenvalues of \(\mathcal{G}\) and \(d(\lambda) = -\lambda + \tfrac{3}{2}\lambda^2 - \tfrac{1}{2}\lambda^3\). This formula reveals a double bottleneck: both \(\lambda\approx 0\) (slow-contracting directions) and \(\lambda\approx 1\) (over-contracted, near-zero Lyapunov weight) yield \(d(\lambda)\approx 0\), while optimal dissipation occurs at \(\lambda\approx 0.423\). (iv) Global convergence follows from a two-stage mechanism: \(\chi^2\)- dissipation drives finite-time entry into a local neighbourhood, after which the spectral gap \(\lambda_*= \lambda_{\min}(\mathcal{G}|_T)\) governs exponential contraction. A self-contained proof with explicit entry time and convergence factor is given in Appendix sec:sec:app:twostage?. (v) The KL convergence factor is made explicit: \(D_{\mathrm{KL}}(q^*\|q_{n+1}) \le (1-\lambda_*)^2\,D_{\mathrm{KL}}(q^*\|q_n) + O(\|v_n\|_*^3)\), where \(v_n = q_n - q^*\in T\) denotes the \(n\)-th iterate’s deviation from the fixed point. The per-iteration fractional improvement \(\gamma := 1-(1-\lambda_*)^2 = \lambda_*(2-\lambda_*)\) is computable from the channel and temperature, and equality holds asymptotically along the slowest eigen-direction. For Gaussian sources with Gaussian initial condition, \(\lambda_*= 1/(2\beta\sigma^2)\) and the Jacobian is diagonalised by Hermite polynomials. The spectral dissipation formula \(d(\lambda)\) makes explicit the per-step Lyapunov structure in the local regime, complementing the global convergence theory of Hayashi [1]–[3] with a constructive, computable rate.
Blahut–Arimoto algorithm, relaxation kernel, \(\chi^2\) dissipation, spectral gap, exponential convergence, KL divergence, Gaussian source.
The Blahut–Arimoto (BA) algorithm [4], [5] computes rate-distortion functions and channel capacities via an iterative Gibbs-type update. Classical analysis [4]–[6] proves convergence via monotone decrease of the free energy, without identifying the underlying dissipation mechanism. Nakagawa et al [7] established sharp \(O(1/n)\) rates for the discrete iteration. Hayashi [1]–[3] reformulated each BA step as a Bregman EM update, obtaining global convergence guarantees within an elegant information-geometric framework.
Despite this progress, two questions remain open. First, existing convergence constants are non-constructive: they depend on abstract geometric quantities and cannot be directly computed from the channel and temperature parameter \(\beta\). Second, the precise spectral mechanism governing the per-step Lyapunov decrease—and in particular its dependence on the non-idempotence of the underlying operator—has not been identified.
The present paper closes both gaps around a single operator:
\[\mathcal{G}= \mathbb{E}_p\bigl[K^*_X \otimes K^*_X\bigr]\]
the relaxation kernel, constructed from the equilibrium Gibbs conditionals \(K^*_x(y) = e^{-\beta d(x,y)}/Z^*_x\). This operator governs, within one unified framework, the continuous-time entropy production, the linearised discrete dynamics, the free-energy curvature, and the per-step Lyapunov dissipation.
Exact \(\chi^2\)-dissipation identity (Section 3). Along the continuous-time BA flow \(\dot{q} = \mathcal{T}(q)-q\), \[\frac{d}{dt}F_\beta(q_t) = -\mathcal{D}(q_t), \qquad \mathcal{D}(q) = \chi^2(\mathcal{T}q \| q) = \sum_y \frac{(\mathcal{T}q(y)-q(y))^2}{q(y)}.\] This is an exact equality, not merely a bound.
Threefold identity of \(\mathcal{G}\) (Section 4). At an interior fixed point \(q^*\), \[\begin{align} \mathcal{G}&= \mathbb{E}_p[K^*_X \otimes K^*_X] && \text{(Gram / correlation matrix)},\\ DV(q^*) &= -\mathcal{G}&& \text{(linearised generator)},\\ \nabla^2_{\mathrm{FR}}F_\beta(q^*) &= \mathcal{G}&& \text{(Fisher--Rao Hessian)}. \end{align}\]
Exact one-step Lyapunov dissipation (Section 5). For the discrete BA iteration, \[\label{eq:dlambda95intro} \Delta\mathcal{L}^{(2)} = \sum_i c_i^2\, d(\lambda_i), \qquad d(\lambda)= -\lambda + \tfrac{3}{2}\lambda^2 - \tfrac{1}{2}\lambda^3,\tag{1}\] where \(\lambda_i\) are eigenvalues of \(\mathcal{G}\) and \(c_i = \langle h,\, e_i\rangle_{q^*}\). The non-idempotence of \(\mathcal{G}\) (i.e.\(\mathcal{G}^2 \neq \mathcal{G}\)) is what drives \(d(\lambda)<0\) for all \(\lambda\in(0,1)\); without non-idempotence the dissipation would vanish identically.
Two-stage global convergence (Section 7). \(\chi^2\)-dissipation drives finite-time entry into a neighbourhood of \(q^*\); thereafter, exponential contraction at rate \(\lambda_*\) takes over.
Explicit KL convergence factor (Section 8). For the discrete iterates \(q_{n+1} = \mathcal{T}(q_n)\), writing \(v_n = q_n - q^*\in T\): \[D_{\mathrm{KL}}(q^*\|q_{n+1}) \le (1-\lambda_*)^2\,D_{\mathrm{KL}}(q^*\|q_n) + O(\|v_n\|_*^3),\] with explicit \(\gamma = \lambda_*(2-\lambda_*)\), computable for every non-degenerate fixed point. Equality holds asymptotically when \(v_n\) is aligned with the \(\lambda_*\)-eigendirection of \(\mathcal{G}\).
Hayashi [1]–[3] developed an influential framework in which each BA step is viewed as an alternating \(e\)/\(m\)-projection in a Bregman divergence system. This yields global monotone decrease of the free energy, an \(O(1/n)\) global convergence rate, and, locally, exponential convergence once the iterate enters a neighbourhood of the fixed point—though the exponential rate constant is not expressed in closed form. The present paper addresses complementary questions: it identifies the relaxation kernel \(\mathcal{G}\) as the central spectral object, derives the exact per-step Lyapunov dissipation formula \(d(\lambda)\), and makes the local convergence factor \(\gamma = \lambda_*(2-\lambda_*)\) explicit and computable from the channel and temperature. The two frameworks are therefore complementary: Hayashi’s provides global guarantees from any starting point; the spectral theory developed here provides fine-grained quantitative information in the local phase.
Csiszár [6] characterised each BA step as an alternating \(I\)-projection, providing a discrete geometric picture. Nakagawa et al [7] proved sharp \(O(1/n)\) rates. Our continuous-time framework complements these results by providing the exact entropy production rate and an explicit local relaxation time.
Beretta and Pelillo [8] proposed a continuous vector flow for channel capacity and proved qualitative exponential convergence via a replicator structure. The present analysis refines their picture by identifying the explicit spectral relaxation rate \(\lambda_*\).
Ramakrishnan et al [9] proved that \(D_{\mathrm{KL}}(q^*\|q_{n+1}) \le D_{\mathrm{KL}}(q^*\|q_n)\) for the quantum BA algorithm; our Corollary 4 upgrades this qualitative monotonicity to a quantitative formula with explicit factor \((1-\lambda_*)^2\).
Section 2 defines the BA flow and weighted Hilbert space. Section 3 proves the \(\chi^2\)-dissipation identity. Section 4 establishes the threefold identity of \(\mathcal{G}\). Section 5 derives the exact one-step dissipation, analyses \(d(\lambda)\), and reports numerical validation. Section 6 proves local exponential convergence. Section 7 assembles the two-stage global theorem, including the uniform moment bound required for the energy argument. Section 8 gives the explicit KL convergence factor. Section 9 specialises to Gaussian sources. Section 10 gives exact solutions for discrete models. Section 11 concludes. Appendix 12 provides the complete proof of two-stage global exponential convergence with explicit entry time and convergence factor. Appendix 13 gives the full proof of the Gaussian attractor theorem, including moment control, Hermite spectral decay, weak convergence, and upgrade to total variation.
This section establishes the objects that appear throughout the paper. We define the BA operator and its continuous-time flow, identify the class of interior fixed points and the dual identity they satisfy, and introduce the weighted Hilbert space in which all subsequent analysis takes place. Readers familiar with the BA algorithm may wish to skim Sections 2.1–2.2 and focus on Section 2.3, which introduces the conditional correlation operator \(\mathcal{G}\) whose threefold identity is the central object of Section 4.
Let \(\mathcal{X},\mathcal{Y}\) be finite sets, \(|\mathcal{X}|=M\), \(|\mathcal{Y}|=N\), \(p\) a strictly positive probability distribution on \(\mathcal{X}\), \(d:\mathcal{X}\times\mathcal{Y}\to\mathbb{R}\) a distortion function, and \(\beta>0\) an inverse temperature. Throughout, \(\Delta(\mathcal{Y})\) denotes the probability simplex on \(\mathcal{Y}\) and \(\operatorname{int}\Delta(\mathcal{Y})\) its relative interior (all coordinates positive). The BA algorithm seeks to minimise the free energy \(F_\beta:\Delta(\mathcal{Y})\to\mathbb{R}\) defined in Section 3 over \(q\in\Delta(\mathcal{Y})\); its iterations act on the output marginal \(q\in\Delta(\mathcal{Y})\), while the joint channel is implicitly optimised at each step via the Gibbs kernel.
Definition 1 (Gibbs kernel and BA operator). Given \(q\in\operatorname{int}\Delta(\mathcal{Y})\) and \(x\in\mathcal{X}\), the Gibbs kernel at \(q\)* is the probability distribution on \(\mathcal{Y}\) \[\label{eq:Kq} K_q(x,y) = \frac{e^{-\beta d(x,y)}\,q(y)}{Z_q(x)}, \qquad Z_q(x) = \sum_{y'\in\mathcal{Y}} q(y')\,e^{-\beta d(x,y')},\tag{2}\] where \(Z_q(x)>0\) is the normalising partition function. For each fixed \(x\), \(K_q(x,\cdot)\in\Delta(\mathcal{Y})\) is the Gibbs posterior over \(\mathcal{Y}\) given source symbol \(x\) and current marginal \(q\).*
The Blahut–Arimoto operator \(\mathcal{T}:\operatorname{int}\Delta(\mathcal{Y})\to\operatorname{int}\Delta(\mathcal{Y})\) is the \(p\)-mixture of these Gibbs posteriors: \[\label{eq:BAop} (\mathcal{T}q)(y) = \sum_{x\in\mathcal{X}} p(x)\,K_q(x,y) = \sum_{x\in\mathcal{X}} p(x)\, \frac{e^{-\beta d(x,y)}\,q(y)}{Z_q(x)}, \quad y\in\mathcal{Y}.\qquad{(1)}\] One verifies directly that \(\sum_y(\mathcal{T}q)(y)=1\) and \((\mathcal{T}q)(y)>0\), so \(\mathcal{T}\) maps \(\operatorname{int}\Delta(\mathcal{Y})\) into itself.
The classical BA algorithm is the discrete iteration \(q_{n+1} = \mathcal{T}(q_n)\). Its fixed points \(\mathcal{T}(q^*) = q^*\) are precisely the KKT stationarity conditions for the rate-distortion optimisation problem \(\min_{q\in\Delta(\mathcal{Y})}F_\beta(q)\): setting the Lagrangian gradient to zero and eliminating the multiplier yields exactly the equation \((\mathcal{T}q)(y) = q(y)\) for all \(y\) [4], [10]. Thus the fixed-point condition is not an auxiliary construct but the optimality condition of the underlying variational problem.
The continuous-time flow studied in this paper is the natural continuous relaxation of the discrete iteration: \[\label{eq:BAflow} \dot{q}_t = \mathcal{T}(q_t) - q_t, \qquad q_0\in\operatorname{int}\Delta(\mathcal{Y}).\tag{3}\] This ODE has the same fixed points as the discrete map and can be understood as the Euler discretisation \(q_{n+1} - q_n = \mathcal{T}(q_n) - q_n\) with unit step size, or equivalently as the replicator-type gradient flow of \(F_\beta\) with respect to the Fisher–Rao metric on \(\operatorname{int}\Delta(\mathcal{Y})\) [8]. The continuous-time formulation is adopted here because it admits exact differential-equation tools—in particular the \(\chi^2\)-dissipation identity of Section 3—that are not directly available for the discrete iteration. All qualitative convergence conclusions transfer back to the discrete map via the free-energy monotonicity \(F_\beta(q_{n+1})\le F_\beta(q_n)\), which holds for both flows [4].
The simplex \(\Delta(\mathcal{Y})\) is forward-invariant under 3 : \(\sum_y\dot{q}_t(y) = \sum_y(\mathcal{T}q_t)(y) - \sum_y q_t(y) = 0\) preserves the sum-to-one constraint, and \(q_t(y)=0\) implies \((\mathcal{T}q_t)(y)\ge 0\), so no coordinate can become negative. The interior \(\operatorname{int}\Delta(\mathcal{Y})\) is likewise forward-invariant since \((\mathcal{T}q)(y)>0\) for all \(q\in\operatorname{int}\Delta(\mathcal{Y})\).
Definition 2 (Interior fixed point). An interior fixed point* is \(q^*\in\operatorname{int}\Delta(\mathcal{Y})\) satisfying \(\mathcal{T}(q^*)=q^*\). At such a point, write \(K^*_x := K_{q^*}(x,\cdot)\in\Delta(\mathcal{Y})\) for the equilibrium Gibbs kernel at source symbol \(x\), i.e. \[K^*_x(y) = K_{q^*}(x,y) = \frac{e^{-\beta d(x,y)}\,q^*(y)}{Z^*_x}, \qquad Z^*_x := Z_{q^*}(x) = \sum_{y'\in\mathcal{Y}} q^*(y')\,e^{-\beta d(x,y')}.\] For each \(x\in\mathcal{X}\), \(K^*_x\in\Delta(\mathcal{Y})\) is a probability distribution on \(\mathcal{Y}\) that depends on \(q^*\), \(\beta\), and \(d\).*
Lemma 1 (Dual fixed-point identity). At an interior fixed point \(q^*\), \[\label{eq:dualFP} \sum_x p(x)\,\frac{K^*_x(y)}{q^*(y)} = 1,\qquad\forall y\in\mathcal{Y}.\qquad{(2)}\]
Proof. From \(q^* = \mathcal{T}(q^*)\): \(q^*(y) = \sum_x p(x)K_q(x,y)\,q^*(y)/q^*(y) = q^*(y)\sum_x p(x)K^*_x(y)/q^*(y)\). Cancel \(q^*(y)>0\). ◻
All perturbation analysis takes place in the tangent space to the simplex. Let \[T = \bigl\{u\in\mathbb{R}^\mathcal{Y}: \textstyle\sum_{y\in\mathcal{Y}} u(y)=0\bigr\}\] be the \((N-1)\)-dimensional subspace of zero-sum vectors, which is the tangent space to \(\Delta(\mathcal{Y})\) at any interior point.
Definition 3 (Fisher–Rao inner product). For \(q\in\operatorname{int}\Delta(\mathcal{Y})\), equip \(T\) with the Fisher–Rao inner product* \[\label{eq:FR95inner} \langle u,v\rangle_q := \sum_{y\in\mathcal{Y}} \frac{u(y)\,v(y)}{q(y)}, \quad u,v\in T.\tag{4}\] Write \(\mathcal{H}_{q}\) for the Hilbert space \((T,\langle\cdot,\cdot\rangle_q)\) and \(\|u\|_q := \langle u,u\rangle_q^{1/2}\) for the induced norm. At the fixed point \(q^*\), abbreviate \(\langle\cdot,\cdot\rangle_* = \langle\cdot,\cdot\rangle_{q^*}\) and \(\|\cdot\|_* = \|\cdot\|_{q^*}\).*
We now introduce the three auxiliary quantities that enter the definition of \(\mathcal{G}\). Fix an interior fixed point \(q^*\) and a tangent perturbation \(h\in T\). Define:
the log-perturbation \(\phi:\mathcal{Y}\to\mathbb{R}\) by \(\phi(y) = h(y)/q^*(y)\);
the conditional mean \(m:\mathcal{X}\to\mathbb{R}\) by \(m(x) = \mathbb{E}_{K^*_x}[\phi] = \sum_{y\in\mathcal{Y}} K^*_x(y)\,\phi(y)\);
the conditional residual \(r:\mathcal{Y}\times\mathcal{X}\to\mathbb{R}\) by \(r(y,x) = \phi(y) - m(x)\), so that \(\mathbb{E}_{K^*_x}[r(\cdot,x)]=0\) for every \(x\).
The quantity \(m(x)\) is the component of \(\phi\) visible to source symbol \(x\) through the equilibrium channel \(K^*_x\), and \(r(\cdot,x)\) is the residual invisible to \(x\).
Definition 4 (Relaxation kernel). The relaxation kernel \(\mathcal{G}:\mathcal{H}_{q^*}\to\mathcal{H}_{q^*}\) is the self-adjoint operator on \((T,\langle\cdot,\cdot\rangle_*)\) defined by \[\label{eq:PiPi} (\mathcal{G}h)(y) := q^*(y)\sum_{x\in\mathcal{X}} w_x(y)\,m(x) = q^*(y)\sum_{x\in\mathcal{X}} \frac{p(x)K^*_x(y)}{q^*(y)}\, \sum_{y'\in\mathcal{Y}} K^*_x(y')\,\frac{h(y')}{q^*(y')},\qquad{(3)}\] where \(w_x(y) = p(x)K^*_x(y)/q^*(y)\) is the posterior weight of source symbol \(x\) given output \(y\) under the equilibrium distribution.
Expanding the right-hand side of ?? , the \((y,y')\)-entry of \(\mathcal{G}\) as a matrix acting on \(\mathbb{R}^\mathcal{Y}\) is \[\label{eq:G95matrix} \mathcal{G}_{yy'} = \sum_{x\in\mathcal{X}} p(x)\,K^*_x(y)\,K^*_x(y'),\tag{5}\] which is the covariance of the random vector \(K^*_X\in\mathbb{R}^\mathcal{Y}\) under the source distribution \(p\). Equivalently, \(\mathcal{G}= \mathbb{E}_p[K^*_X \otimes K^*_X]\) where \(K^*_X\) is viewed as a \(\mathbb{R}^\mathcal{Y}\)-valued random variable. This representation immediately shows that \(\mathcal{G}\) is symmetric (\(\mathcal{G}_{yy'} = \mathcal{G}_{y'y}\)) and positive semidefinite (\(u^\top\mathcal{G}u = \sum_x p(x)\langle K^*_x, u\rangle^2\ge 0\)).
In the language of alternating projections [6], \(\mathcal{G}= \Pi^*\Pi\) where \(\Pi: \mathcal{H}_{q^*}\to \ell^2(\mathcal{X}, p)\) is the conditional expectation operator \((\Pi h)(x) = m(x) = \mathbb{E}_{K^*_x}[\phi]\), mapping tangent vectors \(h\in T\) to functions on \(\mathcal{X}\), and \(\Pi^*\) is its adjoint with respect to the inner products \(\langle\cdot,\cdot\rangle_*\) and \(\langle\cdot,\cdot\rangle_p\). The eigenvalues of \(\mathcal{G}\) therefore lie in \([0,1]\), and \(\mathcal{G}|_T\) has the same eigenvalues as \(\Pi^*\Pi|_T\).
Assumption 1 (Regularity). (i) \(q_t\in\operatorname{int}\Delta(\mathcal{Y})\) for all \(t\ge0\). (ii) \(d(x,y)\le C(1+y^2)\) for some \(C>0\). (iii) \(\mathcal{T}\) is Fréchet differentiable on \(\operatorname{int}\Delta(\mathcal{Y})\).
For finite \(\mathcal{Y}\) these are automatic.
The classical monotonicity \(F_\beta(q_{k+1})\le F_\beta(q_k)\) is well known, but the mechanism behind it has not been identified precisely. This section shows that the rate of decrease of the free energy along the continuous-time flow equals, exactly and at every instant, the Pearson \(\chi^2\)-divergence between \(\mathcal{T}q_t\) and \(q_t\). The derivation is short and rests on the envelope theorem; the result is stronger than a bound. We also record a corollary that will serve as the energy estimate in the global phase of the two-stage convergence argument (Section 7).
The BA free energy is \[\label{eq:F} F_\beta(q) = \sum_x p(x)\log Z_q(x) + \frac{1}{\beta}\sum_y q(y)\log q(y).\tag{6}\] This equals \(\min_P \mathcal{F}_\beta(P;q)\) where \(\mathcal{F}_\beta(P;q) = I_P(X;\hat{X}) + \beta\mathbb{E}_P[d(X,\hat{X})]\) is the standard rate-distortion auxiliary functional, and the minimum is achieved at \(P(\hat{x}|x) = K_q(x,\hat{x})\).
Lemma 2 (Envelope theorem). For every \(q\in\operatorname{int}\Delta(\mathcal{Y})\) and \(h\in T\), \[\label{eq:envelope} \delta F_\beta(q)[h] = -\sum_y \frac{\mathcal{T}q(y)}{q(y)}\,h(y).\qquad{(4)}\]
Proof. Since \(\mathcal{T}(q)\) is the unconstrained interior minimiser of \(\mathcal{F}_\beta(\cdot;q)\), the envelope theorem ([11]) differentiates only the explicit \(q\)-dependence of \(\mathcal{F}_\beta(p;q)\), which is \(-D_{\mathrm{KL}}(p\|q)\). Its gradient is \(-p(y)/q(y)\); substituting \(p=\mathcal{T}(q)\) gives the claim. ◻
Definition 5 (Dissipation functional). \[\mathcal{D}(q) := \chi^2(\mathcal{T}q\|q) = \sum_y \frac{(\mathcal{T}q(y)-q(y))^2}{q(y)}.\]
Theorem 1 (Exact \(\chi^2\)-dissipation identity). Along any trajectory of the BA flow 3 , \[\label{eq:chi2} \frac{d}{dt}F_\beta(q_t) = -\mathcal{D}(q_t).\qquad{(5)}\]
Proof. By the envelope lemma and the flow equation \(\dot{q}_t = \mathcal{T}(q_t)-q_t\): \[\begin{align} \frac{d}{dt}F_\beta(q_t) &= \sum_y \left(-\frac{\mathcal{T}q_t(y)}{q_t(y)}\right) (\mathcal{T}q_t(y) - q_t(y))\\ &= -\sum_y \frac{(\mathcal{T}q_t(y))^2}{q_t(y)} + \sum_y \mathcal{T}q_t(y)\\ &= -\sum_y \frac{(\mathcal{T}q_t(y)-q_t(y))^2}{q_t(y)} - \sum_y q_t(y) + 1 = -\mathcal{D}(q_t), \end{align}\] using \(\sum_y \mathcal{T}q_t(y)=1\) and \(\sum_y q_t(y)=1\). ◻
Remark 2 (Thermodynamic interpretation). Equation ?? is an exact non-equilibrium entropy production law. The classical monotonicity \(F_\beta(q_{k+1})\le F_\beta(q_k)\) for the discrete iteration is the integral form of this identity. The \(\chi^2\)-divergence plays the role of an entropy production rate, analogous to the Fisher information in Fokker–Planck equations [12].
Corollary 1 (Lyapunov property and residual bound). \(F_\beta\) is a strict Lyapunov function with \(\dot{F}_\beta\le 0\), equality iff \(q_t\in\mathcal{F} := \{q:\mathcal{T}q=q\}\). Moreover, \(\mathcal{D}(q)\ge\tfrac{1}{2}\|\mathcal{T}q-q\|_1^2\).
Proof. Non-negativity of \(\chi^2\) and the Cauchy–Schwarz inequality \(\chi^2(r\|q)\ge\tfrac12\|r-q\|_1^2\) [13]. ◻
With the dissipation mechanism in hand, we turn to the operator that governs it. The relaxation kernel \(\mathcal{G}= \mathbb{E}_p[K^*_X\otimes K^*_X]\) was introduced in Definition 4 as a purely algebraic object. This section shows that it in fact plays three distinct and individually natural roles: it is simultaneously the Gram matrix of equilibrium correlations, the linearised generator of the BA flow, and the Fisher–Rao Hessian of the free energy. The identification of these three roles in a single operator is what makes \(\mathcal{G}\) the natural centre of the spectral theory developed in the subsequent sections.
Define the BA vector field \(V:\operatorname{int}\Delta(\mathcal{Y})\to T\) by \(V(q) = \mathcal{T}(q) - q\), so that the BA flow 3 reads \(\dot{q}_t = V(q_t)\). Since \(\mathcal{T}\) maps \(\operatorname{int}\Delta(\mathcal{Y})\) into itself and \(\sum_y(\mathcal{T}q)(y) = \sum_y q(y) = 1\), the difference \(V(q)\) lies in the tangent space \(T\) for every \(q\). The Fréchet derivative \(DV(q^*): T\to T\) is a linear map from \(T\) to itself.
Theorem 3 (Linearised BA vector field). Let \(q^*\) be an interior fixed point. The Fréchet derivative of \(V = \mathcal{T}- \mathrm{id}\) at \(q^*\), viewed as a linear map \(T\to T\), is \[\label{eq:lin} DV(q^*) = -\mathcal{G},\qquad{(6)}\] where \(\mathcal{G}:\mathcal{H}_{q^*}\to\mathcal{H}_{q^*}\) is the relaxation kernel of Definition 4. Equivalently, the Jacobian of the BA map \(\mathcal{T}\) at \(q^*\) satisfies \[\label{eq:DT} D\mathcal{T}(q^*) = I - \mathcal{G}\quad \text{as linear maps } T\to T,\qquad{(7)}\] where \(I\) denotes the identity on \(T\).
Proof. Fix \(h\in T\) and write \(q^\varepsilon = q^*+\varepsilon h\). Expand the partition function: \[\begin{align} Z_{q^\varepsilon}(x) =& Z^*_x + \varepsilon\sum_{y'\in\mathcal{Y}} e^{-\beta d(x,y')}h(y') + O(\varepsilon^2)\\ =& Z^*_x\Bigl(1 + \varepsilon\sum_{y'} K^*_x(y')\,\frac{h(y')}{q^*(y')}\cdot\frac{q^*(y')}{1} \cdot \frac{1}{q^*(y')} + O(\varepsilon^2)\Bigr). \end{align}\] More precisely, \(Z_{q^\varepsilon}(x) = Z^*_x(1 + \varepsilon m(x) + O(\varepsilon^2))\) where \(m(x) = \sum_{y'} K^*_x(y')h(y')/q^*(y') = \mathbb{E}_{K^*_x}[\phi]\) is the conditional mean from Definition 4. Hence \[\frac{1}{Z_{q^\varepsilon}(x)} = \frac{1}{Z^*_x}\bigl(1 - \varepsilon m(x) + O(\varepsilon^2)\bigr).\] Substituting into the BA operator formula ?? : \[\begin{align} (\mathcal{T}q^\varepsilon)(y) &= \sum_x p(x)\,\frac{e^{-\beta d(x,y)}\,(q^*(y)+\varepsilon h(y))}{Z_{q^\varepsilon}(x)}\\ &= \sum_x p(x)\,K^*_x(y)\,(1+\varepsilon\phi(y))\,(1-\varepsilon m(x)) + O(\varepsilon^2)\\ &= q^*(y) + \varepsilon h(y) - \varepsilon\sum_x p(x)\,K^*_x(y)\,m(x) + O(\varepsilon^2), \end{align}\] using \(\sum_x p(x)K^*_x(y) = q^*(y)\) (the dual identity ?? ) and \(K^*_x(y)\,\phi(y) = K^*_x(y)\,h(y)/q^*(y)\). Comparing with \((\mathcal{T}q^\varepsilon)(y) = q^*(y) + \varepsilon(D\mathcal{T}(q^*)h)(y) + O(\varepsilon^2)\) and using the matrix form 5 : \[(D\mathcal{T}(q^*)h)(y) = h(y) - \sum_x p(x)\,K^*_x(y)\,m(x) = h(y) - (\mathcal{G}h)(y),\] so \(D\mathcal{T}(q^*) = I - \mathcal{G}\) and \(DV(q^*) = -\mathcal{G}\) as claimed. ◻
Remark 4. The operator \(\mathcal{G}\) is manifestly symmetric and positive semidefinite: \(u^\top\mathcal{G}u = \sum_x p(x)\langle K^*_x,u\rangle^2\ge 0\). Its null space on \(T\) consists of directions \(h\) invisible to all equilibrium kernels; non-degeneracy (\(\lambda_*>0\)) means no such directions exist.
The representation \(\mathcal{G}= \mathbb{E}_p[K^*_X\otimes K^*_X]\) in 5 identifies \(\mathcal{G}\) as the second moment (Gram) matrix of the \(\mathbb{R}^\mathcal{Y}\)-valued random vector \(K^*_X\) under the source distribution \(p\). Here \(K^*_X\otimes K^*_X\) denotes the \(N\times N\) matrix with \((y,y')\)-entry \(K^*_X(y)\,K^*_X(y')\), and the expectation is taken coordinatewise. This representation shows that \(\mathcal{G}\) encodes the equilibrium correlation structure among output symbols: \(\mathcal{G}_{yy'}\) is the covariance (under \(p\)) between the Gibbs probability of outputting \(y\) and the Gibbs probability of outputting \(y'\), when the source symbol is drawn from \(p\).
The Fisher–Rao Hessian of \(F_\beta\) at \(q\) is the symmetric bilinear form \[\nabla^2_{\mathrm{FR}}F_\beta(q)[h_1,h_2] := \left.\frac{d^2}{d\varepsilon\,d\eta}\right|_{\varepsilon=\eta=0} F_\beta(q+\varepsilon h_1 + \eta h_2).\]
Theorem 5 (Hessian identification). At an interior fixed point \(q^*\), \[\label{eq:Hessian} \nabla^2_{\mathrm{FR}}F_\beta(q^*)[h_1,h_2] = \langle h_1,\, \mathcal{G}h_2\rangle_{q^*}, \quad \forall\,h_1,h_2\in T.\qquad{(8)}\]
Proof. From the envelope lemma, \(\delta F_\beta(q)[h] = -\sum_y (\mathcal{T}q(y)/q(y))h(y)\). Differentiating in direction \(h_2\) and using \(D\mathcal{T}(q^*)=I-\mathcal{G}\): \[\begin{align} \nabla^2_{\mathrm{FR}}F_\beta(q^*)[h_1,h_2] &= -\sum_y \left( \frac{((I-\mathcal{G})h_2)(y)}{q^*(y)} - \frac{q^*(y)}{(q^*(y))^2}h_2(y)\right)h_1(y)\\ &= -\langle h_1,\, (I-\mathcal{G})h_2\rangle_{q^*} + \langle h_1,\, h_2\rangle_{q^*} = \langle h_1,\, \mathcal{G}h_2\rangle_{q^*}. \end{align}\] The second equality uses \(\mathcal{T}q^* = q^*\) so the error term vanishes. ◻
Theorem 5 completes the threefold identity: \(\mathcal{G}\) is (1) Gram matrix of equilibrium correlations, (2) linearised generator of the BA flow, and (3) Fisher–Rao Hessian of the free energy.
Corollary 2 (Local quadratic equivalence). There exists a neighbourhood \(\mathcal{U}\) of \(q^*\) and constants \(\kappa_1,\kappa_2>0\) such that for all \(q\in\mathcal{U}\): \[\label{eq:quad95equiv} \kappa_1\,\|q-q^*\|_*^2 \;\le\; F_\beta(q)-F_\beta(q^*) \;\le\; \kappa_2\,\|q-q^*\|_*^2,\qquad{(9)}\] with constants \(\kappa_1 = \lambda_*/2\) and \(\kappa_2 = \|\mathcal{G}\|/2\).
Proposition 6 (Temperature limits). (a) As \(\beta\to 0\): \(\lambda_*= \beta^2\,\mu_{\min}\bigl(\sum_x p(x)\tilde{d}_x\tilde{d}_x^\top\bigr) + O(\beta^3)\), where \(\tilde{d}_x(y) = d(x,y) - \frac{1}{N}\sum_{y'}d(x,y')\).
(b) As \(\beta\to\infty\), if each \(x\) has a unique optimal output \(y^*(x)\) and the map \(x\mapsto y^*(x)\) is non-constant: \(\lim_{\beta\to\infty}\lambda_*= c_0 > 0\).
Proof. (a) Expand \(K^*_x(y)=\frac{1}{N} - \beta\tilde{d}_x(y)+O(\beta^2)\); restrict the Gram matrix to \(T\). (b) \(K^*_x(y)\to\mathbf{1}_{\{y=y^*(x)\}}\); invoke eigenvalue continuity. ◻
We now turn to the discrete BA iteration and ask how much the Lyapunov functional decreases in a single step. The answer is given by a spectral formula involving the function \(d(\lambda) = -\lambda + \frac{3}{2}\lambda^2 - \frac{1}{2}\lambda^3\), which encodes the effect of the non-idempotence of \(\mathcal{G}\) on the per-step dissipation. Two features of \(d\) deserve attention before the formal analysis: first, \(d(\lambda)<0\) for all \(\lambda\in(0,1)\), confirming strict decrease; second, \(d\) vanishes at both endpoints \(\lambda=0\) and \(\lambda=1\), revealing the double-bottleneck structure discussed in Section 5.4. The section closes with a numerical example on a binary channel that illustrates the spectral dissipation formula concretely.
We now turn to the discrete BA iteration and analyse the one-step decrease of the Lyapunov functional \[\label{eq:Lyap} \mathcal{L}(q) = \sum_x p(x)\,D_{\mathrm{KL}}\bigl(K^*(\cdot|x)\;\|\;K_q(\cdot|x)\bigr).\tag{7}\] This functional satisfies \(\mathcal{L}(q)\ge 0\) with equality iff \(q=q^*\), and decreases along the BA iteration [4].
Remark 7 (Quadratic expansions near \(q^*\)). For a small perturbation \(h\in T\), write \(q = q^*+h\). Then \[\mathcal{L}(q^*+h) = \tfrac12\langle h,\, (I-\mathcal{G})h\rangle_{q^*} + O(\|h\|^3), \qquad F_\beta(q^*+h)-F_\beta(q^*) = \tfrac12\langle h,\, \mathcal{G}h\rangle_{q^*} + O(\|h\|^3).\] The two quadratic leading terms sum to \(\tfrac12\|h\|_*^2\), reflecting their complementary information-geometric roles: \(\mathcal{L}\) measures the average conditional variance of the log-perturbation \(\phi=h/q^*\), while \(F_\beta-F_\beta(q^*)\) measures the squared mean of the conditional expectation.
Fix an interior fixed point \(q^*\) and write \(q = q^*+h\) with \(h\in T\) small. For each \(x\in\mathcal{X}\), the Gibbs kernel at \(q\) satisfies \[K_q(x,y) = K^*_x(y)\,\frac{1 + \varepsilon r(y,x) + \tfrac12\varepsilon^2(r(y,x)^2 - \sigma^2(x)) + O(\varepsilon^3)}{1},\] where \(\varepsilon\sim\|h\|_*\), \(r(y,x) = \phi(y) - m(x)\) is the conditional residual introduced in Section 2 (a centred function of \(y\) under \(K^*_x\)), \(\sigma^2(x) = \mathbb{E}_{K^*_x}[r(\cdot,x)^2]\) is its conditional variance, and \(\mu_3(x) = \mathbb{E}_{K^*_x}[r(\cdot,x)^3]\) is the conditional skewness. Substituting into \(\mathcal{L}(q) = -\sum_x p(x)\mathbb{E}_{K^*_x}[\log(K_q(x,\cdot)/K^*_x(\cdot))]\) and expanding the logarithm, one finds: \[\label{eq:L95expand} \mathcal{L}(q^*+h) = \underbrace{\tfrac12\mathbb{E}_p[\sigma^2(x;h)]}_{\mathcal{L}^{(2)}} + \underbrace{\tfrac16\mathbb{E}_p[\mu_3(x;h)]}_{\mathcal{L}^{(3)}} + \underbrace{\tfrac{1}{24}\mathbb{E}_p[\kappa_4(x;h)]}_{\mathcal{L}^{(4)}} + O(\|h\|^5),\tag{8}\] where \(\kappa_4(x;h)=\mathbb{E}_{K^*_x}[r(\cdot,x)^4]-3\sigma^4(x;h)\) is the conditional excess kurtosis, and all expectations are over \(y\sim K^*_x\). The notation \(\sigma^2(x;h)\) makes explicit the dependence on \(h\) through \(r(y,x) = h(y)/q^*(y) - m(x;h)\).
The quadratic leading term can be expressed directly in terms of \(\mathcal{G}\): \[\label{eq:L295Fisher} \mathcal{L}^{(2)}(h) = \tfrac12\mathbb{E}_p[\sigma^2(x;h)] = \tfrac12\bigl(\|h\|_*^2 - \langle h,\, \mathcal{G}h\rangle_{q^*}\bigr) = \tfrac12\langle h,\, (I-\mathcal{G})h\rangle_{q^*}.\tag{9}\] To see this, note \(\|h\|_*^2 = \sum_y h(y)^2/q^*(y) = \mathbb{E}_{q^*}[\phi^2]\) and \(\langle h,\, \mathcal{G}h\rangle_{q^*} = \mathbb{E}_p[m(x)^2] = \mathbb{E}_p[\mathbb{E}_{K^*_x}[\phi]^2]\); then \(\mathbb{E}_p[\sigma^2] = \mathbb{E}_{q^*}[\phi^2] - \mathbb{E}_p[m^2]\) is the law of total variance.
Write the image of \(q^*+h\) under \(\mathcal{T}\) as \[\mathcal{T}(q^*+h) = q^* + h', \quad h' = h^{(1)} + h^{(2)} + O(\|h\|^3),\] where the terms are determined by expanding \(\mathcal{T}\) around \(q^*\):
\(h^{(1)} = D\mathcal{T}(q^*)\,h = (I-\mathcal{G})h\) is the first-order term (Theorem 3);
\(h^{(2)} = \tfrac{1}{2}D^2\mathcal{T}(q^*)[h,h] \in T\) is the second-order term arising from the curvature of \(\mathcal{T}\) at \(q^*\); it is a symmetric bilinear function of \(h\) with itself, and lies in \(T\) since \(\mathcal{T}\) maps into \(\Delta(\mathcal{Y})\).
The second-order one-step Lyapunov decrement is \[\Delta\mathcal{L}^{(2)} := \mathcal{L}^{(2)}(h') - \mathcal{L}^{(2)}(h),\] which measures the change in the quadratic part of \(\mathcal{L}\) after one BA step. Through order \(\|h\|^2\), only \(h^{(1)}\) contributes.
Theorem 8 (Exact one-step dissipation). With the notation above, the second-order one-step Lyapunov decrement is \[\label{eq:DeltaL2} \Delta\mathcal{L}^{(2)} = \mathcal{L}^{(2)}(h^{(1)}) - \mathcal{L}^{(2)}(h) = -\langle h,\, \mathcal{G}h\rangle_{q^*} + \tfrac32\langle h,\, \mathcal{G}^2 h\rangle_{q^*} - \tfrac12\langle h,\, \mathcal{G}^3 h\rangle_{q^*},\qquad{(10)}\] where all inner products are taken in \(\mathcal{H}_{q^*}= (T, \langle\cdot,\cdot\rangle_*)\) and \(\mathcal{G}^k\) denotes the \(k\)-fold composition of \(\mathcal{G}\) with itself on \(T\).
Proof. Using 9 : \[\Delta\mathcal{L}^{(2)} = \tfrac12\langle h^{(1)},\, (I-\mathcal{G})h^{(1)}\rangle_{q^*} - \tfrac12\langle h,\, (I-\mathcal{G})h\rangle_{q^*}.\] With \(h^{(1)} = (I-\mathcal{G})h\) and symmetry of \(\mathcal{G}\): \[\begin{align} \langle h^{(1)},\, h^{(1)}\rangle_{q^*} &= \|h\|_*^2 - 2\langle h,\, \mathcal{G}h\rangle_{q^*} + \langle h,\, \mathcal{G}^2 h\rangle_{q^*},\\ \langle h^{(1)},\, \mathcal{G}h^{(1)}\rangle_{q^*} &= \langle h,\, \mathcal{G}h\rangle_{q^*} - 2\langle h,\, \mathcal{G}^2 h\rangle_{q^*} + \langle h,\, \mathcal{G}^3 h\rangle_{q^*}. \end{align}\] Substituting and simplifying: \[\Delta\mathcal{L}^{(2)} = \tfrac12\bigl[ {-}2\langle h,\, \mathcal{G}h\rangle_{q^*} + 3\langle h,\, \mathcal{G}^2 h\rangle_{q^*} - \langle h,\, \mathcal{G}^3 h\rangle_{q^*} \bigr],\] which is ?? . ◻
Remark 9 (Role of non-idempotence). If \(\mathcal{G}\) were idempotent (\(\mathcal{G}^2 = \mathcal{G}\), eigenvalues in \(\{0,1\}\)), then \(\mathcal{G}^2 = \mathcal{G}^3 = \mathcal{G}\) and formula ?? would reduce to \((-1+\tfrac32-\tfrac12)\langle h,\, \mathcal{G}h\rangle_{q^*} = 0\). Non-idempotence (eigenvalues in \((0,1)\)) is therefore what drives strictly negative dissipation.
Since \(\mathcal{G}:\mathcal{H}_{q^*}\to\mathcal{H}_{q^*}\) is a self-adjoint operator on the \((N-1)\)-dimensional Hilbert space \((T,\langle\cdot,\cdot\rangle_*)\), the spectral theorem provides an orthonormal eigenbasis \(\{e_i\}_{i=1}^{N-1}\subset T\) with \(\langle e_i, e_j\rangle_* = \delta_{ij}\) and corresponding eigenvalues \(\lambda_i\in[0,1]\) satisfying \(\mathcal{G}e_i = \lambda_i e_i\). The eigenvalues lie in \([0,1]\) because \(\mathcal{G}= \Pi^*\Pi\) implies \(\langle u, \mathcal{G}u\rangle_* = \|\Pi u\|_p^2 \ge 0\) and \(\langle u, \mathcal{G}u\rangle_* \le \|u\|_*^2\) (since \(\Pi\) is a contraction). Any \(h\in T\) decomposes as \(h = \sum_i c_i e_i\) with \(c_i = \langle h, e_i\rangle_*\).
Corollary 3 (Spectral dissipation). \[\label{eq:spectral} \Delta\mathcal{L}^{(2)} = \sum_i c_i^2\,d(\lambda_i), \qquad d(\lambda) = -\lambda + \tfrac32\lambda^2 - \tfrac12\lambda^3.\qquad{(11)}\]
Proof. Since \(\mathcal{G}^k e_i = \lambda_i^k e_i\), substitute into ?? . ◻
Proposition 10 (Properties of \(d\)).
\(d(0) = d(1) = 0\); \(d(\lambda) < 0\) for all \(\lambda\in(0,1)\).
\(d\) attains its minimum at \(\lambda_{\mathrm{opt}} = 1 - 1/\sqrt{3} \approx 0.423\), with \(d(\lambda_{\mathrm{opt}}) = -\frac{1}{3\sqrt{3}} \approx -0.192\).
For small \(\lambda\): \(d(\lambda) = -\lambda + O(\lambda^2)\), so slow-contracting directions (small \(\lambda\)) also dissipate slowly.
For \(\lambda\) near \(1\): \(d(\lambda) = -(1-\lambda)\cdot\frac{\lambda(2-\lambda)}{2} \to 0\), so near-fully-contracted directions also have near-zero dissipation.
Proof. (a) Direct evaluation. (b) \(d'(\lambda) = -1 + 3\lambda - \tfrac32\lambda^2 = 0\) gives \(\lambda = 1 \pm 1/\sqrt{3}\); the interior minimum is \(\lambda_{\mathrm{opt}} = 1-1/\sqrt{3}\). Substituting into \(d\) yields \(d(\lambda_{\mathrm{opt}}) = -\frac{1}{3\sqrt{3}}\). (c)–(d) Taylor expansion. ◻
Remark 11 (Double bottleneck). Proposition 10 reveals a double-sided bottleneck* absent from first-order (linearised) analysis:*
Directions with \(\lambda_i\approx 0\) (weak coupling): the linearised contraction factor \((1-\lambda_i)\approx 1\) is near unity, and \(d(\lambda_i)\approx -\lambda_i\approx 0\), so these directions are slow by both measures.
Directions with \(\lambda_i\approx 1\) (strong coupling): the linearised contraction factor \((1-\lambda_i)\approx 0\) is excellent, but \(\mathcal{L}^{(2)} \propto (1-\lambda_i)\approx 0\), so the Lyapunov mass in these directions is already negligible. Dissipation \(d(\lambda_i)\approx 0\) reflects that there is very little left to dissipate.
The practically relevant bottleneck is the \(\lambda\approx 0\) directions (small spectral gap), which governs the asymptotic convergence rate.
Proposition 12 (Non-idempotence correction). For all \(\lambda\in(0,1)\): \[d(\lambda) > -\lambda, \qquad d(\lambda) - (-\lambda) = \tfrac{1}{2}\lambda^2(3-\lambda) > 0.\] The per-step dissipation is strictly less negative than \(-\lambda\), with the correction of order \(\lambda^2\) for small \(\lambda\).
Proof. Direct: \(\frac{3}{2}\lambda^2 - \frac{1}{2}\lambda^3 = \frac{1}{2}\lambda^2(3-\lambda) > 0\) for \(\lambda\in(0,1)\). ◻
This correction reflects the curvature of \(\mathcal{T}\) at \(q^*\): the non-idempotence of \(\mathcal{G}\) means that the image of the linearised step \((I-\mathcal{G})h\) has strictly smaller Fisher–Rao norm than predicted by the first-order eigenvalue alone.
We illustrate the spectral dissipation formula on the simplest non-trivial instance: a binary channel with asymmetric source. All quantities are computed exactly and can be verified by hand.
Let \(\mathcal{X}= \mathcal{Y}= \{0,1\}\), \(p = (0.6,\,0.4)\), Hamming distortion \(d(x,y) = \mathbf{1}_{\{x\ne y\}}\), and \(\beta = 1\). Set \(r = e^{-\beta} = e^{-1} \approx 0.3679\).
The BA operator acts on \(q = (q_0, q_1)\in\operatorname{int}\Delta(\mathcal{Y})\) via \[(\mathcal{T}q)(0) = p_0\,\frac{q_0}{q_0 + r q_1} + p_1\,\frac{r q_0}{r q_0 + q_1}.\] Setting \((\mathcal{T}q^*)(0) = q^*_0\) and solving gives \[q^* = (0.7163953414,\; 0.2836046586),\] which satisfies \((\mathcal{T}q^*)(0) = q^*_0 = 0.7163953414\) exactly.
With partition functions \(Z^*_0 = q^*_0 + r q^*_1\) and \(Z^*_1 = r q^*_0 + q^*_1\): \[K^*_{x=0} = \bigl(0.8728782667,\; 0.1271217333\bigr), \qquad K^*_{x=1} = \bigl(0.4816709534,\; 0.5183290466\bigr).\]
Since \(|\mathcal{Y}|=2\), the tangent space \(T\) is one-dimensional, spanned by \(h = (1,-1)\). The log-perturbation is \(\phi = h/q^*\), and the conditional means are \[m(x=0) = \mathbb{E}_{K^*_{x=0}}[\phi] \approx 0.7702, \qquad m(x=1) = \mathbb{E}_{K^*_{x=1}}[\phi] \approx -1.1553.\] Computing \(\langle h,\mathcal{G}h\rangle_* = 0.8898\) and \(\|h\|_*^2 = 4.9219\) via Definition 4, the unique eigenvalue of \(\mathcal{G}|_T\) is \[\lambda = \frac{\langle h,\mathcal{G}h\rangle_*}{\|h\|_*^2} = \frac{0.8898}{4.9219} = 0.1808.\] Note that this is not the off-diagonal matrix element \(\mathcal{G}_{01}\), but the Rayleigh quotient of \(\mathcal{G}\) in the Fisher–Rao metric.
The exact formula of Corollary 3 gives \[d(\lambda) = -\lambda + \tfrac{3}{2}\lambda^2 - \tfrac{1}{2}\lambda^3 = -0.1808 + 1.5\times 0.0327 - 0.5\times 0.0059 = -0.1347.\] The correction due to non-idempotence (Proposition 12) accounts for \(0.1808 - 0.1347 = 0.0461\), which is \(\frac{1}{2}\lambda^2(3-\lambda) = 0.5\times 0.0327\times 2.8192 \approx 0.0461\), consistent with the analytic formula. The KL convergence factor for this channel is \(\gamma = \lambda_*(2-\lambda_*) = 0.1808\times 1.8192 \approx 0.329\), meaning approximately \(33\%\) of the KL gap is closed per iteration in the local phase.
The third-order one-step decrement is \[\label{eq:DeltaL3} \Delta\mathcal{L}^{(3)} = 2B(h^{(1)}, h^{(2)}) + \bigl[C(h^{(1)}) - C(h)\bigr],\tag{10}\] where \(B\) is the bilinear form of \(\mathcal{L}^{(2)}\) and \(C(h) = \frac{1}{6}\mathbb{E}_p[\mu_3(h)]\). If \(\mu_3(x;h)=0\) for all \(x\) (conditional symmetry), then \(C(\cdot)\equiv 0\) and the bracket in 10 vanishes, but the term \(2B(h^{(1)},h^{(2)})\) survives because \(h^{(2)}\) encodes conditional variances, not skewness. Thus:
Proposition 13 (Cubic cancellation does not follow from symmetry). Conditional symmetry \(\mu_3(x;h)=0\) does not guarantee \(\Delta\mathcal{L}^{(3)}=0\).
This contrasts with a different notion of “cubic cancellation”—the vanishing of the third-order jet \(\langle h, D^2\mathcal{T}^*(h,h)\rangle_*\) when \(\mu_3=0\) [14]—which refers to the map expansion rather than the Lyapunov decrement.
The spectral dissipation formula of Section 5 shows that convergence is governed, in the slowest directions, by the smallest eigenvalue of \(\mathcal{G}\) on the tangent space. This section makes that connection precise by proving local exponential contraction in a neighbourhood of a non-degenerate fixed point. The rate is exactly \(\lambda_*= \lambda_{\min}(\mathcal{G}|_T)\), and the proof is a direct application of Gronwall’s inequality to the Lyapunov function \(\mathcal{E} = \frac{1}{2}\|q - q^*\|_*^2\). These local results are the second ingredient in the two-stage global argument assembled in Section 7.
Definition 6 (Spectral gap). The spectral gap* of \(\mathcal{G}\) is the smallest eigenvalue of \(\mathcal{G}\) on \((T,\langle\cdot,\cdot\rangle_*)\): \[\label{eq:lam95def} \lambda_*:= \lambda_{\min}(\mathcal{G}|_T) = \min_{\substack{u\in T\\\|u\|_*=1}} \langle u,\mathcal{G}u\rangle_* = \min_{\substack{u\in T\\\|u\|_*=1}} \sum_{x\in\mathcal{X}} p(x)\,\bigl(\mathbb{E}_{K^*_x}[\phi_u]\bigr)^2,\tag{11}\] where \(\phi_u = u/q^*\) is the log-perturbation corresponding to \(u\). The fixed point \(q^*\) is non-degenerate if \(\lambda_*>0\), i.e.if the equilibrium kernels \(\{K^*_x : x\in\mathcal{X}\}\) span \(T\) in the sense that no nonzero tangent direction is invisible to all source symbols.*
Note also that the largest* eigenvalue of \(\mathcal{G}\) on \(T\) satisfies \(\lambda_{\max}(\mathcal{G}|_T) \le 1\), with equality excluded when \(\mathcal{G}^2\ne\mathcal{G}\). This implies \(\|(I-\mathcal{G})|_T\|_{\mathrm{op}} = 1 - \lambda_*\) in the \(\langle\cdot,\cdot\rangle_*\)-operator norm, a fact used in the local contraction argument (Appendix 12).*
The variational form 11 is transparent: \(\lambda_*\) is the minimum variance of the projections \(\langle K^*_x,u\rangle\) over all unit tangent directions. Non-degeneracy means the equilibrium kernels \(\{K^*_x\}\) span the full tangent hyperplane.
Theorem 14 (Spectral contraction). Assume \(q^*\) is non-degenerate (\(\lambda_*>0\)). Let \(L > 0\) be a Lipschitz constant for \(D\mathcal{T}\) near \(q^*\) (exists by Fréchet differentiability of \(\mathcal{T}\)) and set \(\rho_* = \lambda_*/(2L)\). There exist constants \(c_1, c_2, C_0 > 0\) such that for all \(q\in\operatorname{int}\Delta(\mathcal{Y})\) with \(\|q-q^*\|_1\le\rho_*\):
(Bi-Lipschitz) \(c_1\|q-q^*\|_1 \le \|\mathcal{T}q-q\|_1 \le c_2\|q-q^*\|_1\), where \(c_1, c_2\) depend only on \(\lambda_*\), \(L\), and \(\|\mathcal{G}\|\).
(Exponential contraction) For the continuous-time flow, \[\|q(t)-q^*\|_1 \le C_0\,e^{-\lambda_*t/2}\,\|q(0)-q^*\|_1, \quad t\ge 0.\]
(Isolation) \(q^*\) is the unique fixed point of \(\mathcal{T}\) in \(B_1(q^*,\rho_*)\cap\mathcal{A}\), where \(\mathcal{A}\) is the connected component of \(q^*\) in \(\operatorname{int}\Delta(\mathcal{Y})\).
Proof. Since \(DV(q^*)=-\mathcal{G}\) is invertible on \((T,\langle\cdot,\cdot\rangle_*)\) (\(\lambda_*>0\) ensures \(\mathcal{G}\) has trivial kernel on \(T\)), the inverse function theorem applied to \(V:\operatorname{int}\Delta(\mathcal{Y})\to T\) gives (A) and (C) with constants computable from \(\lambda_*\) and \(\|\mathcal{G}\|\).
For (B): write \(u = q-q^*\in T\) and \(\dot{u} = V(q^*+u) = -\mathcal{G}u + R(u)\) where \(\|R(u)\|_* \le L\|u\|_*^2\) by the Lipschitz assumption on \(D\mathcal{T}\).
Define \(\mathcal{E}(t) = \tfrac12\|u(t)\|_*^2\). Then \[\begin{align} \dot{\mathcal{E}} &= \langle u, \dot{u}\rangle_* = -\langle u, \mathcal{G}u\rangle_* + \langle u, R(u)\rangle_* \le -\lambda_*\|u\|_*^2 + L\|u\|_*^3 = -\|u\|_*^2(\lambda_*- L\|u\|_*). \end{align}\] When \(\|u\|_* \le \rho_* = \lambda_*/(2L)\), we have \(\lambda_*- L\|u\|_* \ge \lambda_*/2\), hence \[\dot{\mathcal{E}} \le -\frac{\lambda_*}{2}\|u\|_*^2 = -\lambda_*\,\mathcal{E}.\] Gronwall’s inequality gives \(\mathcal{E}(t) \le \mathcal{E}(0)\,e^{-\lambda_*t}\), so \(\|u(t)\|_* \le \|u(0)\|_*\,e^{-\lambda_*t/2}\).
Finally, since \(\mathcal{Y}\) is finite and all norms on \(\mathbb{R}^\mathcal{Y}\) are equivalent, there exists \(C_0 > 0\) (depending on \(q^*\) through the ratio \(\max_y q^*(y)/\min_y q^*(y)\)) such that \(\|u\|_1 \le C_0\,\|u\|_*\) and \(\|u\|_* \le C_0\,\|u\|_1\). This gives \(\|u(t)\|_1 \le C_0^2\,e^{-\lambda_*t/2}\|u(0)\|_1\); absorbing \(C_0^2\) into \(C_0\) yields (B). ◻
The \(\chi^2\)-dissipation identity and the local spectral contraction are individually useful but each has a limited scope: the dissipation identity is global but does not itself imply exponential decay, while the spectral contraction is exponential but only local. This section joins the two by a finite-entry argument: the dissipation identity guarantees that the BA residual \(\|\mathcal{T}q_t - q_t\|_1\) becomes small in finite time \(T_*\), after which the bi-Lipschitz equivalence of Theorem 14(A) places the trajectory in the exponential basin. The complete proof with explicit constants for the entry time and convergence factor is given in Appendix 12.
The argument requires that the free energy \(F_\beta\) is bounded below along the trajectory. For finite \(\mathcal{Y}\) this is immediate from continuity and compactness of \(\Delta(\mathcal{Y})\). For continuous alphabets with quadratic distortion, the following lemma provides the necessary control.
Lemma 3 (Uniform second-moment bound). Under Assumption 1, for any \(q_0\) with finite second moment, \(\sup_{t\ge 0}\mathbb{E}_{q_t}[\hat{X}^2] \le M < \infty\), and consequently \(F_\beta(q_t)\) is bounded below uniformly in \(t\).
Proof. Let \(V(t) = \operatorname{tr}(\Sigma(t))\) where \(\Sigma(t) = \int \hat{x}\hat{x}^\top q_t\,d\hat{x}\) is the second-moment matrix of \(q_t\). Differentiating along the BA flow: \(\dot{V} = \tilde{V} - V\), where \(\tilde{V} = \operatorname{tr}(\mathbb{E}_{\mathcal{T}(q_t)}[\hat{X}\hat{X}^\top])\) is the second moment of \(\mathcal{T}(q_t)\). For quadratic distortion \(d(x,\hat{x}) = (x-\hat{x})^2\), the posterior \(K_{q_t}(x,\cdot)\) is Gaussian with precision \(\Lambda = \Sigma^{-1}+2\beta I\), giving \(\tilde{V} = \operatorname{tr}(\Lambda^{-1}) + \operatorname{tr}(H\Sigma H)\) with \(H = 2\beta\Lambda^{-1}\) and \(\|H\|_{\mathrm{op}} = 2\beta/(1+2\beta\lambda_{\min}(\Sigma)) < 1\). Hence \(\dot{V} \le C_1 - (1-\|H\|^2_{\mathrm{op}})V\), yielding \(V(t)\le\max(V(0),C_1/(1-\|H\|^2_{\mathrm{op}}))<\infty\). The lower bound on \(F_\beta\) follows since \(F_\beta(q) \ge -\beta M/2 - \log N\) whenever \(\mathbb{E}_q[\hat{X}^2] \le M\). ◻
Theorem 15 (Two-stage global convergence). Let \(q^*\) be a non-degenerate interior fixed point. For every initial condition \(q_0\in\operatorname{int}\Delta(\mathcal{Y})\) in the connected component \(\mathcal{A}\) of \(q^*\), the BA flow satisfies:
(Global phase) There exists a finite time \(T_*<\infty\) such that \(\|\mathcal{T}(q(t))-q(t)\|_1\le\rho\) for all \(t\ge T_*\).
(Local phase) For \(t\ge T_*\): \[\label{eq:global95decay} \|q(t)-q^*\|_1 \le C_0\,e^{-\lambda_*(t-T_*)/2}\,\|q(T_*)-q^*\|_1.\qquad{(12)}\]
Proof. Step 1 (Global phase). By Corollary 1: \(\frac{d}{dt}F_\beta(q_t) = -\mathcal{D}(q_t) \le -\frac{1}{2}\|\mathcal{T}q_t - q_t\|_1^2\). If \(\|\mathcal{T}q_t-q_t\|_1\ge\delta\) for some fixed \(\delta>0\), then \(F_\beta(q_t) \le F_\beta(q_0) - \frac{1}{2}\delta^2 t\to -\infty\), contradicting the boundedness of \(F_\beta\) on \(\Delta(\mathcal{Y})\) (since \(F_\beta\) is continuous on the compact simplex; for continuous alphabets, Lemma 3 provides a uniform lower bound). Hence there exists \(T_*<\infty\) with \(\|\mathcal{T}q(T_*)-q(T_*)\|_1\le\delta\).
Step 2 (Entry into basin). By Theorem 14(A), the bound \(\|\mathcal{T}q-q\|_1\le\delta\) with \(\delta\) sufficiently small implies \(\|q-q^*\|_1\le\rho\) for \(q\) in the connected component \(\mathcal{A}\).
Step 3 (Local phase). Apply Theorem 14(B) from time \(T_*\). ◻
Remark 16 (Mechanism of two-stage convergence). The two stages are dynamically distinct. In the global phase, \(\chi^2\)-dissipation uniformly reduces the BA residual regardless of the trajectory’s geometry. In the local phase, the correlation structure encoded in \(\mathcal{G}\) drives genuine exponential contraction. The bridge between them is Theorem 14(A): a small BA residual and a small distance to \(q^*\) are equivalent near a non-degenerate equilibrium.
Remark 17 (Comparison with Hayashi’s \(O(1/n)\) rate). Hayashi [2] proves global \(O(1/n)\) convergence for the discrete iteration. Theorem 15 gives a complementary picture: the global phase has finite duration \(T_*\) (computable from \(F_\beta(q_0)-F_\beta(q^*)\) and the dissipation lower bound), and the subsequent exponential phase has explicit rate \(\lambda_*\). The two results are consistent and together describe the full trajectory from any initial condition.
The two-stage theorem establishes exponential convergence, but leaves the convergence factor implicit. This section makes it explicit for the discrete BA iteration \(q_{n+1} = \mathcal{T}(q_n)\), \(n = 0, 1, 2, \ldots\) . Write \(v_n = q_n - q^* \in T\) for the signed deviation of the \(n\)-th iterate from the fixed point; this is a well-defined element of the tangent space \(T\) whenever \(q_n \in \operatorname{int}\Delta(\mathcal{Y})\). In the local phase, the KL divergence \(D_{\mathrm{KL}}(q^*\|q_n)\) contracts by a factor \((1-\lambda_*)^2\) per iteration, up to cubic corrections in \(\|v_n\|_*\). The improvement factor \(\gamma = \lambda_*(2-\lambda_*)\) is a simple function of the spectral gap and can be computed directly from the channel and temperature. This upgrades the qualitative monotonicity of Ramakrishnan et al [9] to a quantitative rate.
Corollary 4 (Explicit KL convergence factor). Let \(q^*\) be non-degenerate and \(\{q_n\}\) the discrete BA iterates with \(q_n\) sufficiently close to \(q^*\). Then: \[\label{eq:KL95factor} D_{\mathrm{KL}}(q^*\|q_{n+1}) \le (1-\lambda_*)^2\,D_{\mathrm{KL}}(q^*\|q_n) + O(\|v_n\|_*^3).\qquad{(13)}\] The per-iteration improvement factor is \[\label{eq:gamma} \gamma = 1 - (1-\lambda_*)^2 = \lambda_*(2-\lambda_*).\qquad{(14)}\] Equality in ?? holds asymptotically when \(v_n\) is aligned with the \(\lambda_*\)-eigendirection of \(\mathcal{G}\).
Proof. Step 1. By the spectral contraction (Theorem 14) and the eigenbasis of \(\mathcal{G}\), the slowest-converging component of \(v_n\) satisfies \(v_{n+1} = (I-\mathcal{G})v_n + O(\|v_n\|^2)\). In the \(\lambda_*\)-eigendirection: \(c_{\min}^{(n+1)} = (1-\lambda_*)c_{\min}^{(n)} + O(\|v_n\|^2)\).
Step 2. The KL divergence \(D_{\mathrm{KL}}(q^*\|q)\) has the Fisher–Rao expansion \[\label{eq:KL95FR} D_{\mathrm{KL}}(q^*\|q^*+v) = \tfrac12\|v\|_*^2 + O(\|v\|^3).\tag{12}\] (Standard result: the Fisher–Rao metric is the Hessian of \(D_{\mathrm{KL}}\) at \(q^*\).)
Step 3. Since \(v_{n+1} = (I-\mathcal{G})v_n + O(\|v_n\|^2)\) and \(\|(I-\mathcal{G})v_n\|_*^2 \le (1-\lambda_*)^2\|v_n\|_*^2\) (with equality along the \(\lambda_*\)-direction), \[\begin{align} D_{\mathrm{KL}}(q^*\|q_{n+1}) =& \tfrac12\|v_{n+1}\|_*^2 + O(\|v_n\|^3) \le (1-\lambda_*)^2\,\tfrac12\|v_n\|_*^2 + O(\|v_n\|^3)\\ = &(1-\lambda_*)^2\,D_{\mathrm{KL}}(q^*\|q_n) + O(\|v_n\|^3). \end{align}\] When \(v_n\) is exactly along the \(\lambda_*\)-eigendirection, the inequality becomes an equality asymptotically. ◻
Remark 18 (Relation to monotonicity). Ramakrishnan et al [9] proved that \(D_{\mathrm{KL}}(q^*\|q_{n+1})\le D_{\mathrm{KL}}(q^*\|q_n)\) (qualitative monotonicity). Corollary 4 upgrades this to a quantitative* rate: the convergence factor \((1-\lambda_*)^2 < 1\) is explicit and computable from the channel and temperature. For a Gaussian source with Gaussian initial condition, \(\lambda_*= 1/(2\beta\sigma^2)\) gives \(\gamma = \lambda_*(2-\lambda_*) = \frac{1}{2\beta\sigma^2}(2-\frac{1}{2\beta\sigma^2})\).*
Remark 19 (\(\gamma\) and the spectral function \(d\)). The improvement factor \(\gamma = \lambda_*(2-\lambda_*)\) in KL is related to but distinct from the spectral dissipation \(d(\lambda_*)\) in Corollary 3. The KL factor measures the contraction of the full quadratic form \(\|v\|_*^2\), while \(d(\lambda)\) measures the one-step decrement of the conditional-kernel Lyapunov \(\mathcal{L}\). Both are governed by \(\mathcal{G}\), but via different spectral combinations.
The abstract theory of the preceding sections takes its sharpest form when specialised to Gaussian sources with quadratic distortion. In this case the spectral gap is \(\lambda_*= 1/(2\beta\sigma^2)\), the Jacobian is diagonalised by Hermite polynomials, and the convergence factor \(\gamma = \lambda_*(2-\lambda_*)\) is a simple explicit function of the inverse temperature and source variance. The section also identifies critical slowing down as \(\beta\) approaches the rate-distortion threshold \(1/(2\sigma^2)\), a phenomenon that emerges naturally from the vanishing of \(\lambda_*\).
When the source is Gaussian and distortion is quadratic, the general theory achieves its sharpest quantitative form.
Let \(X\sim\mathcal{N}(0,\sigma^2)\) and \(d(x,\hat{x})=(x-\hat{x})^2\). The BA operator becomes \[\label{eq:Gauss95BA} \mathcal{T}(q)(\hat{x}) = \int \frac{e^{-\beta(x-\hat{x})^2}q(\hat{x})}{\int e^{-\beta(x-\hat{y})^2}q(\hat{y})\,d\hat{y}}\, \frac{e^{-x^2/(2\sigma^2)}}{\sqrt{2\pi\sigma^2}}\,dx.\tag{13}\] For a Gaussian initial distribution \(q_0 = \mathcal{N}(0,s_0)\), \(q_t\) remains Gaussian under the BA flow, and its variance \(s(t) = \mathbb{E}_{q_t}[\hat{X}^2]\) satisfies a closed ODE.
Theorem 20 (Exact variance ODE for Gaussian initial conditions). If \(q_0\) is Gaussian, then \(q_t\) is Gaussian for all \(t\), and its variance \(s(t) = \mathbb{E}_{q_t}[\hat{X}^2]\) satisfies \[\label{eq:var95ODE} \dot{s}(t) = \tilde{s}(s(t),\beta) - s(t),\qquad{(15)}\] with \[\label{eq:s95tilde} \tilde{s}(s,\beta) := \mathbb{E}_{\mathcal{T}(q)}[\hat{X}^2] = \frac{s}{1+2\beta s} + \frac{(2\beta s)^2\sigma^2}{(1+2\beta s)^2},\qquad{(16)}\] independently of the mean of \(q_0\). The unique stable fixed point is \(s^* = \sigma^2 - \frac{1}{2\beta}\), positive iff \(\beta > 1/(2\sigma^2)\).
Proof. For Gaussian \(q\), all integrals are Gaussian, and the expression for \(\tilde{s}\) follows from standard Gaussian integration (completing the square). The fixed point equation \(\tilde{s}(s,\beta)=s\) yields \(s^* = \sigma^2 - 1/(2\beta)\). Linearisation around \(s^*\) gives \(\dot{\delta} = -2\beta s^*\delta + O(\delta^2)\), confirming stability. ◻
Remark 21. Equation ?? holds for Gaussian initial conditions with any mean; the mean decays to zero exponentially at rate \(1/(2\beta s^*)\) and does not affect the variance dynamics. For non-Gaussian initial conditions, the Gaussian family is not invariant, but Theorem 25 below shows convergence to the Gaussian fixed point in total variation.
Proposition 22 (Hermite diagonalisation). At \(q^* = \mathcal{N}(0,s^*)\), the Jacobian \(D\mathcal{T}(q^*)\) is diagonalised in \(\mathcal{H}_{q^*}\) by the Hermite polynomials \(He_n(\hat{x}/\sqrt{s^*})\), \(n=1,2,\ldots\), with eigenvalues \[\label{eq:alpha} \mu_n = \alpha^n, \qquad \alpha := \frac{s^*}{s^*+(2\beta)^{-1}} = 1 - \frac{1}{2\beta\sigma^2} \in(0,1).\qquad{(17)}\] The eigenvalues of \(\mathcal{G}= I - D\mathcal{T}(q^*)\) are therefore \(\lambda_n = 1-\alpha^n\).
Proof. The linearised BA map at a Gaussian fixed point is a Mehler-type convolution; its eigenfunctions in the Gaussian \(L^2\) space are Hermite polynomials [15], with eigenvalues \(\alpha^n\) by direct computation on the generating function. ◻
Theorem 23 (Gaussian spectral gap). \[\label{eq:Gauss95gap} \lambda_*= \lambda_1 = 1-\alpha = \frac{1}{2\beta\sigma^2}.\qquad{(18)}\] The local relaxation time is \(\tau_{\mathrm{relax}} = 1/\lambda_*= 2\beta\sigma^2\).
Remark 24 (Critical slowing down). As \(\beta\to 1/(2\sigma^2)^+\) (approaching the rate-distortion threshold), \(s^*\to 0\) and \(\alpha\to 1\), so \(\lambda_*\to 0\). This is the phenomenon of critical slowing down: near the phase transition, equilibration time \(\tau_{\mathrm{relax}}\) diverges.
At the Gaussian fixed point, the \(\mathcal{G}\)-eigenvalues are \(\lambda_n = 1-\alpha^n\in(0,1)\). The spectral function \(d(\lambda_n) = -\lambda_n + \frac{3}{2}\lambda_n^2 - \frac{1}{2}\lambda_n^3\) is negative for all \(n\ge 1\), confirming strict dissipation. The optimal dissipation occurs at the eigenvalue closest to \(\lambda_{\mathrm{opt}}\approx 0.423\). Since \(\lambda_n = 1-\alpha^n\) is increasing in \(n\), the optimal dissipation index is \(n^* = \lfloor \log(1-\lambda_{\mathrm{opt}})/ \log\alpha\rfloor\).
The KL convergence factor is explicitly \[\gamma = \lambda_*(2-\lambda_*) = \frac{1}{2\beta\sigma^2}\!\left(2 - \frac{1}{2\beta\sigma^2}\right),\] which increases from \(0\) (at threshold) toward \(2-1 = 1\) (as \(\beta\to\infty\)).
Theorem 25 (Gaussian attractor). Under the BA flow with \(X\sim\mathcal{N}(0,\sigma^2)\) and quadratic distortion, if \(\beta>1/(2\sigma^2)\), then for any initial \(q_0\) with finite second moment, \(q_t\to\mathcal{N}(0,s^*)\) in total variation as \(t\to\infty\).
Proof. The proof proceeds in five steps: uniform moment control (Lemma 7), spectral decay of Hermite modes (Lemma 8), variance convergence (Lemma 9), weak convergence via subsequential compactness (Lemma 10), and upgrade to total variation via Pinsker’s inequality and KL convergence (Lemma 11). The complete argument is given in Appendix 13. ◻
The Gaussian distribution emerges here as a dynamical consequence, not a variational assumption. The BA flow is a spectral filter suppressing non-Gaussian cumulants hierarchically.
The two-point and three-cluster models below serve a dual purpose. They provide closed-form expressions for \(\lambda_*\), making the spectral gap directly verifiable, and they illustrate the two-stage convergence structure in settings simple enough to be analysed by hand. Both examples confirm the general theory and offer concrete parameter dependence that may guide intuition for larger channels.
Let \(\mathcal{X}=\mathcal{Y}=\{0,1\}\), \(p = (p_0,p_1)\), and \(d(x,y)=\mathbf{1}_{x\ne y}\). The BA operator with \(\beta>0\) reduces to a scalar equation in \(q = q(1)\in(0,1)\). The fixed point satisfies \[q^* = \frac{p_0 e^{-\beta\cdot 0} \cdot q^*}{Z_0(q^*)} + \frac{p_1 e^{-\beta\cdot 1}\cdot q^*}{Z_1(q^*)},\] and the spectral gap is \[\lambda_*= p_0 K^*_0(0)K^*_0(1) + p_1 K^*_1(0)K^*_1(1)\cdot(-1)^2\] (sum of squared equilibrium kernels evaluated on the single tangent direction). This gives \(\lambda_*\) as an explicit function of \(p\) and \(\beta\), with \(\lambda_*\to 0\) at the hard-decision boundary \(\beta\to\infty\).
Let \(\mathcal{X}= \{1,2,3\}\), \(\mathcal{Y}= \{1,2,3\}\), uniform \(p\), and distortion \(d(x,y) = 0\) if \(x=y\), \(d(x,y) = 1\) otherwise. This symmetric model has a uniform fixed point \(q^* = (1/3,1/3,1/3)\). By symmetry, the spectral gap of the two-dimensional tangent space is \[\lambda_*= \frac{2 e^{-\beta}}{(2e^{-\beta}+1)^2} = \frac{2r}{(2r+1)^2}, \qquad r = e^{-\beta}.\] For this model, the two-stage structure is clearly visible in simulation: from a highly asymmetric initial distribution, the global \(\chi^2\)-phase redistributes mass over a timescale \(O(1/\delta^2)\) (where \(\delta\) is the initial asymmetry), after which exponential contraction at rate \(\lambda_*\) takes over.
We have developed a unified spectral theory for BA dynamics, centred on the relaxation kernel \(\mathcal{G}= \mathbb{E}_p[K^*_X\otimes K^*_X]\).
Summary of results.
The exact \(\chi^2\)-dissipation identity ?? reveals the precise entropy production mechanism of the continuous-time flow.
The threefold identity (Theorems 3 and 5) unifies statistical, operator-theoretic, and information-geometric views of \(\mathcal{G}\).
The spectral dissipation formula ?? with \(d(\lambda) = -\lambda+\frac{3}{2}\lambda^2-\frac{1}{2}\lambda^3\) quantifies the exact one-step Lyapunov decrease and reveals the double bottleneck.
The two-stage global convergence theorem combines finite-time basin entry (from \(\chi^2\)-dissipation) with local exponential contraction (from \(\lambda_*\)).
The explicit KL convergence factor \(\gamma = \lambda_*(2-\lambda_*)\) upgrades Ramakrishnan et al’s monotonicity to a constructive rate.
Relation to prior work. The results of this paper are intended as a complement to the frameworks of Hayashi [1]–[3] and Nakagawa et al [7]. Hayashi’s Bregman-EM analysis establishes elegant global convergence guarantees and an \(O(1/n)\) rate from any initial condition; the present contribution identifies the relaxation kernel \(\mathcal{G}\) as the central spectral object and makes the local convergence factor \(\gamma = \lambda_*(2-\lambda_*)\) explicit and computable from the channel and temperature. Nakagawa et al’s sharp \(O(1/n)\) rates govern the transient phase before the local exponential phase described here. Together, these results give a more complete quantitative picture of the full BA trajectory: global sublinear approach, finite-time entry, and explicit exponential contraction.
Algorithmic implications. The spectral theory developed here has direct consequences for the design and acceleration of BA-type algorithms. The double-bottleneck structure of \(d(\lambda)\), the explicit dependence of the convergence factor on \(\lambda_*\), and the two-stage entry mechanism all suggest concrete strategies for improving the iteration in practice. These algorithmic questions—including spectral preconditioning, momentum-based acceleration targeting the \(\lambda\approx 0\) directions, and online estimation of the spectral gap during the global phase—are taken up in my companion paper currently in preparation.
This research received no formal funding and was conducted based on the author’s independent academic interest.
This appendix provides the complete proof of the two-stage global exponential convergence theorem for the discrete BA iteration, with all constants made explicit. The argument assembles three ingredients from the main text: the \(\chi^2\)-dissipation identity (Theorem 1), the uniform moment bound (Lemma 3), the bi-Lipschitz equivalence of BA residual and distance to the fixed point (Theorem 14(A)), and the spectral contraction (Theorem 14(B)).
Theorem 26 (Two-Stage Global Exponential Convergence). Let \(q^*\in\operatorname{int}\Delta(\mathcal{Y})\) be a non-degenerate interior fixed point with spectral gap \(\lambda_*> 0\), Lipschitz constant \(L\) for \(D\mathcal{T}\) near \(q^*\), and bi-Lipschitz constant \(c_1\) from Theorem 14(A). Set \(\rho = \lambda_*/(4L)\). Then for every initial condition \(q_0\in\operatorname{int}\Delta(\mathcal{Y})\) in the connected component \(\mathcal{A}\) of \(q^*\), there exist explicit constants \(N_*(q_0) < \infty\) and \(C(q_0) > 0\) such that for all \(n \ge N_*\): \[D_{\mathrm{KL}}(q^*\|q_n) \le C(q_0)\cdot\left(1 - \frac{3\lambda_*}{4}\right)^{2(n - N_*)}.\] The entry time satisfies \[N_*(q_0) \le \left\lceil\frac{2(F_\beta(q_0) - F_\beta(q^*))}{c_1^2\rho^2}\right\rceil.\]
Lemma 4. There exists \(N_* < \infty\) such that \(\|q_{N_*} - q^*\|_1 \le \rho\).
Proof. By the \(\chi^2\)-dissipation identity (Theorem 1) and Corollary 1, the free energy satisfies \[F_\beta(q_{n+1}) - F_\beta(q_n) \le -\tfrac{1}{2}\|\mathcal{T}q_n - q_n\|_1^2\] along the discrete iteration. Suppose for contradiction that \(\|\mathcal{T}q_n - q_n\|_1 \ge \delta_0 := c_1\rho\) for all \(n \le M\). Summing over \(n = 0,\ldots,M-1\): \[F_\beta(q_M) \le F_\beta(q_0) - \frac{c_1^2\rho^2}{2}\,M.\] Choosing \(M = \lceil 2(F_\beta(q_0)-F_\beta(q^*))/(c_1^2\rho^2)\rceil\) gives \(F_\beta(q_M) < F_\beta(q^*)\), contradicting \(F_\beta\ge F_\beta(q^*)\) on \(\mathcal{A}\). Hence there exists \(N_* \le M\) with \(\|\mathcal{T}q_{N_*} - q_{N_*}\|_1 < c_1\rho\).
By Theorem 14(A) (bi-Lipschitz), this implies \(\|q_{N_*} - q^*\|_1 \le \rho\). ◻
For \(n \ge N_*\), write \(v_n = q_n - q^* \in T\). The BA iteration gives \[v_{n+1} = \mathcal{T}(q^*+v_n) - q^* = (I - \mathcal{G})v_n + R(v_n),\] where \(R(v_n) = \mathcal{T}(q^*+v_n) - q^* - D\mathcal{T}(q^*)v_n \in T\) is the nonlinear remainder satisfying \(\|R(v_n)\|_* \le L\|v_n\|_*^2\). The linear term \((I-\mathcal{G})v_n\) satisfies \[\|(I-\mathcal{G})v_n\|_* \le \|(I-\mathcal{G})|_T\|_{\mathrm{op}}\|v_n\|_* = (1-\lambda_*)\|v_n\|_*,\] because the operator norm of \((I-\mathcal{G})\) on \((T,\langle\cdot,\cdot\rangle_*)\) equals \(1 - \lambda_{\min}(\mathcal{G}|_T) = 1-\lambda_*\) (since eigenvalues of \(I-\mathcal{G}\) are \(1-\lambda_i\in[0,1-\lambda_*]\)).
Lemma 5. For \(\rho = \lambda_*/(4L)\), whenever \(\|v_n\|_* \le \rho\): \[\|v_{n+1}\|_* \le \left(1 - \frac{3\lambda_*}{4}\right)\|v_n\|_*.\] Moreover, \(\|v_{n+1}\|_* \le \rho\), so the neighbourhood \(\{\|q - q^*\|_* \le \rho\}\) is forward-invariant.
Proof. By the triangle inequality and spectral contraction: \[\|v_{n+1}\|_* \le \|(I-\mathcal{G})v_n\|_* + \|R(v_n)\|_* \le (1-\lambda_*)\|v_n\|_* + L\|v_n\|_*^2.\] Since \(\|v_n\|_* \le \rho = \lambda_*/(4L)\): \[\|v_{n+1}\|_* \le \left(1 - \lambda_*+ \frac{\lambda_*}{4}\right)\|v_n\|_* = \left(1 - \frac{3\lambda_*}{4}\right)\|v_n\|_*.\] Forward invariance follows since \((1 - 3\lambda_*/4) < 1\). ◻
By induction, for all \(n \ge N_*\): \[\label{eq:vn95decay} \|v_n\|_* \le \left(1 - \frac{3\lambda_*}{4}\right)^{n-N_*}\|v_{N_*}\|_*.\tag{14}\]
Lemma 6. There exists \(M > 0\) such that for all \(\|v\|_* \le \rho\): \[D_{\mathrm{KL}}(q^*\|q^*+v) \le \frac{1+M\rho}{2}\,\|v\|_*^2.\]
Proof. The standard Taylor expansion of KL at \(q^*\) gives \(D_{\mathrm{KL}}(q^*\|q^*+v) = \frac{1}{2}\|v\|_*^2 + O(\|v\|_*^3)\). The cubic remainder is bounded by \(M\|v\|_*^3 \le M\rho\|v\|_*^2\) on \(\{\|v\|_* \le \rho\}\). ◻
Combining 14 and Lemma 6: \[\begin{align} D_{\mathrm{KL}}(q^*\|q_n) &\le \frac{1+M\rho}{2}\|v_n\|_*^2 \le \frac{1+M\rho}{2} \left(1-\frac{3\lambda_*}{4}\right)^{2(n-N_*)}\|v_{N_*}\|_*^2. \end{align}\] Similarly, \(D_{\mathrm{KL}}(q^*\|q_{N_*}) \ge \frac{1-M\rho}{2}\|v_{N_*}\|_*^2\) (lower bound from the same Taylor expansion), so \[\|v_{N_*}\|_*^2 \le \frac{2}{1-M\rho}\,D_{\mathrm{KL}}(q^*\|q_{N_*}).\] Setting \[C(q_0) = \frac{1+M\rho}{1-M\rho}\,D_{\mathrm{KL}}(q^*\|q_{N_*}),\] which is finite since \(M\rho < 1\) for \(\rho = \lambda_*/(4L)\) sufficiently small, completes the proof. 0◻
Hayashi [2] establishes global \(O(1/n)\) convergence for the discrete BA iteration via the Bregman-EM framework, a result that is both powerful and applicable from the very first iteration. Theorem 26 provides a complementary perspective: after the finite entry time \(N_*\), the iteration converges exponentially with an explicit rate determined by \(\lambda_*\).
The two frameworks address complementary aspects of the convergence problem: Hayashi’s [2] global \(O(1/n)\) result governs the full trajectory from any initial condition, while Theorem 26 provides an explicit exponential rate for the asymptotic phase \(n \ge N_*\). The entry time \(N_*\) and convergence factor \((1-3\lambda_*/4)^2\) are both determined by \(\lambda_*\) and the initial free energy gap, giving a complete and quantitative picture of the trajectory.
For \(X\sim\mathcal{N}(0,\sigma^2)\) with quadratic distortion and \(\beta > 1/(2\sigma^2)\), the spectral gap is \(\lambda_*= 1/(2\beta\sigma^2)\) (Theorem 23). The convergence factor is \[\left(1-\frac{3\lambda_*}{4}\right)^2 = \left(1 - \frac{3}{8\beta\sigma^2}\right)^2,\] and the entry time bound becomes \[N_*(q_0) \le \left\lceil\frac{128\,\beta^2\sigma^4 L^2 (F_\beta(q_0)-F_\beta(q^*))}{c_1^2}\right\rceil.\] As \(\beta\to 1/(2\sigma^2)^+\), the factor \((1-3\lambda_*/4)^2\to 1\) and the entry time diverges, consistent with critical slowing down at the rate-distortion threshold (Remark following Theorem 23).
This appendix provides a complete proof of Theorem 25. Throughout, we consider the continuous-time BA flow under quadratic distortion \(d(x,\hat{x})=|x-\hat{x}|^2\), Gaussian source \(X\sim\mathcal{N}(0,\sigma^2)\), and inverse temperature \(\beta>1/(2\sigma^2)\). We write \(q_t\) for the evolving reproduction law and \(q_* = \mathcal{N}(0,s^*)\) for the Gaussian fixed point of Theorem 20. The goal is to prove \(\|q_t - q_*\|_{\mathrm{TV}}\to 0\) as \(t\to\infty\).
Lemma 7 (Uniform second-moment control). Let \(q_0\) have finite second moment. Then \(\sup_{t\ge 0}\int_\mathbb{R}\hat{x}^2\,q_t(d\hat{x}) < \infty\).
Proof. Define \(V(t) = \int \hat{x}^2\,q_t(d\hat{x})\). By Lemma 3, there exist constants \(C < \infty\) and \(c > 0\) such that \(\dot{V}(t) \le C - cV(t)\). Grönwall’s inequality gives \(V(t) \le \max\{V(0),\, C/c\} < \infty\) for all \(t \ge 0\). ◻
Hence the family \(\{q_t\}_{t\ge 0}\) is tight in \(\mathcal{P}(\mathbb{R})\).
Fix the equilibrium measure \(q_* = \mathcal{N}(0,s^*)\) and let \(\{H_n\}_{n\ge 0}\) denote the orthonormal Hermite polynomial basis in \(L^2(q_*)\). Since \(q_t \ll q_*\) for all \(t > 0\), define the density ratio \(f_t = dq_t/dq_*\) and expand \[f_t = 1 + \sum_{n=1}^{\infty} a_n(t)\,H_n, \qquad a_n(t) = \langle f_t - 1,\, H_n\rangle_{L^2(q_*)}.\]
Lemma 8 (Modewise exponential decay). For every \(n \ge 3\) there exists \(C_n < \infty\) such that \(|a_n(t)| \le C_n\,\alpha^{nt}\) for all \(t \ge 0\).
Proof. By Proposition 22, the linearised BA operator acts diagonally in the Hermite basis: \(D\mathcal{T}(q_*)\,H_n = \alpha^n H_n\) with \(0 < \alpha < 1\). The nonlinear remainder is quadratic in \(q_t - q_*\) and hence lower order once the trajectory enters the local basin of \(q_*\). Variation-of-constants then gives \(a_n(t) = \alpha^{nt}a_n(0) + o(\alpha^{nt})\), from which \(|a_n(t)| \le C_n\,\alpha^{nt}\). ◻
Because the source is centred and the distortion is symmetric, the BA flow preserves zero mean: \(\int \hat{x}\,q_t(d\hat{x}) = 0\) for all \(t\), so the \(n=1\) mode vanishes identically.
Lemma 9 (Variance convergence). \(s(t) := \int \hat{x}^2\,q_t(d\hat{x}) \to s^*\) as \(t\to\infty\).
Proof. By Theorem 20, \(\dot{s}(t) = F(s(t)) + r(t)\), where \(F\) has unique stable fixed point \(s^*\) and \(r(t)\to 0\). Standard theory of asymptotically autonomous ODEs implies \(s(t)\to s^*\). ◻
Hence the \(n=2\) Hermite component also converges to equilibrium.
Lemma 10 (Weak convergence). \(q_t \Rightarrow q_*\) as \(t\to\infty\).
Proof. By Lemma 7, \(\{q_t\}\) is tight, so every sequence \(t_k\to\infty\) admits a weakly convergent subsequence \(q_{t_k}\Rightarrow\mu\). By Lemma 8, all Hermite coefficients of degree \(n \ge 3\) satisfy \(\int H_n\,d\mu = 0\). By Lemma 9, \(\int\hat{x}^2\,d\mu = s^*\). The \(n=1\) coefficient vanishes by zero-mean preservation. Hence \(\mu\) has vanishing Hermite coefficients for all \(n \ge 1\) relative to \(q_*\), which characterises \(\mu = q_*\) uniquely. Since every subsequential limit equals \(q_*\), the full family converges: \(q_t \Rightarrow q_*\). ◻
Lemma 11 (Pinsker upgrade). If \(q_t \Rightarrow q_*\) and \(D_{\mathrm{KL}}(q_t\|q_*)\to 0\), then \(\|q_t - q_*\|_{\mathrm{TV}}\to 0\).
Proof. Pinsker’s inequality gives \(\|q_t - q_*\|_{\mathrm{TV}}^2 \le 2\,D_{\mathrm{KL}}(q_t\|q_*)\). ◻
It remains to verify \(D_{\mathrm{KL}}(q_t\|q_*)\to 0\).
The free energy \(F_\beta\) is a strict Lyapunov functional (Theorem 1) with \(q_* = \arg\min F_\beta\), so \(F_\beta(q_t)\downarrow F_\beta(q_*)\). By Theorem 5, near \(q^*\) there is local quadratic equivalence \(F_\beta(q) - F_\beta(q_*) \asymp D_{\mathrm{KL}}(q\|q_*)\). Since \(q_t \Rightarrow q_*\) by Lemma 10, eventually \(q_t\) lies in this local neighbourhood, and therefore \(D_{\mathrm{KL}}(q_t\|q_*)\to 0\).
Combining Lemmas 7–11 yields \(q_t \Rightarrow q_*\) and \(D_{\mathrm{KL}}(q_t\|q_*)\to 0\), hence \(\|q_t - q_*\|_{\mathrm{TV}}\to 0\) as \(t\to\infty\). This completes the proof of Theorem 25. 0◻
Q. Wang is with the School of Information Science and Engineering, and the School of Economics and Management, Southeast University, Nanjing, China. E-mail: qiaowang@seu.edu.cn↩︎