Feedback Cycles in Exploratory Equilibria2


Abstract

Entropy regularization smooths equilibrium policies in time-inconsistent stochastic control. At low temperature, the same Gibbs response can strongly amplify errors in learned rewards and dynamics. We show that the derivative of an exploratory equilibrium is governed by a backward Volterra–parabolic resolvent. Along an aligned positive mode, a lower bound has the same exponential order. A block decomposition identifies the source of the amplification: causal paths contribute powers of \(1/\tau\), whereas a positive feedback cycle can produce exponential growth.

At fixed temperature, a local equilibrium branch is twice differentiable with respect to finite-dimensional model parameters, which yields a function-valued delta method. A bounded uniformly elliptic diffusion realizes this path–cycle distinction in every finite dimension. Closing one positive cycle changes the root-\(n\) linear-response boundary from a power law to order \(1/\log n\); along the cyclic Perron mode, right-endpoint discretization is relatively consistent exactly when \(N\tau^2\to\infty\). An affine model also gives an exact nonlinear transition at the Lambert-\(W\) temperature \(\beta T/W(\beta T\sqrt n)\). Numerical calculations illustrate these rates.

time-inconsistent stochastic control, reinforcement learning, exploratory equilibrium HJB equation, entropy regularization, Volterra operator, feedback cycles, statistical conditioning, Mittag–Leffler function

93E20, 49L20, 68T05, 60H30, 65M12, 62M05

1 Introduction↩︎

Entropy regularization replaces a point action by a relaxed policy and turns policy improvement into a Gibbs map. In continuous-time reinforcement learning (RL), this gives a smooth Hamiltonian and a tractable policy iteration. Convergence results for time-consistent exploratory HJB equations include [1][4]. Entropy annealing and robustness have also been studied for time-consistent continuous-time control [5], [6].

For a time-inconsistent objective, a policy chosen at \((t,x)\) need not remain optimal at a later time. One instead seeks an equilibrium among temporal selves, described by an extended HJB system [7][11]. That system is nonlocal: the policy at time \(t\) depends on a diagonal derivative of an auxiliary value function whose reference variables are frozen at \((t,x)\).

Huang, Yu, and Zhang [12] recently proposed an exploratory equilibrium HJB (EEHJB) system for a broad class of these problems, proved fixed-temperature convergence of policy iteration, and constructed a classical solution. A related vanishing-entropy result obtains unregularized equilibria from subsequential EEHJB limits [13]. These works address well-posedness, fixed-model policy iteration, and the zero-temperature limit, but not how learned-model or policy-evaluation errors are amplified as the temperature falls. A fixed-temperature iteration can converge quickly while its equilibrium is poorly conditioned.

Our analysis is self-contained relative to that well-posedness theory. We state the fixed-point system and its data hypotheses in [sec:model] 2 3, construct the mild evaluation map and local equilibrium branch directly, and solve the explicit models separately.

Adjacent perturbation theory covers time-consistent relaxed control [14] and equilibrium-induced values for time-inconsistent stopping [15], [16]. Nonlocal equilibrium HJB equations, equilibrium dynamic programming, and related BSDE representations are studied in [17], [18]. Relaxed equilibria and sample-based learning for time-inconsistent control and decision problems appear in [19][22]. None of these results gives a temperature-explicit derivative for a continuous-time exploratory control equilibrium and propagates it through a model estimator.

The organizing point is that low-temperature conditioning is governed not only by the local \(1/\tau\) Gibbs factor, but also by the topology and smoothing of causal returns. A future policy changes the continuation score of every earlier self. In discrete time this derivative is strictly triangular. In continuous time it becomes a backward Volterra operator with infinitely many nonzero causal powers. The series converges for each \(\tau>0\), but its norm may grow exponentially as \(\tau\downarrow0\).

Treating future decision nodes as agents with logit responses goes back to agent quantal response equilibrium in extensive-form games [23]. Static graph QRE can also exhibit Ising-type phase changes [24]. The analogy here is limited: temporal selves form a continuum coupled by a controlled diffusion, and the directed graph describes a causal derivative rather than equilibrium multiplicity.

Let \(\alpha\) denote the smoothing order of the causal kernel. The equilibrium derivative is bounded by \(C\tau^{-1}E_\alpha(\kappa T^\alpha/\tau)\). If the forcing and return kernel preserve an aligned sign, the lower bound has the same small-temperature exponent. Without that sign condition, negative feedback may damp the response.

The block form shows where repeated amplification enters. On a directed acyclic graph, the Neumann series stops at the longest path. Under aligned forcing, a positive \(q\)-cycle with total smoothing order \(\beta\) contributes the lower factor \(E_\beta(C/\tau^q)\) and forces divergence below the scale \((\log n)^{-\beta/q}\). When the corresponding one-step upper bound matches the cycle return, this scale is the linear-response boundary.

To connect conditioning with estimation, we treat a finite-dimensional coefficient family at fixed \(\tau>0\). The mild EEHJB evaluation map is \(C^2\), its causal derivative is invertible, and the resulting equilibrium branch satisfies a function-valued delta method. The parabolic argument gives the general upper scale \(e^{C/\tau^2}\); whether a state-dependent model attains this rate remains open.

Two exact models show what the causal rates mean. In a bounded uniformly elliptic diffusion of arbitrary finite dimension, the tied-policy response is \[\frac{v}{\tau}\exp\left\{\left(-\nu I+\frac{v}{\tau} K\right)(T-t)\right\}.\] For \(K\ge0\), a graph of longest path \(L\) has susceptibility \(\Theta(\tau^{-(L+1)})\), while a positive cycle produces an \(e^{C/\tau}\) factor along a Perron direction. The model also has a nonlinear counterpart: exponentially small signals vanish on a DAG, whereas cooperative feedback selects extreme actions under the positive-row-sum, support, and short-horizon conditions of 5. In the affine model, root-\(n\) reward noise has the exact critical temperature \[\tau_n^\star=\frac{\beta T}{W(\beta T\sqrt n)} \sim\frac{2\beta T}{\log n}.\] Here \(W\) denotes the principal real branch of the Lambert \(W\) function. At this scale the initial policy has a nondegenerate random limit although every fixed later policy converges to the reference mixture.

Time discretization retains this graph dependence. Along a cyclic Perron mode, a uniform \(N\)-step rule has vanishing relative error exactly when \(N\tau^2\to\infty\). For a fixed DAG, \(N\to\infty\) suffices without coupling \(N\) to \(\tau\). The general stability and delta-method results keep \(\tau>0\) fixed; joint small-temperature limits are stated only when an explicit bound or exact model supports them.

The paper follows this causal chain. 2 derives the Gibbs fixed point from infinitesimal spikes. 3 proves the abstract resolvent and graph dichotomy. 4 establishes EEHJB stability and the fixed-temperature delta method. 5 realizes the sharp path–cycle and nonlinear statistical transitions, and 6 closes with discretization and reproducible numerical illustrations.

2 Time-inconsistent relaxed control and the EEHJB score map↩︎

Fix a horizon \(T>0\), state dimension \(d\), and a compact metric action space \(\mathsf A\) equipped with a reference probability measure \(\mu\) of full support. A Markov relaxed policy is a kernel \(\pi(t,x,da)\) absolutely continuous with respect to \(\mu\). Put \[\bar b^\pi(t,x)=\int_\mathsf Ab(t,x,a)\pi(t,x,da), \qquad \bar r_{\tau,y}^\pi(t,x)=\int_\mathsf Ar(\tau,y,t,x,a)\pi(t,x,da).\] Finite action spaces are included: if \(\mu\) gives positive mass to each action, every policy has a density with respect to \(\mu\), and the integrals and relative entropy below are finite sums. For a known, uncontrolled diffusion matrix \(\sigma(t,x)\), the exploratory state is \[\label{eq:state-sde} dX_s^\pi=\bar b^\pi(s,X_s^\pi)\,ds+\sigma(s,X_s^\pi)\,dW_s, \qquad X_t^\pi=x.\tag{1}\] The self whose reference time-state is \((\tau,y)\) evaluates a continuation from \((t,x)\) by \[\begin{align} V^\pi(\tau,t,y,x) :=\mathbb{E}_{t,x}\bigg[&\int_t^T \Big\{r(\tau,y,s,X_s^\pi,A_s) -\tau_e\log\frac{d\pi(s,X_s^\pi)}{d\mu}(A_s)\Big\}\,ds\notag\\ &+F(\tau,y,X_T^\pi)\bigg]. \label{eq:aux-value} \end{align}\tag{2}\] Here, conditionally on the state, \(A_s\) has distribution \(\pi(s,X_s^\pi,\cdot)\). We write the entropy temperature as \(\tau_e>0\) temporarily to distinguish it from the reference-time variable; from 3 onward it is denoted simply by \(\tau\). Nonexponential discounting can be absorbed into \(r\).

The policy-evaluation equation for a fixed admissible \(\pi\) is the family of linear parabolic PDEs \[\begin{align} \partial_tV^\pi+\tfrac12\operatorname{tr} (\sigma\sigma^\top D^2_xV^\pi) +\bar b^\pi\cdot D_xV^\pi+\bar r_{\tau,y}^\pi -\tau_e\operatorname{KL}(\pi(t,x)\Vert\mu)&=0, \label{eq:policy-evaluation-pde}\\ V^\pi(\tau,T,y,x)&=F(\tau,y,x). \notag \end{align}\tag{3}\] The variables \((\tau,y)\) are frozen parameters in this PDE. The acting self at \((t,x)\) sets \((\tau,y)=(t,x)\) but uses the flow-state partial derivative \(D_xV^\pi(t,t,x,x)\), not the total derivative of the diagonal map. Its action score is \[\label{eq:score-map} Q^\pi(t,x,a)=b(t,x,a)\cdot D_xV^\pi(t,t,x,x) +r(t,x,t,x,a).\tag{4}\] The Gibbs variational identity gives the unique pointwise best response \[\label{eq:gibbs-map} \mathcal{G}_{\tau_e}(q)(da) =\frac{\exp(q(a)/\tau_e)\mu(da)}{\int_\mathsf A\exp(q(a')/\tau_e)\mu(da')}.\tag{5}\] Define the best-response operator \[\label{eq:equilibrium-fixed-point} \mathcal{T}_{\tau_e,M}(\pi):=\mathcal{G}_{\tau_e}(\mathcal{Q}_M(\pi)), \qquad M=(b,r,F).\tag{6}\] A regularized Markov equilibrium is a fixed point \(\pi^\star=\mathcal{T}_{\tau_e,M}(\pi^\star)\); when it is unique we write \(\Phi_{\tau_e}(M)=\pi^\star\). Here \(\mathcal{Q}_M\) means: solve 3 and form 4 . This is a compact score-map representation of the EEHJB system. The stability section adds a terminal statistic and a nonlinear function \(G\) and states that extended system explicitly.

The next result anchors this fixed-point formulation in the infinitesimal-spike equilibrium criterion. Its local hypotheses justify the first-order spike limit; the conclusion identifies the Gibbs fixed point as exactly the no-profitable-spike condition.

Fix \((t,x)\in[0,T)\times\mathbb{R}^d\). Let \(\rho\ll\mu\) satisfy \(\operatorname{KL}(\rho\Vert\mu)<\infty\), and set \[\pi^{\epsilon,\rho}(s,z)= \begin{cases} \rho, &t\le s<t+\epsilon,\\ \pi(s,z), &t+\epsilon\le s\le T. \end{cases}\] Let \(J_{t,x}(\eta):=V^\eta(t,t,x,x)\) denote the payoff of the self whose reference variables are frozen at \((t,x)\). Suppose that, for some \(\delta>0\), \(u(s,z):=V^\pi(t,s,x,z)\) belongs to \(C^{1,2}([t,t+\delta]\times\mathbb{R}^d)\) and satisfies 3 on this strip. Assume that \(b\), \(\sigma\), and \(r\) are continuous at \((t,x)\), uniformly in \(a\in\mathsf A\), that the spike SDE is locally well posed, and that \(u\) and its derivatives in Itô’s formula have polynomial growth controlled by corresponding moments of the spike process. Assume also that \[(s,z)\longmapsto \big(\bar b^\pi(s,z),\bar r_{t,x}^\pi(s,z), \operatorname{KL}(\pi(s,z)\Vert\mu)\big)\] is right-continuous at \((t,x)\). Assume \(\operatorname{KL}(\pi(t,x)\Vert\mu)<\infty\), \(Q^\pi(t,x,\cdot)\in C(\mathsf A)\), and, for some \(\epsilon_0\in(0,\delta]\), the entire bracketed integrand in 8 is uniformly integrable over \(0<\epsilon<\epsilon_0\) and \(s\in[t,t+\epsilon]\). Then \[\label{eq:spike-gain} \begin{align} \lim_{\epsilon\downarrow0} \frac{J_{t,x}(\pi^{\epsilon,\rho})-J_{t,x}(\pi)}{\epsilon} &=\Gamma_{t,x}^\pi(\rho)-\Gamma_{t,x}^\pi(\pi(t,x)),\\ \Gamma_{t,x}^\pi(\rho) &=\int_\mathsf AQ^\pi(t,x,a)\rho(da)-\tau_e\operatorname{KL}(\rho\Vert\mu). \end{align}\tag{7}\] If these hypotheses hold for every finite-entropy \(\rho\) at each point, then \(\pi\) admits no infinitesimal spike with positive limit in 7 if and only if it satisfies 6 . The maximizer is unique up to \(\mu\)-null sets.

Proof. Let \(X^{\epsilon,\rho}\) be the state process under the spike policy, and write \[\begin{align} \bar b^\rho(s,z)&=\int_\mathsf Ab(s,z,a)\rho(da),& \bar r_{t,x}^\rho(s,z)&=\int_\mathsf Ar(t,x,s,z,a)\rho(da),\\ k^\pi(s,z)&=\operatorname{KL}(\pi(s,z)\Vert\mu).&& \end{align}\] Since the policy returns to \(\pi\) at \(t+\epsilon\), the Markov property gives \[\begin{align} J_{t,x}(\pi^{\epsilon,\rho}) =\mathbb{E}\bigg[&\int_t^{t+\epsilon} \big\{\bar r_{t,x}^\rho(s,X_s^{\epsilon,\rho}) -\tau_e\operatorname{KL}(\rho\Vert\mu)\big\}\,ds\\ &+u(t+\epsilon,X_{t+\epsilon}^{\epsilon,\rho})\bigg]. \end{align}\] For \(0<\epsilon<\delta\), apply Itô’s formula to \(u(s,X_s^{\epsilon,\rho})\) and use the policy-evaluation equation for \(u\). Because the diffusion matrix in 1 is independent of the action, the second-order terms in the two generators cancel exactly. We obtain \[\begin{align} J_{t,x}(\pi^{\epsilon,\rho})-u(t,x) =\mathbb{E}\int_t^{t+\epsilon}\bigg[& (\bar b^\rho-\bar b^\pi)\cdot D_zu +(\bar r_{t,x}^\rho-\bar r_{t,x}^\pi)\notag\\ &-\tau_e\big\{\operatorname{KL}(\rho\Vert\mu) -k^\pi\big\} \bigg](s,X_s^{\epsilon,\rho})\,ds . \label{eq:exact-spike-difference} \end{align}\tag{8}\] Here \(k^\pi\) is evaluated at \((s,X_s^{\epsilon,\rho})\), whereas \(\operatorname{KL}(\rho\Vert\mu)\) is constant on the spike interval. A stopped Burkholder–Davis–Gundy estimate, recorded in the supplement, gives \(\sup_{t\le s\le t+\epsilon}|X_s^{\epsilon,\rho}-x|\to0\) in probability. Right-continuity identifies the limit of the bracketed integrand, and its assumed uniform integrability permits passage to the limit by Vitali’s theorem in 8 . Dividing by \(\epsilon\) gives 7 , now with the baseline drift, reward, and entropy terms displayed explicitly.

Set \(Z=\int_\mathsf Ae^{Q^\pi(t,x,a)/\tau_e}\mu(da)\). Compactness of \(\mathsf A\) and continuity of \(Q^\pi(t,x,\cdot)\) give \(0<Z<\infty\). If \(G=\mathcal{G}_{\tau_e}(Q^\pi(t,x,\cdot))\), direct calculation yields \[\Gamma_{t,x}^\pi(\rho) =\tau_e\log Z-\tau_e\operatorname{KL}(\rho\Vert G).\] Thus \(G\) is the unique maximizer. Requiring the derivative in 7 to be nonpositive for every finite-entropy \(\rho\) is equivalent to \(\pi(t,x)=G\). Applying this pointwise proves the fixed-point equivalence. ◻

This equivalence concerns the local equilibrium criterion defined by infinitesimal spikes. It does not assert precommitment optimality or the strong finite-spike equilibrium property.

A perturbation of \(\pi(s,\cdot)\) at \(s>t\) changes \(V^\pi(t,t,\cdot,\cdot)\), hence the score and policy at \(t\). Causality prevents a policy perturbation at an earlier time from changing a later score. Thus \(D_\pi\mathcal{Q}_M\) is a backward Volterra operator.

3 The causal Volterra resolvent↩︎

We first separate the causal argument from the PDE estimates. Let \(X\) and \(Y\) be Banach spaces for policy tangents and centered action scores at one time, respectively, and put \(\mathcal{X}=C([0,T];X)\) and \(\mathcal{Y}=C([0,T];Y)\) with supremum norms. In finite actions, \(X\) is the product of simplex tangent spaces with the \(\ell^1\) norm and \(Y\) carries the oscillation norm. For a signed measure \(\nu\), our total variation convention is \(\left\lVert \nu\right\rVert_{\rm TV}=|\nu|(\mathsf A)\), so the distance between two probability measures is at most two. Let \(\mathsf H\) be a Banach space of primitive model directions.

For \(\alpha>0\), define the right-sided Riemann–Liouville integral \[\label{eq:fractional-integral} (I_{T-}^\alpha f)(t) :=\frac{1}{\Gamma(\alpha)}\int_t^T(s-t)^{\alpha-1}f(s)\,ds.\tag{9}\] We use the convention \(I_{T-}^0=\mathrm{Id}\). The fractional integrals satisfy \(I_{T-}^\alpha I_{T-}^\gamma=I_{T-}^{\alpha+\gamma}\) and, for every integer \(m\ge0\), \[\label{eq:fractional-constant} (I_{T-}^{m\alpha}{\boldsymbol{1}})(t) =\frac{(T-t)^{m\alpha}}{\Gamma(m\alpha+1)}.\tag{10}\] The Mittag–Leffler function is \[\label{eq:mittag-leffler} E_\alpha(z):=\sum_{m=0}^\infty\frac{z^m}{\Gamma(m\alpha+1)}.\tag{11}\] See [25][27] for the fractional-integral and Volterra background. For positive \(z\) and \(0<\alpha<2\), \(E_\alpha(z)=\alpha^{-1}e^{z^{1/\alpha}}+O(z^{-1})\); also \(E_1(z)=e^z\) and \(E_2(z)=\cosh\sqrt z\). For the block argument we also use \(\log E_\alpha(z)\sim z^{1/\alpha}\) on the positive axis for every \(\alpha>0\). A Stirling–Laplace proof of this all-orders relation is recorded in the supplement.

Lemma 1 (Gibbs differential and global Lipschitz bound). For \(q,h\in L^\infty(\mu)\) and \(\tau>0\), \[\label{eq:gibbs-derivative} D\mathcal{G}_\tau(q)[h](da) =\frac{\mathcal{G}_\tau(q)(da)}{\tau} \left(h(a)-\int h\,d\mathcal{G}_\tau(q)\right).\qquad{(1)}\] Moreover, for all \(q,\tilde{q}\), \[\label{eq:gibbs-global-lipschitz} \left\lVert \mathcal{G}_\tau(q)-\mathcal{G}_\tau(\tilde{q})\right\rVert_{\rm TV} \le \min\left\{2,\frac{\operatorname{osc}(q-\tilde{q})}{\tau}\right\}, \qquad \operatorname{osc}f:=\operatorname*{ess\,sup}_{\mu}f -\operatorname*{ess\,inf}_{\mu}f.\qquad{(2)}\] Here \(D\mathcal{G}_\tau(q)\) is the Fréchet derivative into the finite signed measures equipped with the total-variation norm.

Proof. Differentiating the normalized exponential gives ?? . For the global bound, interpolate \(q_u=\tilde{q}+u(q-\tilde{q})\), integrate the derivative from \(u=0\) to one, and use \(\int|h-\int h\,d\rho|d\rho\le\operatorname{osc}h\). The bound by two is the diameter of the probability simplex in our convention. ◻

At a reference-tied model \(M^\circ\), the equilibrium policy equals a reference \(\mu\) and every acting-self score is constant over actions. For \(h\in\mathsf H\), let \(c[h]\in\mathcal{Y}\) be the direct score derivative, with the future policy held fixed. Thus \(c:\mathsf H\to\mathcal{Y}\) is a bounded linear map. Let \(\mathcal{K}:\mathcal{X}\to\mathcal{Y}\) be the future-policy-to-current-score derivative and let \(\mathcal{S}_t:Y\to X\) be the centered covariance operator in ?? at the reference policy. Its bounded pointwise lift \(\mathcal{S}:\mathcal{Y}\to\mathcal{X}\) is \((\mathcal{S}y)(t)=\mathcal{S}_t y(t)\). The equilibrium tangent \(u=D_M\Phi_\tau(M^\circ)[h]\) formally satisfies \[\label{eq:abstract-tangent} u(t)=\frac{1}{\tau}\mathcal{S}_t\{c[h](t)+(\mathcal{K}u)(t)\}.\tag{12}\]

Assumption 1 (Fractional causal influence). There are \(\alpha\in(0,2]\), \(\kappa\ge0\), and \(C_c<\infty\) such that, for every \(v\in\mathcal{X}\), every \(h\in\mathsf H\), and every \(t\in[0,T]\), \[\begin{align} \left\lVert \mathcal{S}_t(\mathcal{K}v)(t)\right\rVert_X &\le \kappa(I_{T-}^\alpha\left\lVert v(\cdot)\right\rVert_X)(t), \label{eq:fractional-kernel-bound}\\ \sup_t\left\lVert \mathcal{S}_tc[h](t)\right\rVert_X&\le C_c\left\lVert h\right\rVert_{\mathsf H}. \label{eq:direct-bound} \end{align}\] {#eq: sublabel=eq:eq:fractional-kernel-bound,eq:eq:direct-bound}

For a bounded Volterra kernel, \(\alpha=1\). A kernel that vanishes linearly on the diagonal has \(\alpha=2\). The \(D_x\) estimate for a uniformly elliptic parabolic semigroup has the weak singularity \((s-t)^{-1/2}\), hence \(\alpha=1/2\).

Theorem 1 (Fractional Volterra susceptibility). Under 1, 12 has a unique solution for every \(\tau>0\), given by the convergent causal resolvent \[\label{eq:volterra-neumann} u=\frac{1}{\tau}\sum_{m=0}^\infty \tau^{-m}(\mathcal{S}\mathcal{K})^m\mathcal{S}c[h].\qquad{(3)}\] For \(0\le t\le T\), \[\label{eq:mittag-bound} \left\lVert u(t)\right\rVert_X \le\frac{C_c\left\lVert h\right\rVert_{\mathsf H}}{\tau} E_\alpha\!\left(\frac{\kappa(T-t)^\alpha}{\tau}\right).\qquad{(4)}\] In particular, as \(\tau\downarrow0\), the upper bound has exponential scale \[\label{eq:mittag-asymptotic-scale} \frac{1}{\tau}\exp\left\{ \kappa^{1/\alpha}(T-t)\tau^{-1/\alpha}\right\}.\qquad{(5)}\]

Proof. Write \(A=\mathcal{S}\mathcal{K}\). Repeated use of ?? , the semigroup identity for fractional integrals, and 10 gives \[\left\lVert A^m\mathcal{S}c[h](t)\right\rVert_X \le C_c\left\lVert h\right\rVert_{\mathsf H}\, \frac{\kappa^m(T-t)^{m\alpha}}{\Gamma(m\alpha+1)}.\] The majorant series is finite for every \(\tau>0\) and sums to ?? ; hence ?? converges absolutely in \(\mathcal{X}\) and solves 12 . If \(z=\tau^{-1}Az\), then for every \(m\ge1\), \[\|z(t)\|_X\le \|z\|_\infty \frac{(\kappa/\tau)^m(T-t)^{m\alpha}}{\Gamma(m\alpha+1)}.\] The right-hand side tends to zero, which proves uniqueness. When \(\kappa>0\) and \(t<T\), the positive-axis Mittag–Leffler asymptotic yields ?? . ◻

The norm estimate in 1 need not be sharp without a sign condition. For example, the scalar operator \[(Av)(t)=-\kappa\int_t^T v(s)\,ds\] satisfies the order-one kernel bound, but the response to unit forcing is \(\tau^{-1}e^{-\kappa(T-t)/\tau}\). Negative feedback damps the perturbation. A matching lower bound requires a direction in which the causal feedback keeps its sign.

Let \(X_+\subset X\) be a closed convex pointed cone. Write \(x\preceq y\) when \(y-x\in X_+\) and use the pointwise order on \(\mathcal{X}\). For policy tangents, \(X_+\) may be the ray generated by one action contrast rather than the usual order on signed measures.

Theorem 2 (Aligned-mode lower bound). Suppose the conditions of 1 hold. Set \[A=\mathcal{S}\mathcal{K},\qquad w_h(t)=\mathcal{S}_tc[h](t).\] Suppose that for one oriented direction \(h\in\mathsf H\setminus\{0\}\) there are constants \(c_->0\) and \(0<\kappa_-\le\kappa\), a path \(\psi\in C([0,T];X_+)\), and positive functionals \(\ell_t\in X^*\) with \(\left\lVert \ell_t\right\rVert_{X^*}\le1\) and \(q(t):=\ell_t(\psi(t))>0\). Assume that these objects do not depend on \(\tau\) and that \[\begin{align} A\mathcal{X}_+&\subset\mathcal{X}_+, \label{eq:A-positive}\\ w_h(t)&\succeq c_-\left\lVert h\right\rVert_{\mathsf H}\,\psi(t), \label{eq:forcing-aligned}\\ A(f\psi)(t)&\succeq \kappa_-\bigl(I_{T-}^{\alpha}f\bigr)(t)\psi(t) \label{eq:mode-subeigen} \end{align}\] {#eq: sublabel=eq:eq:A-positive,eq:eq:forcing-aligned,eq:eq:mode-subeigen} for every nonnegative \(f\in C([0,T])\). Then \[\label{eq:matched-mittag-bound} \frac{c_-q(t)\left\lVert h\right\rVert_{\mathsf H}}{\tau} E_\alpha\!\left(\frac{\kappa_-(T-t)^\alpha}{\tau}\right) \le \left\lVert u(t)\right\rVert_X \le \frac{C_c\left\lVert h\right\rVert_{\mathsf H}}{\tau} E_\alpha\!\left(\frac{\kappa(T-t)^\alpha}{\tau}\right).\qquad{(6)}\] For each fixed \(t<T\), \[\begin{align} \kappa_-^{1/\alpha}(T-t) &\le \liminf_{\tau\downarrow0}\tau^{1/\alpha} \log\frac{\tau\left\lVert u(t)\right\rVert_X}{\left\lVert h\right\rVert_{\mathsf H}} \notag\\ &\le \limsup_{\tau\downarrow0}\tau^{1/\alpha} \log\frac{\tau\left\lVert u(t)\right\rVert_X}{\left\lVert h\right\rVert_{\mathsf H}} \le \kappa^{1/\alpha}(T-t). \label{eq:matched-log-rate} \end{align}\qquad{(7)}\] If \(\kappa_-=\kappa\), the upper exponential rate is attained.

Proof. Let \(f_m=I_{T-}^{m\alpha}{\boldsymbol{1}}\). Induction gives \[\label{eq:positive-iterate-lower} A^mw_h(t)\succeq c_-\left\lVert h\right\rVert_{\mathsf H}\,\kappa_-^m f_m(t)\psi(t),\qquad m\ge0.\tag{13}\] The case \(m=0\) is ?? . If it holds at \(m\), positivity of \(A\) and ?? give \[A^{m+1}w_h \succeq c_-\left\lVert h\right\rVert_{\mathsf H}\,\kappa_-^{m+1} I_{T-}^{\alpha}f_m\,\psi =c_-\left\lVert h\right\rVert_{\mathsf H}\,\kappa_-^{m+1}f_{m+1}\psi.\] Insert 13 into ?? . Both series converge uniformly and the cone is closed, hence \[u(t)\succeq \frac{c_-\left\lVert h\right\rVert_{\mathsf H}}{\tau}\psi(t) E_\alpha\!\left(\frac{\kappa_-(T-t)^\alpha}{\tau}\right).\] Apply \(\ell_t\) and use \(\left\lVert \ell_t\right\rVert\le1\). This proves the lower bound; the upper bound is ?? . The positive-axis Mittag–Leffler asymptotic gives ?? . ◻

Corollary 1 (Fractional statistical resolution scale). Under the hypotheses of 2, suppose \(\kappa_-=\kappa>0\). Fix \(t<T\), let \(h_n=n^{-1/2}h\), and denote its tangent response by \(u_{\tau}[h_n]\). Put \[c_\alpha=2^\alpha\kappa(T-t)^\alpha.\] If \(\tau_n(\log n)^\alpha\to c\in(0,\infty)\), then \[\left\lVert u_{\tau_n}[h_n](t)\right\rVert_X\longrightarrow \begin{cases} 0, &c>c_\alpha,\\ \infty, &0<c<c_\alpha. \end{cases}\] Thus the root-\(n\) linear resolution boundary has order \((\log n)^{-\alpha}\).

Proof. By ?? and the positive-axis asymptotic, \[\log\left\lVert u_{\tau_n}[h_n](t)\right\rVert_X =-\frac{1}{2}\log n+\kappa^{1/\alpha}(T-t)\tau_n^{-1/\alpha} +O(\log\log n).\] The coefficient of \(\log n\) is negative for \(c>c_\alpha\) and positive for \(c<c_\alpha\). ◻

For \(\kappa>0\) and fixed \(t<T\), consider the canonical balance \[\tau^{-1}E_\alpha\!\left(\frac{\kappa(T-t)^\alpha}{\tau}\right)=\sqrt n,\] and set \(a=\kappa^{1/\alpha}(T-t)\). With \(W\) as above, its asymptotic solution is \[\label{eq:fractional-critical-temperature} \bar\tau_n= \left[ \frac{a}{\alpha W\!\left(a\alpha^{1/\alpha-1}n^{1/(2\alpha)}\right)} \right]^\alpha\{1+o(1)\} \sim\left(\frac{2a}{\log n}\right)^\alpha.\tag{14}\] This is a linearized resolution boundary. A nonlinear phase law also requires control of the small-temperature remainder; 8 supplies that control in the affine model.

To identify the relevant return channels, decompose \(X=X_1\times\cdots\times X_d\), write \(A:=\mathcal{S}\mathcal{K}\in\mathcal{L}(\mathcal{X})\) and \(\mathsf C:=\mathcal{S}c\in\mathcal{L}(\mathsf H,\mathcal{X})\), and set \(A=(A_{ij})\). Give \(A\) the directed graph with an edge \(j\to i\) when \(A_{ij}\ne0\). Directed walks index the blocks of \(A^m\) [28], allowing the Neumann series to distinguish finite causal paths from repeated cyclic returns.

Theorem 3 (Block feedback dichotomy). Let \(D_\tau h\) denote the solution of \[\label{eq:block-tangent} D_\tau h=\frac{1}{\tau}\{\mathsf C h+A D_\tau h\}.\qquad{(8)}\] Thus \(D_\tau:\mathsf H\to\mathcal{X}\). The block operators \(A\) and \(\mathsf C\), their graph, and every cone, vector, functional, and constant in the hypotheses below are independent of \(\tau\).

  1. If the block graph is acyclic with longest directed path \(L\), then \(A^{L+1}=0\) and \[\label{eq:dag-block-resolvent} D_\tau=\frac{1}{\tau}\sum_{m=0}^{L}\tau^{-m}A^m\mathsf C.\qquad{(9)}\] If \(\mathsf C=0\), then \(D_\tau=0\). Otherwise, let \(m_\star=\max\{m\le L:A^m\mathsf C\ne0\}\). Then \[\label{eq:dag-block-order} \|D_\tau\|_{\mathcal{L}(\mathsf H,\mathcal{X})}=\tau^{-(m_\star+1)} \{\|A^{m_\star}\mathsf C\|_{\mathcal{L}(\mathsf H,\mathcal{X})}+o(1)\}.\qquad{(10)}\] In particular, the order is \(\tau^{-(L+1)}\) when the forcing excites a nonzero longest path.

  2. Suppose each \(X_i\) has a closed pointed cone, every block \(A_{ij}\) is positive, and the causal Neumann series for \(\tau^{-1}A\) converges in \(\mathcal{X}\) for every \(\tau>0\). Consider a directed cycle \[i_0\to i_1\to\cdots\to i_{q-1}\to i_0\] with return operator \[B=A_{i_0i_{q-1}}A_{i_{q-1}i_{q-2}}\cdots A_{i_1i_0}.\] Assume \(\mathsf C h_\star\) belongs to the product cone for some \(h_\star\in\mathsf H\setminus\{0\}\). Suppose there are \(c_\star,g,\beta>0\), a path \(\psi\in C([0,T];X_{i_0,+})\), and positive functionals \(\ell_t\) of norm at most one, such that, for every nonnegative \(f\in C([0,T])\), \[\begin{align} (\mathsf C h_\star)_{i_0}(t)&\succeq c_\star\|h_\star\|_{\mathsf H}\psi(t), \label{eq:cycle-forcing}\\ B(f\psi)(t)&\succeq g(I_{T-}^{\beta}f)(t)\psi(t), \label{eq:cycle-return} \end{align}\] {#eq: sublabel=eq:eq:cycle-forcing,eq:eq:cycle-return} and \(\varrho(t):=\ell_t(\psi(t))>0\). Then, for \(t<T\), \[\label{eq:cycle-mittag-lower} \|(D_\tau h_\star)_{i_0}(t)\| \ge\frac{c_\star\varrho(t)\|h_\star\|_{\mathsf H}}{\tau} E_\beta\!\left(\frac{g(T-t)^\beta}{\tau^q}\right).\qquad{(11)}\] Consequently, \[\label{eq:cycle-exponential-rate} \liminf_{\tau\downarrow0}\tau^{q/\beta} \log\{\tau\|(D_\tau h_\star)_{i_0}(t)\|\} \ge g^{1/\beta}(T-t).\qquad{(12)}\]

Proof. The \((i,j)\) block of \(A^m\) is a sum of operator products indexed by directed walks of length \(m\) from \(j\) to \(i\). An acyclic graph has no walk longer than \(L\), so \(A^{L+1}=0\) and the Neumann series gives ?? . Multiplying by \(\tau^{m_\star+1}\) leaves \(A^{m_\star}\mathsf C\) plus terms that vanish in operator norm, which proves ?? .

For the cycle, positivity implies that every Neumann term is nonnegative. The \((i_0,i_0)\) block of \(A^q\) contains \(B\); for every product-cone-valued \(z\), the remaining walk terms are positive and hence \((A^qz)_{i_0}\succeq Bz_{i_0}\). Thus induction and ?? –?? give \[(A^{qm}\mathsf C h_\star)_{i_0} \succeq c_\star\|h_\star\|_{\mathsf H}g^m I_{T-}^{m\beta}{\boldsymbol{1}}\,\psi .\] Keeping the terms indexed by \(qm\) in the causal series and applying \(\ell_t\) proves ?? . The logarithmic positive-axis asymptotic \(\log E_\beta(x)\sim x^{1/\beta}\), valid for every \(\beta>0\), gives ?? . ◻

Under the positivity hypotheses above, the return condition can be checked edge by edge. Let \(i_q=i_0\), choose paths \(\psi_r\in C([0,T];X_{i_r,+})\) with \(\psi_q=\psi_0=: \psi\), and let \(a_r,\alpha_r>0\). If, for every nonnegative \(f\in C([0,T])\), \[\label{eq:edgewise-fractional-lower} A_{i_{r+1}i_r}(f\psi_r) \succeq a_r I_{T-}^{\alpha_r}f\,\psi_{r+1}, \qquad r=0,\ldots,q-1,\tag{15}\] then ?? holds with \(g=\prod_{r=0}^{q-1}a_r\) and \(\beta=\sum_{r=0}^{q-1}\alpha_r\).

Corollary 2 (Cycle resolution scale). Under the cyclic hypotheses of 3, fix \(t<T\), put \(H=T-t\), and let \(h_n=n^{-1/2}h_\star\). If \(\tau_n(\log n)^{\beta/q}\to\vartheta\in(0,\infty)\), then \[\label{eq:cycle-resolution-lower} \liminf_{n\to\infty} \frac{\log\|(D_{\tau_n}h_n)_{i_0}(t)\|}{\log n} \ge-\frac{1}{2}+g^{1/\beta}H\vartheta^{-q/\beta}.\qquad{(13)}\] Thus the response diverges when \[\label{eq:cycle-resolution-critical-lower} \vartheta<2^{\beta/q}g^{1/q}H^{\beta/q}.\qquad{(14)}\] If, in addition, 1 holds with \(\alpha=\beta/q\in(0,2]\) and \(g=\kappa^q\), then the matching upper bound shows that the response vanishes when the inequality in ?? is reversed. The root-\(n\) linear-response boundary is therefore \((\log n)^{-\beta/q}\).

Proof. Apply the logarithmic Mittag–Leffler asymptotic to ?? ; the exterior factor contributes \(\log(\tau_n^{-1})=O(\log\log n)=o(\log n)\). The upper half follows from ?? with \(\alpha=\beta/q\). ◻

The ratio \[\label{eq:cycle-mean-order} \bar\alpha_C=\frac{1}{q_C}\sum_{e\in C}\alpha_e\tag{16}\] is the mean causal smoothing order of a positive cycle \(C\). Among several cycles, the smallest \(\bar\alpha_C\) gives the strongest power of \(1/\tau\) in the exponential; ties are broken by the geometric return gain. Among the lower bounds certified by 15 , suppose a mode \(\psi_i\) is fixed at each vertex and every edge bound uses these same vertex modes. Finding the strongest guaranteed exponent is then a minimum-mean-cycle problem in the sense of [29]. This statement does not exclude another operator component from producing a larger full-resolvent norm. For edge kernels bounded below by positive constants along the cycle, \(\alpha_e=1\), and the cycle guarantees inverse-log linear-response divergence below its cycle constant; a matching upper bound identifies the full boundary.

The same causal iteration gives a nonlinear bound for model and numerical error. For two policies define \(d(t)=\sup_{x}\left\lVert \pi(t,x)-\tilde{\pi}(t,x)\right\rVert_{\rm TV}\).

Theorem 4 (Nonlinear model-and-residual stability). Let \(\pi\) be an exact causal Gibbs equilibrium for \(M\). Let \(\eta\in L^\infty(0,T)\) be nonnegative, and let \(\tilde{\pi}\) have fixed-point residual at most \(\eta(t)\) for \(\widetilde{M}\): \[\label{eq:fixed-point-residual} \sup_x\left\lVert \tilde{\pi}(t,x)- \mathcal{G}_\tau(\mathcal{Q}_{\widetilde{M}}(\tilde{\pi})(t,x,\cdot))\right\rVert_{\rm TV} \le\eta(t).\qquad{(15)}\] Suppose the score maps obey the following bound pairwise for every two candidates in the class, for some \(\alpha>0\), \(L_M,\kappa\ge0\), and primitive distance \(\varepsilon=d(M,\widetilde{M})\): \[\label{eq:nonlinear-score-bound} \sup_x\operatorname{osc}\{\mathcal{Q}_M(\pi)-\mathcal{Q}_{\widetilde{M}}(\tilde{\pi})\}(t,x,\cdot) \le L_M\varepsilon+\kappa(I_{T-}^\alpha d)(t).\qquad{(16)}\] Then \[\label{eq:nonlinear-stability-bound} d(t)\le \left(\left\lVert \eta\right\rVert_\infty+\frac{L_M\varepsilon}{\tau}\right) E_\alpha\!\left(\frac{\kappa(T-t)^\alpha}{\tau}\right).\qquad{(17)}\] In particular, exact equilibria are unique within any class satisfying ?? with \(\varepsilon=0\).

The constants \(L_M\) and \(\kappa\) are understood to be uniform over any temperature range for which the displayed dependence is used. The quantity \(\eta\) is a policy fixed-point residual. A sup-norm score residual \(\rho\) first passes through the Gibbs map and contributes at most \(2\rho/\tau\) to \(\eta\) (or \(\rho/\tau\) when \(\rho\) denotes score oscillation).

Proof. The triangle inequality, ?? , and 1 give \[d(t)\le\eta(t)+\frac{L_M\varepsilon}{\tau} +\frac{\kappa}{\tau}(I_{T-}^\alpha d)(t).\] Put \(a=\|\eta\|_\infty+L_M\varepsilon/\tau\). Since \(d\le2\), \(N\) iterations give \[d(t)\le a\sum_{m=0}^{N-1} \frac{[\kappa(T-t)^\alpha/\tau]^m}{\Gamma(m\alpha+1)} +2\frac{[\kappa(T-t)^\alpha/\tau]^N}{\Gamma(N\alpha+1)}.\] The last term tends to zero as \(N\to\infty\), which proves ?? . If both the model discrepancy and the residual vanish, the same estimate gives \(d(t)=0\) for every \(t\). ◻

For exact equilibria, so that \(\eta_n=0\), if \(\varepsilon_n=O_{\mathbb{P}}(n^{-1/2})\), a sufficient condition for the model-error bound in ?? to vanish in probability is \[\label{eq:sufficient-annealing} \frac{1}{\sqrt n\,\tau_n} E_\alpha\!\left(\frac{\kappa T^\alpha}{\tau_n}\right)\longrightarrow0.\tag{17}\] For an approximate fixed point one must additionally require \[\|\eta_n\|_\infty E_\alpha\!\left(\frac{\kappa T^\alpha}{\tau_n}\right) \longrightarrow0\] in probability, or deterministically when the residual is deterministic. When \(\kappa>0\), the exponential term gives the critical order \((\log n)^{-\alpha}\). More precisely, a sufficient condition is \[\label{eq:fractional-log-barrier} \liminf_{n\to\infty}\tau_n(\log n)^\alpha >2^\alpha\kappa T^\alpha.\tag{18}\] If \(\kappa=0\), there is no logarithmic barrier and \((\sqrt n\,\tau_n)^{-1}\to0\) is the corresponding sufficient condition. This is only a sufficient general law: stable action gaps can make the Gibbs covariance shrink with \(\tau\), whereas tied modes can attain the full amplification. The exact construction in 5 proves a sharp nonlinear law for \(\alpha=1\).

4 EEHJB stability and influence equations↩︎

This section turns the abstract causal mechanism into two EEHJB results. First, it derives a temperature-explicit stability estimate for learned-model perturbations. Second, at fixed positive temperature, it constructs a twice differentiable local equilibrium branch and its influence equations. The weak space below is sufficient for the mild fixed-point and stability arguments; the stronger Hölder space is used only to recover classical solutions. A separate comparison envelope records the additional uniformity needed when the temperature also varies.

4.1 Temperature-explicit stability↩︎

We work directly with the following diffusion system. Let \(r=r(y,t,x,a)\), let \(\delta\) be a fixed bounded \(C^1\) discount kernel with \(\delta(0)=1\), and let \(G(\vartheta,y,z)\) be fixed and smooth. Its first two arguments are reference variables, not flow-state variables. For a policy \(\pi\), set \[\begin{align} B^\pi(t,x)&=\int_\mathsf Ab(t,x,a)\pi(t,x,da),& R_y^\pi(t,x)&=\int_\mathsf Ar(y,t,x,a)\pi(t,x,da). \label{eq:general-policy-averages} \end{align}\tag{19}\] Its policy-evaluation pair is the solution of \[\label{eq:general-V1} \begin{align} 0={}&\partial_tV^{1,\pi}+\tfrac12\operatorname{tr} (\sigma\sigma^\top D_x^2V^{1,\pi})+B^\pi\cdot D_xV^{1,\pi}\\ &+\delta(t-\vartheta)\{R_y^\pi-\tau\operatorname{KL}(\pi\Vert\mu)\}, \qquad V^{1,\pi}(\vartheta,T,y,x)=F(\vartheta,y,x). \end{align}\tag{20}\] \[\label{eq:general-V2} \begin{align} 0={}&\partial_tV^{2,\pi}+\tfrac12\operatorname{tr} (\sigma\sigma^\top D_x^2V^{2,\pi})+B^\pi\cdot D_xV^{2,\pi},\\ &V^{2,\pi}(T,x)=h(x). \end{align}\tag{21}\] The variables \((\vartheta,y)\) are frozen in 20 and in \(G\). The acting self’s criterion is \[J(t,x;\pi):=V^{1,\pi}(t,t,x,x)+G(t,x,V^{2,\pi}(t,x)).\] Its policy-relevant flow-state gradient is \[\label{eq:Z-general} Z_V(t,x)=D_xV^1(t,t,x,x) +G_z(t,x,V^2(t,x))D_xV^2(t,x).\tag{22}\] For a model \(M=(b,r,F,h)\), put \[\begin{align} q_{M,V}(t,x,a)&=b(t,x,a)\cdot Z_V(t,x)+r(x,t,x,a), \tag{23}\\ \pi_{M,V}(t,x)&=\mathcal{G}_\tau(q_{M,V}(t,x,\cdot)). \tag{24} \end{align}\] An equilibrium is a classical pair \(V=(V^1,V^2)\) satisfying 2021 under \(\pi=\pi_{M,V}\). This is the full system needed below. The spike calculation in [prop:spike-verification] applies with the extra flow-state gradient \(G_zD_xV^2\) already included in \(Z_V\). The first two arguments of \(G\) remain frozen, so no reference-state derivative enters. Thus the PDE fixed point is equivalent to the local equilibrium condition.

Stability and the implicit-function argument use the weak value space \[\mathcal{D}_1=\{(\vartheta,t,y,x):0\leq\vartheta\leq t\leq T, \;y,x\in\mathbb{R}^d\}, \qquad \mathcal{Q}=[0,T]\times\mathbb{R}^d.\] Let \(\mathfrak V\) be the product space of pairs \(V=(V^1,V^2)\) for which \(V^1,D_xV^1\) are bounded and uniformly continuous on \(\mathcal{D}_1\), and \(V^2,D_xV^2\) are bounded and uniformly continuous on \(\mathcal{Q}\). Its norm is \[\begin{align} \|V\|_{\mathfrak V}:={}& \sup_{\mathcal{D}_1}(|V^1|+|D_xV^1|) +\sup_{\mathcal{Q}}(|V^2|+|D_xV^2|). \label{eq:X1-norm} \end{align}\tag{25}\] For \(t_0\in[0,T]\), its tail norm is \[\begin{align} \|V\|_{\mathfrak V,[t_0,T]}:={}& \sup_{\substack{(\vartheta,t,y,x)\in\mathcal{D}_1\\t\ge t_0}} (|V^1|+|D_xV^1|) +\sup_{\substack{(t,x)\in\mathcal{Q}\\t\ge t_0}} (|V^2|+|D_xV^2|). \label{eq:X1-tail-norm} \end{align}\tag{26}\] The derivatives are classical flow-state derivatives. Uniform convergence of a function and its derivative preserves this relation along line segments, so \(\mathfrak V\) is complete. The diagonal flow-gradient map \[\label{eq:diagonal-trace-weak} \mathsf T V^1(t,x):=D_xV^1(t,t,x,x)\tag{27}\] has operator norm at most one from the first component of \(\mathfrak V\) into \(C_b(\mathcal{Q};\mathbb{R}^d)\). It is the flow-state derivative restricted to the diagonal, not the total derivative of \(x\mapsto V^1(t,t,x,x)\).

The coefficient norm is defined next. If \(E\) contains a distinguished flow-state variable \(x\), let \(\mathcal{B}_x^1(E)\) consist of functions for which \(f\) and \(D_xf\) are bounded and uniformly continuous in all variables, including the action variable when present, with norm \(\|f\|_{\mathcal{B}_x^1}=\sup_E(|f|+|D_xf|)\). Vector-valued versions use the Euclidean norm. Set \[\begin{align} \mathfrak C^1:={}& \mathcal{B}_x^1(\mathcal{Q}\times\mathsf A;\mathbb{R}^d) \times\mathcal{B}_x^1(\mathbb{R}^d\times\mathcal{Q}\times\mathsf A) \times\mathcal{B}_x^1([0,T]\times\mathbb{R}^d\times\mathbb{R}^d) \times\mathcal{B}_x^1(\mathbb{R}^d), \label{eq:coefficient-space} \end{align}\tag{28}\] for \((b,r,F,h)\) in that order. Supremum norms over reference variables and actions are part of this definition.

Classical regularity uses a stronger space. Put \(\rho_r=|y-y'|+|t-t'|^{1/2}+|x-x'|\), \(\rho_{\rm ref}=|\vartheta-\vartheta'|^{1/2}+|y-y'|\), and \(F^\sharp(\vartheta,y)=F(\vartheta,y,\cdot)\). The space \(\mathfrak C^{\rm cl}_\gamma\subset\mathfrak C^1\) requires, uniformly in \(a\), \(b\in C_b^{\gamma/2,1+\gamma}\), \(r,D_xr\) to be \(\gamma\)-Hölder in \(\rho_r\), \(F^\sharp\) to be bounded in \(C_b^{2+\gamma}\) and \(\gamma\)-Hölder in \(\rho_{\rm ref}\), and \(h\in C_b^{2+\gamma}\). Its exact norm is in the supplement.

The hypotheses needed for the fixed-temperature calculus are separated from the additional envelope needed for uniform small-temperature comparison.

Assumption 2 (Data and diffusion evolution). Fix \(\gamma\in(0,1)\). The action space is compact metric and \(\mu\) has full support. Put \(a=\sigma\sigma^\top\) and \(\mathcal{L}_t^0=\tfrac12\operatorname{tr}(a(t,\cdot)D_x^2)\). The following hold on the parameter neighborhood under consideration.

  1. The uncontrolled matrix \(a\) is bounded and uniformly elliptic. It belongs to \(C_b^{\gamma/2,1+\gamma}(\mathcal{Q})\), and its evolution family \(P^0_{t,s}\) satisfies \[\begin{align} \|P^0_{t,s}f\|_\infty&\le C_0\|f\|_\infty,\notag\\ \|D_xP^0_{t,s}f\|_\infty &\le C_0(s-t)^{-1/2}\|f\|_\infty,\notag\\ \|P^0_{t,s}g\|_{\mathcal{B}_x^1}&\le C_0\|g\|_{\mathcal{B}_x^1}. \label{eq:base-semigroup} \end{align}\qquad{(18)}\] For each \(\eta\in(0,\gamma)\) and \(0\le t\le t'<s\le T\), it also satisfies \[\begin{align} [D_xP^0_{t,s}f]_{C_x^\eta} &\le C_\eta(s-t)^{-(1+\eta)/2}\|f\|_\infty,\notag\\ \|D_xP^0_{t,s}f-D_xP^0_{t',s}f\|_\infty &\le C_\eta|t-t'|^{\eta/2}(s-t')^{-(1+\eta)/2}\|f\|_\infty. \label{eq:base-holder} \end{align}\qquad{(19)}\] For the classical conclusion, we also use the standard whole-space linear Schauder estimates for this evolution family under the same coefficient assumptions.

  2. The coefficient quadruple belongs to \(\mathfrak C^1\). For a finite-dimensional family, \(\theta\mapsto M_\theta\) is \(C^2\) from an open subset of \(\mathbb{R}^p\) into \(\mathfrak C^1\).

  3. The discount kernel is \(C^1\). For every \(R<\infty\) and \(0\le j\le3\), \(\partial_z^jG\) is bounded and uniformly continuous on \(\mathcal{Q}\times[-R,R]\). In particular, \[\omega_{G,R}(\varepsilon):= \sup_{\substack{(t,x)\in\mathcal{Q},\;|z|\vee|z'|\le R\\|z-z'|\le\varepsilon}} |G_{zzz}(t,x,z)-G_{zzz}(t,x,z')|\longrightarrow0.\]

  4. For the classical conclusion, \(M_\theta\) and its first two parameter derivatives are locally bounded in \(\mathfrak C^{\rm cl}_\gamma\), and, on bounded \(z\)-ranges, \(\partial_z^jG\), \(0\le j\le3\), are uniformly bounded in \(C_b^{\gamma/2,\gamma}(\mathcal{Q})\).

Items (i)(iii) are the mild clauses; item (iv) is the classical add-on. All bounds are local and uniform in the stated parameter neighborhood.

Assumption 3 (Uniform comparison envelope). Let \(\mathcal{I}\subset(0,\bar\tau]\). The selected classical equilibria compared in 5 obey \[\sup_{\tau\in\mathcal{I}} \bigl(\|V_{M,\tau}^\star\|_{\mathfrak V} +\|V_{\widetilde{M},\tau}^\star\|_{\mathfrak V}\bigr)\le B_V,\] and the mild-clause bounds in 2 are uniform over the same collection. Existence is asserted only for these selected solutions. No evaluation on an open value-function ball is included in this assumption.

Taking \(\mathcal{I}=\{\tau\}\) gives a fixed-temperature comparison. If \(0\in\overline{\mathcal{I}}\), the common bound is an extra small-temperature hypothesis; fixed-temperature well-posedness does not supply it.

For two models with the same \((\sigma,G)\), use the primitive distance \[d_1(M,\widetilde{M})=\|M-\widetilde{M}\|_{\mathfrak C^1}. \label{eq:model-distance}\tag{29}\] For \(t_0\in[0,T]\), set \[\label{eq:policy-tail-distance} d_{{\rm pol},t_0}(\pi,\widetilde{\pi}) :=\sup_{\substack{t\in[t_0,T]\\x\in\mathbb{R}^d}} \left\lVert \pi(t,x)-\widetilde{\pi}(t,x)\right\rVert_{\rm TV}, \qquad d_{\rm pol}(\pi,\widetilde{\pi}) :=d_{{\rm pol},0}(\pi,\widetilde{\pi}).\tag{30}\]

For a candidate \(V\) and model \(M\), abbreviate \[\begin{align} H_\tau(q)&:=\tau\log\int_\mathsf Ae^{q(a)/\tau}\mu(da),\notag\\ B_{M,V}&:=B^{\pi_{M,V}},& R_{M,y,V}&:=R_y^{\pi_{M,V}},\notag\\ \Phi^1_{M,V} &:=\delta(t-\vartheta)\{H_\tau(q_{M,V}) +R_{M,y,V}-R_{M,x,V}-B_{M,V}\cdot Z_V\}. \label{eq:cancelled-source} \end{align}\tag{31}\] Here \(R_{M,x,V}\) means that the frozen reward state is set equal to the current state. The Gibbs log-partition identity rewrites the first evaluation equation with drift \(B_{M,V}\) and source \(\Phi^1_{M,V}\); its short derivation is recorded in the supplement.

The next estimate uses the uncontrolled diffusion evolution and requires no spatial derivative bound on the Gibbs-induced drift.

Lemma 2 (Bounded-drift parabolic estimate). Suppose 2(i) holds. Let \(u\) be the mild solution of \[\partial_tu+\mathcal{L}_t^0u+c\cdot D_xu+f=0,\qquad u(T)=g,\] where \(\|c\|_\infty\le B_c\), \(f\) is bounded, and \(g\in\mathcal{B}_x^1\). Then \[\label{eq:parabolic-gradient-estimate} \left\lVert u(t)\right\rVert_\infty+\left\lVert D_xu(t)\right\rVert_\infty \le C\left\lVert g\right\rVert_{\mathcal{B}_x^1} +C\int_t^T\{1+(s-t)^{-1/2}\}\left\lVert f(s)\right\rVert_\infty\,ds,\qquad{(20)}\] where \(C\) depends on \((C_0,B_c,T)\) and not on derivatives of \(c\). The same estimate holds for a difference equation with a bounded drift forcing \(\dot{c}\cdot D_xu\).

Proof. Duhamel’s formula and ?? , with \(e(t)=\|u(t)\|_\infty+\|D_xu(t)\|_\infty\) and \(k(r)=1+r^{-1/2}\), give \[e(t)\le C\|g\|_{\mathcal{B}_x^1} +C\int_t^Tk(s-t)\|f(s)\|_\infty\,ds +CB_c\int_t^Tk(s-t)e(s)\,ds .\] Since \(k(r)\le C_Tr^{-1/2}\), the \(m\)th resolvent iterate is bounded by \((CB_c)^mT^{m/2}/\Gamma(m/2+1)\). Summation proves the estimate and uniqueness; short-interval contraction and continuation give existence. For induced drifts, \(B_c=\sup_a\|b(\cdot,a)\|_\infty\), with no temperature factor. The supplement records the full resolvent closure. ◻

Theorem 5 (Temperature-explicit learned-model stability). Let \(M\) and \(\widetilde{M}\) satisfy the mild clauses of 2 3, with equilibrium pairs \(V_M^\star,V_{\widetilde{M}}^\star\) and policies \(\pi_M^\star, \pi_{\widetilde{M}}^\star\). For every \(\tau\in\mathcal{I}\), \[\begin{align} \left\lVert V_M^\star-V_{\widetilde{M}}^\star\right\rVert_{\mathfrak V} &\le S_\tau d_1(M,\widetilde{M}), \label{eq:value-stability}\\ d_{\rm pol}(\pi_M^\star,\pi_{\widetilde{M}}^\star) &\le \min\left\{2,\frac{C(1+S_\tau)}{\tau} d_1(M,\widetilde{M})\right\}, \label{eq:policy-stability} \end{align}\] {#eq: sublabel=eq:eq:value-stability,eq:eq:policy-stability} where one may take \[\label{eq:S-tau} S_\tau=C E_{1/2}\!\left(C(1+\tau^{-1})T^{1/2}\right) \le C\exp\{CT(1+\tau^{-2})\}.\qquad{(21)}\] Along a sequence \(\tau\downarrow0\), this estimate requires \(\tau\in\mathcal{I}\) and the uniform envelope in 3. For a fixed model, any two classical solutions in the envelope coincide; hence the selected solution is unique in that class.

Proof. Let \(E(t)=\|V_M^\star-V_{\widetilde{M}}^\star\|_{\mathfrak V,[t,T]}\) and \(\varepsilon=d_1(M,\widetilde{M})\). Bounded derivatives of \(G\) give \(\|Z_{V_M^\star}-Z_{V_{\widetilde{M}}^\star}\|_\infty\le CE(t)\), so the score difference is bounded by \(C\{E(t)+\varepsilon\}\) in oscillation. Hence 1 yields \[\label{eq:policy-difference-intermediate} d_{{\rm pol},t}(\pi_M^\star,\pi_{\widetilde{M}}^\star)\le \min\left\{2,C\tau^{-1}\{E(t)+\varepsilon\}\right\}.\tag{32}\] Write \(\Delta\) for corresponding differences and take all suprema below over the tail \([t,T]\). Directly from the action averages, the log-partition identity, and the common envelope, \[\begin{align} \|\Delta B\|_\infty &\le C\{\varepsilon+d_{{\rm pol},t} (\pi_M^\star,\pi_{\widetilde{M}}^\star)\},\notag\\ \|\Delta H_\tau\|_\infty&\le C\{\varepsilon+E(t)\},\notag\\ \|\Delta\Phi^1\|_\infty &\le C\{\varepsilon+E(t)+d_{{\rm pol},t} (\pi_M^\star,\pi_{\widetilde{M}}^\star)\}. \label{eq:source-differences} \end{align}\tag{33}\] Keep \(B_{M,V_M^\star}\) as the principal drift. The remaining drift forcing is \(\Delta B\cdot D_xV_{\widetilde{M}}^\star\) and is controlled by the first line of 33 , without differentiating an induced drift. Subtracting the two cancelled systems, applying 2, and then using 32 gives \[\label{eq:fractional-value-inequality} E(t)\le C\varepsilon+C(1+\tau^{-1}) \int_t^T\{1+(s-t)^{-1/2}\}\{E(s)+\varepsilon\}\,ds.\tag{34}\] Since \(E\) is nonincreasing, the pointwise estimate passes to its tail norm; on a finite horizon the nonsingular kernel is dominated by the order-\(1/2\) kernel. The fractional Gronwall iteration from 1 proves ?? –?? . Substitution in 32 proves ?? . ◻

At fixed temperature the factor in ?? can be absorbed into a constant. If the learned model and temperature vary together within a common envelope, the theorem gives the sufficient condition \[\label{eq:eehjb-sufficient-consistency} \varepsilon_n\tau_n^{-1}\exp(C/\tau_n^2)\longrightarrow0.\tag{35}\] This rate is a worst-case certificate, not a sharp universal boundary. The exact model in 5 has a bounded Volterra kernel and a different, sharp \(e^{C/\tau}\) law.

4.2 The EEHJB influence system and a delta method↩︎

Let \(\theta\in\Theta\subset\mathbb{R}^p\) parameterize \(M_\theta=(b_\theta,r_\theta, F_\theta,h_\theta)\) as in 2; \((\sigma,G)\) remain known. Fix \(\tau>0\) and suppress \((\theta,\tau)\) from the equilibrium notation. Put \[\begin{align} B(t,x)&=\int_\mathsf Ab(t,x,a)\pi(t,x,da),\notag\\ R_y(t,x)&=\int_\mathsf Ar(y,t,x,a)\pi(t,x,da). \label{eq:B-R} \end{align}\tag{36}\] For a direction \(u\in\mathbb{R}^p\), dots denote derivatives at \(\theta\) in direction \(u\). Define \[\begin{align} \dot{Z}={}&D_x\dot{V}^1(t,t,x,x) +G_{zz}(t,x,V^2)[\dot{V}^2]D_xV^2 +G_z(t,x,V^2)D_x\dot{V}^2, \tag{37}\\ \dot{q}(a)={}&\dot{b}(a)\cdot Z+b(a)\cdot\dot{Z} +\dot{r}(x,t,x,a), \tag{38}\\ \dot{\pi}(da)={}&\frac{\pi(da)}{\tau} \left\{\dot{q}(a)-\int_\mathsf A\dot{q}(a')\pi(da')\right\}, \tag{39}\\ \dot{B}={}&\int_\mathsf A\dot{b}(a)\pi(da)+\int_\mathsf Ab(a)\dot{\pi}(da), \tag{40}\\ \dot{R}_y={}&\int_\mathsf A\dot{r}(y,t,x,a)\pi(da) +\int_\mathsf Ar(y,t,x,a)\dot{\pi}(da). \tag{41} \end{align}\]

For the classical conclusion, let \(\mathfrak V_\eta\subset\mathfrak V\) consist of pairs with uniform \(C_b^{1+\eta/2,2+\eta}\) flow regularity and bounded \(\eta\)-Hölder dependence of \(V^1,D_xV^1\) on the frozen variables, measured with \(|t-t'|^{1/2}+|x-x'|+|\vartheta-\vartheta'|^{1/2}+|y-y'|\). The precise Banach norm is recorded in the supplement. The only trace fact used below is \[\label{eq:diagonal-holder-trace} \|\mathsf T V^1\|_{C_b^{\eta/2,\eta}(\mathcal{Q})} \le C\|V^1\|_{\mathfrak V^1_\eta}, \qquad 0<\eta\le\gamma.\tag{42}\]

Policy tangents are signed kernels, whereas the KL term is evaluated only on probability kernels. Let \[\mathfrak P=C_b([0,T]\times\mathbb{R}^d;\mathcal{M}(\mathsf A))\] be the Banach space of finite signed kernels with norm \(\|\eta\|_{\mathfrak P}:=\sup_{t,x}\|\eta(t,x)\|_{\rm TV}\). Policy derivatives lie in its closed zero-mass subspace \(\mathfrak P_0\).

For a candidate \(V\in\mathfrak V\), use \(M=M_\theta\) in 31 , set \(\pi_{\theta,V}=\mathcal{G}_\tau(q_{\theta,V})\), and define \(U=\mathcal{E}_{\theta,\tau}(V)\) by \[\begin{align} U^1(\vartheta,t,y,\cdot) ={}&P^0_{t,T}F_\theta(\vartheta,y,\cdot) +\int_t^TP^0_{t,s} \{B_{\theta,V}\cdot D_xU^1+\Phi^1_{\theta,V}\}(s)\,ds, \tag{43}\\ U^2(t,\cdot) ={}&P^0_{t,T}h_\theta +\int_t^TP^0_{t,s} \{B_{\theta,V}(s)\cdot D_xU^2(s,\cdot)\}\,ds. \tag{44} \end{align}\] Flow variables are suppressed inside the integrands. Backward contraction and continuation give a unique mild solution. The Gibbs identity shows that the fixed points of 4344 are exactly those of the original Gibbs-substituted system; they are classical whenever they lie in \(\mathfrak V_\eta\). Define \[\label{eq:value-residual-map} \mathfrak F(\theta,V) :=V-\mathcal{E}_{\theta,\tau}(V).\tag{45}\]

Fix \(\tau>0\) and suppose the mild clauses of 2 hold. Locally on every bounded open parameter-value neighborhood on which the stated coefficient bounds hold, the mild evaluation map \(\mathcal{E}_\tau\) is uniquely defined into \(\mathfrak V\) and is twice continuously Fréchet differentiable. On each sufficiently small bounded convex neighborhood, \[\begin{align} \|\mathcal{E}_\tau(z+\zeta)-\mathcal{E}_\tau(z)\|_{\mathfrak V} &\le C_\tau\|\zeta\|,\notag\\ \|\mathcal{E}_\tau(z+\zeta)-\mathcal{E}_\tau(z) -D\mathcal{E}_\tau(z)[\zeta]\|_{\mathfrak V} &\le C_\tau\|\zeta\|^2, \label{eq:E-C2-main}\\ \|D\mathcal{E}_\tau(z+\zeta)-D\mathcal{E}_\tau(z)\|_{\rm op} &\le C_\tau\|\zeta\|.\notag \end{align}\tag{46}\] For a pure value increment \(W\), the terminal variation is zero and \[\label{eq:parabolic-calculus-causal} \|D_V\mathcal{E}_{\theta,\tau}(V)W\|_{\mathfrak V,[t,T]} \le C(1+\tau^{-1}) I_{T-}^{1/2}\!\left(\|W\|_{\mathfrak V,[\,\cdot,T]}\right)(t).\tag{47}\] The constants in the differentiability statements are fixed-temperature constants.

Proof. The bounded diagonal trace and fixed-temperature Gibbs derivative bounds make the induced drift and cancelled source \(C^2\) in the relevant supremum norms. Differentiating once and twice gives linear parabolic equations, to which 2 applies, proving 46 ; the remainder equations and continuity of the second derivative are in the supplement. For a pure value direction the terminal variation vanishes and the forcing at time \(s\) is bounded by \(C(1+\tau^{-1})\|W\|_{\mathfrak V,[s,T]}\), which gives 47 . ◻

Suppose the mild clauses and the classical add-on of 2 hold. If \(V\in\mathfrak V\) satisfies \(V=\mathcal{E}_{\theta,\tau}(V)\) for fixed \(\tau>0\), then \(V\in\mathfrak V_\eta\) for every \(\eta\in(0,\gamma)\). Hence the mild fixed point is a classical solution of the entropy-cancelled system and of the original Gibbs-substituted EEHJB system.

The parabolic bootstrap and the reference-variable difference estimate are given in the supplementary material.

Theorem 6 (Forced tangent EEHJB and statistical delta method). Under the mild clauses of 2, let \(\theta_0\) be interior to \(\Theta\) and suppose \(\theta\mapsto M_\theta\) is \(C^2\) into the stated coefficient spaces on an open neighborhood of \(\theta_0\). Fix \(\tau>0\) and suppose \((V_0,\pi_0)\) is one mild equilibrium at \(\theta_0\). Then there is a neighborhood \(U\) of \(\theta_0\) and a unique local mild equilibrium branch \(\Psi_\tau(\theta):=(V^\star_{\theta,\tau},\pi^\star_{\theta,\tau})\) from \(U\) into \(\mathfrak V\times\mathfrak P\), with \(\Psi_\tau(\theta_0)=(V_0,\pi_0)\). Uniqueness is among zeros whose value lies in a fixed \(\mathfrak V\)-neighborhood of \(V_0\). The branch is \(C^2\) at fixed temperature. Its derivative is the unique mild solution of 3741 and \[\begin{align} 0={}&\partial_t\dot{V}^1+\tfrac12\operatorname{tr} (\sigma\sigma^\top D_x^2\dot{V}^1)+B\cdot D_x\dot{V}^1 +\dot{B}\cdot D_xV^1\notag\\ &+\delta(t-\vartheta)\left[ \dot{R}_y -\tau\int_\mathsf A\log\frac{d\pi}{d\mu}(a)\,\dot{\pi}(da) \right], \qquad \dot{V}^1(\vartheta,T,y,x)=\dot{F}(\vartheta,y,x), \label{eq:tangent-V1}\\ 0={}&\partial_t\dot{V}^2+\tfrac12\operatorname{tr} (\sigma\sigma^\top D_x^2\dot{V}^2)+B\cdot D_x\dot{V}^2 +\dot{B}\cdot D_xV^2, \qquad \dot{V}^2(T,x)=\dot{h}(x). \label{eq:tangent-V2} \end{align}\] {#eq: sublabel=eq:eq:tangent-V1,eq:eq:tangent-V2} Here \(\delta\) is the fixed discount kernel; omit it when discounting has already been absorbed into \(r\). Under the classical add-on of 2, every value on this branch belongs to \(\mathfrak V_\eta\), \(0<\eta<\gamma\), and is classical.

More explicitly, there is a modulus \(\omega_\tau(r)\downarrow0\) as \(r\downarrow0\) such that \[\label{eq:frechet-remainder} \|\Psi_\tau(\theta_0+v)-\Psi_\tau(\theta_0) -D\Psi_\tau(\theta_0)[v]\|_{\mathfrak V\times\mathfrak P} \le \omega_\tau(\|v\|)\|v\|.\qquad{(22)}\] No uniformity of \(\omega_\tau\) as \(\tau\downarrow0\) is asserted.

The statistical input below is the finite-dimensional parameter \(\theta\in\mathbb{R}^p\); the output is function valued. We do not claim a nonparametric coefficient-process delta method.

If an estimator satisfies \[\label{eq:theta-clt} \sqrt n(\widehat\theta_n-\theta_0)\Rightarrow\Xi\quad\text{in }\mathbb{R}^p,\qquad{(23)}\] then \(\mathbb{P}(\widehat\theta_n\in U)\to1\), and on this event \[\label{eq:equilibrium-functional-clt} \sqrt n\left\{(V^\star_{\widehat\theta_n,\tau}, \pi^\star_{\widehat\theta_n,\tau}) -(V^\star_{\theta_0,\tau},\pi^\star_{\theta_0,\tau})\right\} \Rightarrow(\dot{V}[\Xi],\dot{\pi}[\Xi]).\qquad{(24)}\] The explicit local Gibbs factor in 39 contributes \(1/\tau^2\) to a variance only when a nonvanishing tied score contrast survives under the Gibbs law. The full tangent variance can additionally contain the square of the Volterra factor in 48 ; conversely, Gibbs concentration at a separated, strongly concave maximizer can cancel part of the local factor.

Proof. For fixed \(\tau\), [prop:parabolic-calculus] makes the entropy-cancelled residual \(C^2\). Its value derivative is \(D_V\mathfrak F=I-\mathcal{L}_\tau\), where 47 bounds every causal power of \(\mathcal{L}_\tau\) by the corresponding Gamma-denominator term. Induction gives \[\|\mathcal{L}_\tau^mW\|_{\mathfrak V,[t,T]} \le \frac{[C(1+\tau^{-1})]^m(T-t)^{m/2}}{\Gamma(m/2+1)}\|W\|_{\mathfrak V}.\] Thus the Neumann series converges in operator norm and is the inverse of \(I-\mathcal{L}_\tau\); in particular, \[\label{eq:tangent-inverse-bound} \|(D_V\mathfrak F)^{-1}\| \le C E_{1/2}\!\left(C(1+\tau^{-1})T^{1/2}\right).\tag{48}\] The \(C^2\) Banach implicit-function theorem applies and gives \[\|\dot{V}[u]\|_{\mathfrak V}\le S_\tau\|u\|, \qquad \|\dot{\pi}[u]\|_{\mathfrak P}\le C\tau^{-1}(1+S_\tau)\|u\|.\] The \(C^1\) implicit map gives ?? ; for \(v_n=\widehat\theta_n-\theta_0=O_{\mathbb{P}}(n^{-1/2})\), its scaled remainder is \(o_{\mathbb{P}}(1)\). Direct differentiation gives 37 –?? ; the constant term in the KL derivative vanishes because \(\int\dot{\pi}=0\). The continuous mapping theorem on the finite-dimensional range of \(D\Psi_\tau(\theta_0)\) then proves ?? , without a separability claim for the ambient supremum spaces [30]. Classicality follows from [prop:mild-to-classical]. ◻

Corollary 3 (A self-contained open class). Fix \(\tau>0\) and let a finite-dimensional family \(M_\theta\) satisfy 2. At \(\theta=0\), suppose \[b_0(t,x,a)=\bar b(t,x), \qquad r_0(y,t,x,a)=\bar r(y,t,x).\] The two linear terminal problems at \(\theta=0\) have a unique classical pair \(V_0\), and \(\pi_0=\mu\). There is \(\varepsilon>0\) such that every \(|\theta|<\varepsilon\) has a unique equilibrium in a fixed \(\mathfrak V\)-neighborhood of \(V_0\). It is classical, and the value and policy are \(C^2\) in \(\theta\).

Proof. At the base model the score is constant over actions for every candidate, so \(\pi_{0,V}=\mu\) and entropy cancellation makes \(\mathcal{E}_{0,\tau}(V)=V_0\) independent of \(V\). Hence \(D_V\mathfrak F(0,V_0)=I\). The mild-evaluation proposition and the Banach implicit-function theorem give the local branch and its uniqueness. Classicality follows from [prop:mild-to-classical]. ◻

The class is genuinely state dependent: bounded smooth perturbations of both reward and controlled drift give a spatially varying first policy derivative. A concrete two-parameter construction, including direct verification of 2, is given in the supplement.

The theorem keeps \(\tau>0\) fixed and does not justify a triangular array \(\tau_n\downarrow0\) without uniform control of the local neighborhood and quadratic remainder. Precise sufficient conditions are stated in the supplement. Likewise, a numerical equilibrium error \(o_{\mathbb{P}}(n^{-1/2})\) preserves the limit law, but turning this observation into a stopping rule requires a separate a posteriori error bound.

5 Exact models and feedback graphs↩︎

Section 3 leaves the upper bound unmatched unless a positive return mode is present. The two models below make that return mechanism explicit. The affine model gives a closed nonlinear statistical transition, but its linearly growing payoffs place it outside the bounded framework of 4. The trigonometric model has bounded smooth data and a uniformly nondegenerate diffusion.

5.1 An affine reference-state diffusion↩︎

Let \(\mathsf A=\{-1,+1\}\), let \(\mu\) be uniform, and write \[\label{eq:binary-mean} m^\pi(t,x)=\int_\mathsf Aa\,\pi(t,x,da).\tag{49}\] The two-point space is compact and \(\mu\) charges both points, so this is a direct instance of the action-space setup in 2; no limiting argument is used. For \(\sigma>0\), consider \[\label{eq:affine-state} dX_s=m^\pi(s,X_s)\,ds+\sigma\,dW_s.\tag{50}\] The self with frozen reference state \(y\) evaluates \[\begin{align} J^{\pi,\xi}(y;t,x)=\mathbb{E}_{t,x}\bigg[\int_t^T \bigg\{\beta(X_s-y)m^\pi(s,X_s) -\tau\operatorname{KL}\bigl(\pi(s,X_s)\Vert\mu\bigr)\bigg\}\,ds +\xi X_T\bigg], \label{eq:affine-objective} \end{align}\tag{51}\] where \(\beta>0\). Immediate rewards are tied on the acting diagonal \(y=x\). Nevertheless, the current drift changes the state from which future selves evaluate their reference-state reward.

Theorem 7 (Exact affine cascade). For every \((\tau,\xi)\in(0,\infty)\times\mathbb{R}\), the model 5051 has a unique equilibrium in the affine class \[\label{eq:affine-value-ansatz} V^\xi(y;t,x)=w_\xi(t)(x-y)+\xi y+c_\xi(t).\qquad{(25)}\] It is state independent and satisfies \[\begin{align} m_\tau^\xi(t)&=\tanh\!\left(\frac{w_\xi(t)}{\tau}\right), &w_\xi'(t)+\beta m_\tau^\xi(t)&=0, &w_\xi(T)&=\xi. \label{eq:affine-ode} \end{align}\qquad{(26)}\] More explicitly, \[\begin{align} \sinh\!\left(\frac{w_\xi(t)}{\tau}\right) &=e^{\beta(T-t)/\tau}\sinh(\xi/\tau), \label{eq:exact-w}\\ m_\tau^\xi(t) &=\frac{z_\tau^\xi(t)}{\sqrt{1+z_\tau^\xi(t)^2}}, &z_\tau^\xi(t)&=e^{\beta(T-t)/\tau}\sinh(\xi/\tau). \label{eq:exact-m} \end{align}\] {#eq: sublabel=eq:eq:exact-w,eq:eq:exact-m} At the tied model \(\xi=0\), the policy susceptibility is \[\label{eq:exact-affine-susceptibility} \chi_\tau(t):=\left.\partial_\xi m_\tau^\xi(t)\right|_{\xi=0} =\frac{1}{\tau} e^{\beta(T-t)/\tau}.\qquad{(27)}\]

Proof. Under a state-independent policy with mean \(m(t)\), substituting ?? into the linear evaluation PDE and comparing the coefficient of \(x-y\) gives \(w'(t)+\beta m(t)=0\) and \(w(T)=\xi\). The flow-state partial derivative is \(V_x=w(t)\). By contrast, the total derivative of the diagonal map is \(dV(x;t,x)/dx=V_x+V_y=\xi\); replacing the former by the latter would incorrectly erase the intertemporal effect.

At the acting diagonal, action \(a\) has score \(aw(t)\), so the binary Gibbs identity gives \(m(t)=\tanh(w(t)/\tau)\) and ?? . The remaining scalar coefficient solves a linear terminal equation and does not affect the policy. Global Lipschitz continuity gives uniqueness; the resulting bounded deterministic drift makes \(X\) Gaussian with finite moments, while \(\operatorname{KL}(\pi_s\Vert\mu)\le\log2\), proving admissibility. Finally, \[\frac{d}{dt}\log\left|\sinh\frac{w_\xi(t)}{\tau}\right| =-\frac{\beta}{\tau}\] away from zero; continuity covers the zero solution. Integration, followed by \(\tanh u=\sinh u/\sqrt{1+\sinh^2u}\), proves ?? –?? . Differentiation at \(\xi=0\) gives ?? . ◻

The formula yields an exact joint statistical and annealing limit, without requiring a uniform delta-method remainder.

Theorem 8 (Root-\(n\) phase transition). Let \(\widehat\xi_n\) satisfy \(Z_n:=\sqrt n\,\widehat\xi_n\Rightarrow Z\), where \(Z\) is nondegenerate and \(\mathbb{P}(Z=0)=0\), and let \(\tau_n\downarrow0\) with \(\sqrt n\,\tau_n\to\infty\). Define \[\label{eq:a-n} a_n=\frac{e^{\beta T/\tau_n}}{\sqrt n\,\tau_n}.\qquad{(28)}\] Then the time-zero equilibrium obeys \[\label{eq:three-phase-law} m_{\tau_n}^{\widehat\xi_n}(0)\Rightarrow \begin{cases} 0, & a_n\to0,\\[1mm] \displaystyle\frac{aZ}{\sqrt{1+a^2Z^2}}, &a_n\to a\in(0,\infty),\\[3mm] \operatorname{sign}(Z),&a_n\to\infty. \end{cases}\qquad{(29)}\] The exact unit-amplification temperature and its first-order expansion are \[\label{eq:lambert-critical-temperature} \tau_n^\star=\frac{\beta T}{W(\beta T\sqrt n)} \sim\frac{2\beta T}{\log n},\qquad{(30)}\] where \(W\) is the same principal real branch as above. Equivalently, a nonunit critical limit \(a\in(0,\infty)\) is characterized by \[\label{eq:exact-critical-centering} \frac{\beta T}{\tau_n}-\frac{1}{2}\log n-\log\tau_n\longrightarrow\log a.\qquad{(31)}\] At any such critical scale and every fixed \(t\in(0,T]\), \(m_{\tau_n}^{\widehat\xi_n}(t)\to0\) in probability, although the time-zero policy has the nondegenerate limit in the middle line of ?? .

Proof. Since \(Z_n=O_{\mathbb{P}}(1)\) and \(\sqrt n\tau_n\to\infty\), \(\widehat\xi_n/\tau_n\to0\) in probability. Consequently, \[\sinh\left(\frac{\widehat\xi_n}{\tau_n}\right) =\frac{Z_n}{\sqrt n\tau_n}\{1+o_{\mathbb{P}}(1)\}.\] Thus ?? gives \(z_{\tau_n}^{\widehat\xi_n}(0)=a_nZ_n\{1+o_{\mathbb{P}}(1)\}\). The three conclusions follow from the continuous mapping theorem; in the last case use \(\mathbb{P}(Z=0)=0\). Solving \(a_n=1\) gives \(x_ne^{x_n}=\beta T\sqrt n\) with \(x_n=\beta T/\tau_n\), proving ?? . Taking logarithms gives ?? . Finally, at a fixed \(t>0\) the corresponding amplification is \(a_ne^{-\beta t/\tau_n}\to0\). ◻

The leading and two-term logarithmic approximations used below are \[\label{eq:critical-temperature-approximations} \tau_{n,1}=\frac{2\beta T}{\log n},\qquad L_n=\log(\beta T\sqrt n),\qquad \tau_{n,2}=\frac{\beta T}{L_n-\log L_n}.\tag{52}\] The leading approximation misses the critical centering. Since \(W(x)=\log x-\log\log x+o(1)\), the log–log term determines whether the amplified noise has a finite nonzero limit. The transition is localized near the initial time: at the critical scale, the initial policy remains sample dependent while each fixed later-time policy converges to the reference mixture.

Although \(1/\log n\) also occurs in simulated annealing [31], there it is an algorithmic-time cooling law tied to energy barriers. Here it balances root-\(n\) statistical noise against a causal equilibrium condition number and changes the policy’s statistical limit; no mixing-time claim is involved.

5.2 Coupled bounded modes and feedback cycles↩︎

The scalar example extends to coupled modes without losing its closed linear response. Let \(d\ge1\), \(\mathsf A=[-1,1]^d\), and \(\mu=\mu_0^{\otimes d}\), where \(\mu_0\) is a symmetric nondegenerate probability measure with full support on \([-1,1]\). Put \[\label{eq:ell-v} v=\int_{-1}^1 a^2\mu_0(da)>0,\qquad \ell(z)=\frac{\int_{-1}^1 a e^{za}\mu_0(da)}{\int_{-1}^1 e^{za}\mu_0(da)}.\tag{53}\] For the uniform law, \(v=1/3\) and \(\ell(z)=\coth z-z^{-1}\), with the value at zero defined by continuity. Thus this model has a genuinely continuous action space.

For \(K\in\mathbb{R}^{d\times d}\), consider \[\begin{align} dX_s&=m^\pi(s,X_s)\,ds+\sigma\,dW_s, &m^\pi(t,x)&=\int_\mathsf Aa\,\pi(t,x,da), \tag{54}\\ r(y,s,x,a)&=\sin(x-y)^\top Ka, &F_{\boldsymbol{\varepsilon}}(y,x) &=\boldsymbol{\varepsilon}^\top\sin(x-y), \tag{55} \end{align}\] where sine acts componentwise, \(\sigma>0\), and \(W\) is \(d\)-dimensional. The data are bounded and smooth in the state variables, and the diffusion is uniformly elliptic. In the notation of 4, we take \(\delta\equiv1\), \(h\equiv0\), and \(G\equiv0\).

Theorem 9 (Coupled bounded modes). Write \(\nu=\sigma^2/2\). The model 5455 has a unique continuous equilibrium within the state-independent class. It satisfies \[\begin{align} \pi_t^{\boldsymbol{\varepsilon}}(da) &=\bigotimes_{i=1}^d \frac{e^{a_iZ_i(t)/\tau}\mu_0(da_i)}{\int_{-1}^1e^{uZ_i(t)/\tau}\mu_0(du)},\notag\\ m_i(t)&=\ell\!\left(\frac{Z_i(t)}{\tau}\right), \label{eq:coupled-gibbs}\\ Z_i(t)&=\varepsilon_i e^{-\nu(T-t)}\cos M_i(t,T) +\int_t^T e^{-\nu(s-t)}\cos M_i(t,s)(Km(s))_i\,ds,\notag\\ M_i(t,s)&=\int_t^s m_i(u)\,du. \label{eq:coupled-volterra} \end{align}\] {#eq: sublabel=eq:eq:coupled-gibbs,eq:eq:coupled-volterra} At \(\boldsymbol{\varepsilon}=0\), the susceptibility matrix is \[\label{eq:matrix-susceptibility} \chi_\tau(t):= D_{\boldsymbol{\varepsilon}}m_{\boldsymbol{\varepsilon}}(t)\big|_0 =\frac{v}{\tau} \exp\left\{\left(-\nu I+\frac{v}{\tau} K\right)(T-t)\right\}.\qquad{(32)}\] If \(K\) is symmetric, then \[\label{eq:symmetric-matrix-norm} \left\lVert \chi_\tau(t)\right\rVert_2 =\frac{v}{\tau} \exp\left\{\left(\frac{v\lambda_{\max}(K)}{\tau}-\nu\right)(T-t)\right\}.\qquad{(33)}\] If \(K\) and \(\boldsymbol{\varepsilon}\) range over compact sets and \(\mathcal{I}\subset(0,\bar\tau]\), these equilibria satisfy 2 3 with comparison constants independent of \(\tau\in\mathcal{I}\). Parameter derivatives are asserted only at fixed temperature.

Proof. For deterministic \(m\), Gaussian convolution gives, coordinatewise, \[\mathbb{E}_{t,x}\sin(X_s^i-y_i) =e^{-\nu(s-t)}\sin\{x_i-y_i+M_i(t,s)\}.\] Differentiating policy evaluation in the flow state and setting \(y=x\) gives ?? . The acting score is \(a^\top Z(t)\). Since the reference measure is a product, the Gibbs law factorizes and its mean is ?? .

A complex score turns the Volterra equation into a finite-dimensional ODE. Define \[\zeta_i(t)=\varepsilon_i e^{-\nu(T-t)+iM_i(t,T)} +\int_t^T e^{-\nu(s-t)+iM_i(t,s)}(Km(s))_i\,ds.\] Then \(Z_i=\operatorname{Re}\zeta_i\). In reverse time, writing \(\zeta_i(T-u)=x_i(u)+iy_i(u)\) turns the Volterra equation into \[\label{eq:coupled-score-ode} \begin{align} \dot{x}=-\nu x-m\odot y+Km, &\qquad \dot{y}=-\nu y+m\odot x,\\ m_i=\ell(x_i/\tau), &\qquad (x(0),y(0))=(\boldsymbol{\varepsilon},0). \end{align}\tag{56}\] The vector field is locally Lipschitz and has linear growth because \(|m_i|\le1\). Thus 56 , and equivalently the Volterra system, has one global solution. This proves existence and uniqueness in the state-independent class.

Symmetry of \(\mu_0\) gives \(\ell(0)=0\) and \(\ell'(0)=v\). Differentiation at the tied model yields \[\chi_\tau(t)=\frac{v}{\tau} e^{-\nu(T-t)}I +\frac{v}{\tau}\int_t^T e^{-\nu(s-t)}K\chi_\tau(s)\,ds.\] In reverse time this is the mild equation with generator \(-\nu I+(v/\tau)K\), proving ?? . Diagonalizing a symmetric \(K\) gives ?? .

For the final assertion, the uncontrolled evolution is the heat semigroup, the trigonometric primitives are bounded and smooth, and the induced drift is spatially constant with \(|m_i|\le1\). Gaussian convolution gives an explicit bounded value, while \(\tau\operatorname{KL}(\pi_t\Vert\mu)\le\operatorname{osc}_a(a^\top Z(t))\) and \(|Z(t)|\le C(|\boldsymbol{\varepsilon}|+T\|K\|)\). These bounds verify 2 3; the complete value formula and derivative check are recorded in the supplement. ◻

Corollary 4 (Acyclic and cyclic feedback). Assume \(K\ge0\) and fix \(t<T\). Give \(K\) the directed graph with an edge \(j\to i\) when \(K_{ij}>0\), and put \(h=T-t\).

  1. If the graph is acyclic with longest directed path of length \(L\), then \[\label{eq:dag-susceptibility} \chi_\tau(t)=\frac{v}{\tau} e^{-\nu h} \sum_{r=0}^L\frac{(vh)^r}{r!\tau^r}K^r, \qquad \left\lVert \chi_\tau(t)\right\rVert_2=\Theta(\tau^{-(L+1)}).\qquad{(34)}\] If \(K^Lc\ne0\), the response to \(n^{-1/2}c\) has boundary \[\label{eq:dag-statistical-scale} \tau_n\asymp n^{-1/\{2(L+1)\}}.\qquad{(35)}\]

  2. If the graph contains a directed cycle, then \(\rho(K)>0\) and \[\label{eq:cycle-susceptibility} \rho(\chi_\tau(t)) =\frac{v}{\tau}\exp\left\{-\nu h+\frac{v\rho(K)h}{\tau}\right\}.\qquad{(36)}\] Every induced matrix norm is bounded below by the right-hand side. For a Perron vector, the response to \(n^{-1/2}r\) has unit amplification at \[\label{eq:cycle-critical-temperature} \tau_{n,\rm cyc}^\star =\frac{v\rho(K)h}{W\!\left(\rho(K)h e^{\nu h}\sqrt n\right)} \sim\frac{2v\rho(K)h}{\log n}.\qquad{(37)}\]

For the full policy derivative, \[\label{eq:policy-matrix-lower} \left\|D_{\boldsymbol{\varepsilon}}\pi_t\big|_0 \right\|_{\ell^\infty\to{\rm TV}} \ge \|\chi_\tau(t)\|_{\infty\to\infty} \ge \rho(\chi_\tau(t)).\qquad{(38)}\] The matching upper bound differs by at most \(d/v\); hence the policy derivative has the same polynomial order in the acyclic case and exponential rate in the cyclic case. The norm comparison is proved in the supplement.

Proof. An acyclic nonnegative matrix is nilpotent after a simultaneous permutation of rows and columns. Path expansion gives \(K^{L+1}=0\) and \(K^L\ne0\), so ?? follows by expanding the exponential. A positive directed cycle implies \(\rho(K)>0\). Spectral mapping applied to ?? gives ?? . Scaling the leading DAG term by \(n^{-1/2}\) proves ?? . Along a Perron vector, equating the multiplier in ?? to \(\sqrt n\) gives \(xe^x=\rho(K)he^{\nu h}\sqrt n\) with \(x=v\rho(K)h/\tau\), which proves ?? . ◻

The graph also controls the nonlinear response.

Corollary 5 (Nonlinear acyclic decay and cooperative selection). Assume \(K\ge0\).

  1. If the graph is acyclic with longest path \(L\), then \[\label{eq:dag-nonlinear-bound} |m_{\boldsymbol{\varepsilon}}(t)| \le\frac{1}{\tau}\sum_{r=0}^L \frac{(T-t)^r}{r!\tau^r}K^r|\boldsymbol{\varepsilon}|.\qquad{(39)}\] In particular, \(\|\boldsymbol{\varepsilon}_\tau\|=O(e^{-\gamma/\tau})\) with \(\gamma>0\) implies \(m_{\boldsymbol{\varepsilon}_\tau}(t)\to0\).

  2. Suppose \(\kappa_*:=\min_i(K\mathbf{1})_i>0\), \(\sup\operatorname{supp}\mu_0=1\), and \(T<\pi/2\). Set \(a_\sigma=e^{-\nu T}\cos T\). For \(\boldsymbol{\varepsilon}_\tau=\tau e^{-\gamma/\tau}\mathbf{1}\) and fixed \(t<T\), if \[\label{eq:collective-selection-condition} 0<\gamma<a_\sigma v\kappa_*(T-t),\qquad{(40)}\] then \(m_{\boldsymbol{\varepsilon}_\tau,i}(t)\to1\) for every \(i\).

Proof. Since \(\ell'(z)\) is the variance under an exponential tilt of \(\mu_0\), \(|\ell'(z)|\le1\). Equations ?? –?? give \[|m(t)|\le\tau^{-1}|\boldsymbol{\varepsilon}| +\tau^{-1}\int_t^T K|m(s)|\,ds.\] Positive iteration stops after \(K^L\) and proves ?? .

For the second claim, the positive orthant is invariant in 56 . On a face \(x_i=0\), one has \(m_i=0\) and \(\dot{x}_i=(Km)_i\ge0\). Thus \(x,m\ge0\). On a face \(y_i=0\), one has \(\dot{y}_i=m_ix_i\ge0\), so \(y\ge0\) as well. Since \(0\le M_i(t,s)\le s-t<T<\pi/2\), the score satisfies, componentwise, \[\label{eq:coupled-score-lower} Z_i(t)\ge a_\sigma\varepsilon_i +a_\sigma\int_t^T(Km(s))_i\,ds.\tag{57}\] Let \(\mathcal{P}\) be the positive Volterra map obtained by applying \(\ell(\cdot/\tau)\) to the right-hand side of 57 . The exact solution satisfies \(m\ge\mathcal{P}m\). Since \(m\ge0\) and \(\mathcal{P}\) is positive and monotone, \(0\le\mathcal{P}0\le\mathcal{P}m\le m\); induction gives \(m\ge\mathcal{P}^k0\) for every \(k\). The row-sum bound \(K\mathbf{1}\ge\kappa_*\mathbf{1}\) shows inductively that these vector iterates dominate the scalar Picard iterates whose limit is \(\underline m\). Hence \(m_i(t)\ge\underline m(t)\), where \[\underline w(t)=a_\sigma\tau e^{-\gamma/\tau} +a_\sigma\kappa_*\int_t^T\ell(\underline w(s)/\tau)\,ds, \qquad \underline m(t)=\ell(\underline w(t)/\tau).\] For \(q_\tau(u)=\underline w(T-u)/\tau\), \[q_\tau'(u)=\frac{a_\sigma\kappa_*}{\tau}\ell(q_\tau(u)), \qquad q_\tau(0)=a_\sigma e^{-\gamma/\tau}.\] Since \(\ell(q)/q\to v\) at zero, the time needed to reach any fixed small \(\delta>0\) is \[\frac{\tau}{a_\sigma\kappa_*} \int_{a_\sigma e^{-\gamma/\tau}}^\delta\frac{dq}{\ell(q)} =\frac{\gamma}{a_\sigma v\kappa_*}+o(1).\] Condition ?? leaves positive time after this hitting point. On that interval \(\ell(q)\ge\ell(\delta)>0\), so \(q_\tau(T-t)\to\infty\). The support assumption gives \(\ell(q)\to1\). ◻

The restriction \(T<\pi/2\) is used only in part (ii): it keeps the cosine kernel positive. On a longer horizon the return can change sign, so the monotone comparison above gives no selection conclusion.

If \(K\ge0\) is irreducible, let \(r\gg0\) be its Perron vector and equip \(\mathbb{R}^d\) with the weighted norm \(\|x\|_r=\max_i|x_i|/r_i\). Then \(\|K\|_r=\rho(K)\), and the tangent operator in ?? satisfies ?? with \(\alpha=1\), \(\psi(t)=e^{\nu t}r\), and \(\kappa_-=\kappa=v\rho(K)\). Thus the abstract upper and lower exponential rates are both attained by a bounded uniformly elliptic model in every state dimension. A reducible matrix gives the same conclusion after restriction to an irreducible Perron class.

6 Time discretization↩︎

A finite time grid has a strictly triangular influence matrix and hence a polynomial condition number at fixed \(N\). The continuous-time operator can have infinitely many nonzero causal powers and exponential conditioning, so refinement need not commute with \(\tau\downarrow0\). We quantify this mismatch by applying a uniform right-endpoint rule to the coupled tangent equation after factoring out the Brownian damping.

For the coupled model with \(K\ge0\), set \(Y_i=e^{\nu(T-t_i)}\chi_{\tau,N}(t_i)\) and use the right-endpoint rule \[Y_i=\frac{v}{\tau} I+\frac{v\Delta}{\tau} \sum_{j=i}^{N-1}K Y_{j+1}, \qquad Y_N=\frac{v}{\tau} I.\] The resulting \(N\)-step time-zero susceptibility is \[\label{eq:matrix-discrete-susceptibility} \chi_{\tau,N}(0)=\frac{v}{\tau} e^{-\nu T} \left(I+\frac{vT}{N\tau}K\right)^N.\tag{58}\] For any fixed induced matrix norm, write \[\label{eq:matrix-relative-mesh-error} \mathfrak e_{\tau,N} :=\frac{\|\chi_{\tau,N}(0)-\chi_\tau(0)\|}{\|\chi_\tau(0)\|}.\tag{59}\] If the graph is acyclic with longest path \(L\), then \[\label{eq:dag-discrete-susceptibility} \chi_{\tau,N}(0)=\frac{v}{\tau} e^{-\nu T} \sum_{r=0}^L\binom Nr\left(\frac{vT}{N\tau}\right)^rK^r.\tag{60}\] For fixed \(N\ge L\), the ratio of the leading low-temperature coefficient to its continuous counterpart is \[\label{eq:dag-leading-grid-ratio} q_{N,L}=\frac{L!\binom NL}{N^L} =\prod_{j=0}^{L-1}\left(1-\frac{j}{N}\right).\tag{61}\] When \(L\le1\), the discrete and continuous susceptibilities agree for every \(N\ge1\). When \(L\ge2\) and \(N\ge L\), one has \(\mathfrak e_{\tau,N}\to0\) jointly with \(\tau\downarrow0\) if and only if \(N\to\infty\); no coupling between \(N\) and \(\tau\) is needed.

If the graph contains a cycle and \(r\) is a Perron vector, then along \(r\) the ratio of the discrete to continuous response is \[\label{eq:cycle-grid-ratio} \frac{\|\chi_{\tau,N}(0)r\|}{\|\chi_\tau(0)r\|} =\left(1+\frac{x}{N}\right)^Ne^{-x}, \qquad x=\frac{v\rho(K)T}{\tau}.\tag{62}\] Along any sequence \(\tau\downarrow0\) and \(N\to\infty\), this ratio tends to one if and only if \(N\tau^2\to\infty\).

Proof. Backward substitution of the linearized grid equation gives 58 . If \(K^{L+1}=0\), the binomial series stops at \(L\), giving 61 ; coefficientwise comparison of the two finite sums proves the stated DAG equivalence. For a Perron vector, 58 reduces to 62 , and \(u-u^2/2\le\log(1+u)\le u-u^2/[2(1+u)]\) gives \[-\frac{x^2}{2N}\le \log\left\{\left(1+\frac{x}{N}\right)^Ne^{-x}\right\} \le-\frac{x^2}{2N(1+x/N)}.\] The bounds are equivalent to \(x^2/N\to0\), hence to \(N\tau^2\to\infty\); the short necessity and subsequence arguments for both cases are recorded in the supplement. ◻

6.1 Numerical illustrations↩︎

All numerical calculations are reproducible. The Python file sicon_volterra_experiments.py generates every reported number and both figures. The Monte Carlo calculation uses \(100{,}000\) common Gaussian draws with seed \(20260718\); all remaining checks are deterministic. No constant in a theoretical curve is fitted to the output.

For 1, \(d=6\), \(T=1\), \(\sigma=0.8\), and \(\mu_0\) is uniform on \([-1,1]\), so \(v=1/3\). The chain matrix has \(K_{i+1,i}=1\) for \(i=1,\ldots,5\) and all other entries zero. The cyclic matrix adds \(K_{1,6}=1\).

Figure 1: Susceptibility and grid cost for a six-node chain and the cycle formed by adding one closing edge. Left: \log\|\chi_\tau(0)\|_2 is asymptotically linear in \log(1/\tau) for the chain and in 1/\tau for the cycle. Right: the smallest N for which the relative response error is at most 0.01. The chain uses the full matrix 2-norm error in 59 , while the cycle uses the Perron-mode response error in 62 . The chain cost approaches a temperature-independent level, whereas the cyclic Perron cost scales as \tau^{-2}.

For 2, \(\beta=T=1\), \(Z\sim N(0,1)\), and \(\widehat\xi_n=Z/\sqrt n\). The same Gaussian draws are used across the three temperature schedules at each \(n\).

At the largest simulated size \(n=2^{26}\), the mean absolute time-zero magnetizations are \(0.0503\), \(0.5229\), and \(0.9738\) in the stable, critical, and selection schedules. At \(t=T/2\) they are \(0.0048\), \(0.0233\), and \(0.1465\), consistent with the fixed-time convergence in 8. At \(n=10^{14}\), the relative errors of the leading and two-term logarithmic approximations in 52 are \(16.2\%\) and \(1.32\%\).

Figure 2: Critical-temperature transition. Left: the exact Lambert-W scale and the approximations in 52 . Right: Monte Carlo estimates of \mathbb{E}|m_{\tau_n}^{\widehat\xi_n}(0)| for \tau_n/\tau_n^\star\in\{1.5,1,0.7\}.

7 Discussion↩︎

On a DAG, a perturbation passes through only finitely many future selves. Under aligned forcing, a positive cycle permits repeated returns and can change polynomial susceptibility into exponential susceptibility. The bounded diffusion realizes both rates in arbitrary finite dimension. It also shows why a grid that is accurate at fixed temperature may fail under cooling: along a cyclic Perron mode, relative consistency requires \(N\tau^2\to\infty\).

Two limits of the present theory remain. The general parabolic argument gives \(e^{C/\tau^2}\), whereas the explicit graph model attains \(e^{C/\tau}\); whether state dependence can attain the larger rate is open. The statistical result assumes finite-dimensional model estimation and fixed positive temperature, except in the exactly solvable affine model. Action-dependent diffusion, nonsmooth estimators, and graph criteria for state-dependent equilibria remain unresolved.

Declaration of generative-AI assistance↩︎

OpenAI ChatGPT/Codex (July 2026) assisted with literature searches, drafting, editing, mathematical checks, and numerical code. The author assumes responsibility for all content.

Code availability↩︎

The accompanying reproducibility archive contains the seeded Python code and numerical output used to produce both figures and every reported numerical value.

Supplementary material. The supplement contains the detailed technical closures and classical bootstrap used in Sections 4–6.

References↩︎

[1]
W. Tang, Y. P. Zhang, and X. Y. Zhou, Exploratory HJB equations and their convergence, SIAM J. Control Optim., 60 (2022), pp. 3191–3216, https://doi.org/10.1137/21M1448185.
[2]
Y.-J. Huang, Z. Wang, and Z. Zhou, Convergence of policy iteration for entropy-regularized stochastic control problems, SIAM J. Control Optim., 63 (2025), pp. 752–777, https://doi.org/10.1137/24M1638744.
[3]
H. V. Tran, Z. Wang, and Y. P. Zhang, Policy iteration for exploratory Hamilton–Jacobi–Bellman equations, Appl. Math. Optim., 91 (2025), 50, https://doi.org/10.1007/s00245-025-10249-3.
[4]
J. Ma, G. Wang, and J. Zhang, Convergence analysis for entropy-regularized control problems: A probabilistic approach, SIAM J. Control Optim., 64 (2026), pp. 816–842, https://doi.org/10.1137/24M1680039.
[5]
D. Sethi, D. Šiška, and Y. Zhang, Entropy annealing for policy mirror descent in continuous time and space, SIAM J. Control Optim., 63 (2025), pp. 3006–3041, https://doi.org/10.1137/24M166591X.
[6]
J. Cao, F. Acero, D. Šiška, and Y. Zhang, Entropy regularization improves policy robustness in continuous-time reinforcement learning, 2026, https://arxiv.org/abs/2607.03168.
[7]
J. Yong, Time-inconsistent optimal control problems and the equilibrium HJB equation, Math. Control Relat. Fields, 2 (2012), pp. 271–329, https://doi.org/10.3934/mcrf.2012.2.271.
[8]
T. Björk, M. Khapko, and A. Murgoci, On time-inconsistent stochastic control in continuous time, Finance Stoch., 21 (2017), pp. 331–360, https://doi.org/10.1007/s00780-017-0327-5.
[9]
T. Björk, M. Khapko, and A. Murgoci, Time-Inconsistent Control Theory with Finance Applications, Springer Finance, Springer, Cham, 2021, https://doi.org/10.1007/978-3-030-81843-2.
[10]
Q. Lei and C. S. Pun, Nonlocal fully nonlinear parabolic differential equations arising in time-inconsistent problems, J. Differential Equations, 358 (2023), pp. 339–385, https://doi.org/10.1016/j.jde.2023.02.025.
[11]
Q. Lei and C. S. Pun, Nonlocality, nonlinearity, and time inconsistency in stochastic differential games, Math. Finance, 34 (2024), pp. 190–256, https://doi.org/10.1111/mafi.12420.
[12]
Y.-J. Huang, X. Yu, and K. Zhang, Policy iteration achieves regularized equilibrium under time inconsistency, 2026, https://arxiv.org/abs/2603.06145.
[13]
Z. Wang, X. Yu, J. Zhang, and Z. Zhou, Equilibrium under time-inconsistency: A new existence theory by vanishing entropy regularization, 2026, https://arxiv.org/abs/2603.10321.
[14]
C. Reisinger and Y. Zhang, Regularity and stability of feedback relaxed controls, SIAM J. Control Optim., 59 (2021), pp. 3118–3151, https://doi.org/10.1137/20M1312435.
[15]
E. Bayraktar, Z. Wang, and Z. Zhou, Stability of equilibria in time-inconsistent stopping problems, SIAM J. Control Optim., 61 (2023), pp. 674–696, https://doi.org/10.1137/22M1496955.
[16]
E. Bayraktar, Z. Wang, and Z. Zhou, Short communication: Stability of time-inconsistent stopping for one-dimensional diffusions, SIAM J. Financial Math., 13 (2022), pp. SC123–SC135, https://doi.org/10.1137/22M1510005.
[17]
Q. Lei and C. S. Pun, On the well-posedness of Hamilton–Jacobi–Bellman equations of the equilibrium type, 2023, https://arxiv.org/abs/2307.01986. Revised May 2026.
[18]
D. Possamaï and M. Rodriguez Polo, Here, there and everywhere: State-dependent time-inconsistent stochastic control, 2026, https://arxiv.org/abs/2603.22022.
[19]
E. Bayraktar, Y.-J. Huang, Z. Wang, and Z. Zhou, Relaxed equilibria for time-inconsistent Markov decision processes, Math. Oper. Res., 50 (2025), pp. 2666–2687, https://doi.org/10.1287/moor.2023.0209. Published online October 23, 2024.
[20]
M. Dai, Y. Dong, and Y. Jia, Learning equilibrium mean–variance strategy, Math. Finance, 33 (2023), pp. 1166–1212, https://doi.org/10.1111/mafi.12402.
[21]
N. S. Lesmana and C. S. Pun, A subgame perfect equilibrium reinforcement learning approach to time-inconsistent problems, SIAM J. Financial Math., 16 (2025), pp. 68–122, https://doi.org/10.1137/23M1594510.
[22]
X. Guo, Y. Huang, and X. Yu, Deterministic policy gradient for learning equilibrium in time-inconsistent control problems, 2026, https://arxiv.org/abs/2606.11798.
[23]
R. D. McKelvey and T. R. Palfrey, Quantal response equilibria for extensive form games, Experimental Economics, 1 (1998), pp. 9–41, https://doi.org/10.1023/A:1009905800005.
[24]
A. Leonidov, A. Savvateev, and A. G. Semenov, Quantal response equilibria in binary choice games on graphs, 2019, https://arxiv.org/abs/1912.09584.
[25]
K. Diethelm, The Analysis of Fractional Differential Equations, vol. 2004 of Lecture Notes in Mathematics, Springer, Berlin, 2010, https://doi.org/10.1007/978-3-642-14574-2.
[26]
I. Podlubny, Fractional Differential Equations, vol. 198 of Mathematics in Science and Engineering, Academic Press, San Diego, 1999.
[27]
G. Gripenberg, S.-O. Londen, and O. Staffans, Volterra Integral and Functional Equations, vol. 34 of Encyclopedia of Mathematics and its Applications, Cambridge University Press, Cambridge, 1990, https://doi.org/10.1017/CBO9780511662805.
[28]
A. Berman and R. J. Plemmons, Nonnegative Matrices in the Mathematical Sciences, vol. 9 of Classics in Applied Mathematics, Society for Industrial and Applied Mathematics, Philadelphia, 1994, https://doi.org/10.1137/1.9781611971262.
[29]
R. M. Karp, A characterization of the minimum cycle mean in a digraph, Discrete Mathematics, 23 (1978), pp. 309–311, https://doi.org/10.1016/0012-365X(78)90011-0.
[30]
A. W. van der Vaart, Asymptotic Statistics, vol. 3 of Cambridge Series in Statistical and Probabilistic Mathematics, Cambridge University Press, Cambridge, 1998, https://doi.org/10.1017/CBO9780511802256.
[31]
B. Hajek, Cooling schedules for optimal annealing, Math. Oper. Res., 13 (1988), pp. 311–329, https://doi.org/10.1287/moor.13.2.311.

  1. Independent Researcher, Redwood City, CA (, https://orcid.org/0009-0008-1026-1083).↩︎