An Adjoint-Sensitivity Framework for Lost-in-the-Middle Phenomena in Causal Residual Transformers


Abstract

We develop an adjoint-sensitivity framework for positional influence in causal residual Transformers and separate unconditional analytic results from conditional boundary-shape conclusions. The principal unconditional theorem is the residual-to-depth-flow estimate for layer controls converging in \(L^1\), complemented by a finite-token-to-Volterra attention estimate that explicitly controls the first cells near the causal endpoint. We define a normalized adjoint-energy influence density and derive its exact evolution along full-batch gradient flow. The adjoint admits an exact generator-term decomposition into residual transmission, nonlocal Volterra, and local channels, including all covariance cross terms. Causal masking can amplify early-position sensitivity and residual identity paths can transmit a right-localized terminal bias, but neither mechanism alone forces a U-shaped profile. We therefore state boundary advantages under independently checkable energy, correlation, and local-channel bounds; these conditions are sufficient rather than necessary. Finite-token influence balancing, positional reweighting, and task-aligned observability are presented as diagnostics or regularizers with explicit differentiation requirements, computational costs, and limitations. Controlled simulations illustrate that each intervention controls its designated surrogate, while observability balance or outer-loop reweighting need not monotonically reduce the influence-based Lost-in-the-Middle diagnostic.

Keywords: positional influence density; Lost-in-the-Middle; primacy and recency; position bias; deterministic influence control; deterministic gradient flow; continuous-depth causal residual Transformers; causal masking; Volterra adjoint; observability regularization.

MSC 2020: 68T07; 49K15; 65L20.

1 Introduction↩︎

Long-context language models can accept large prompts without using every position equally. Empirical work documents substantial position sensitivity in retrieval and question answering, often with stronger performance near the beginning or end of a context than in its middle [1], [2]. We use Lost-in-the-Middle for this double boundary advantage and ask a narrower and purely theoretical question:

Can primacy, recency, and lost-in-the-middle be derived as consequences of deterministic continuous-time training dynamics and continuous-depth causal residual dynamics, and can one write principled influence-control laws that remove them?

Our answer is affirmative in a precise but conditional sense. We do not claim that the model reproduces every feature of a production Transformer or that one mechanism explains every long-context failure. Rather, the paper first establishes the analytic framework in which positional sensitivity can be compared: the residual recursion converges to the continuous-depth flow when the interpolated layer controls converge in \(L^1\), as quantified by Theorem 7, finite-token causal attention converges to the Volterra operator by Theorem 6, and the backward equation 32 defines the adjoint-energy influence measure in Definition 1. Proposition 11 then justifies differentiating its density along the gradient flow 35 , leading to the exact evolution identity 37 .

Within this framework, causal masking and residual propagation provide two distinct available channels. The Volterra adjoint formula 56 transports covectors from later outputs to earlier input positions; Proposition 14 yields a quantitative cone-channel primacy effect under joint cross-data and cross-depth coercivity, together with a middle-channel upper bound that rules out the comparison-of-lower-bounds fallacy. The residual Duhamel identity 68 , together with the localized stability estimate 87 , shows that a right-localized terminal adjoint can persist to depth zero, but does not assert that the residual path creates such localization by itself. Proposition 15 gives a dimensionally consistent squared-energy condition for this persistence. Finally, the exact regional decomposition 95 retains all channel cross-covariances, and Theorem 16 gives explicit sufficient conditions under which the middle average is smaller than both boundary averages, equivalently when the index 42 is positive.

The unconditional results concern approximation, adjoint propagation, and influence-measure dynamics, whereas primacy, recency, and a U-shaped lost-in-the-middle profile are conditional mechanism conclusions. This distinction is essential: causal geometry enlarges the future cone of early positions and residual dynamics preserve terminal sensitivity, but neither feature alone forces a U-shaped influence profile for every parameter state, readout, and data law.

1.1 Relation to prior work and novelty↩︎

1.1.0.1 Gradient attribution and influence in sequence models.

Input gradients and integrated gradients quantify prediction sensitivity to input features [3]; attention rollout and attention flow instead propagate attention-based relevance across layers [4]. Our density is a normalized squared input-adjoint energy averaged over the data law. It is not an attention-weight explanation and does not satisfy the path axioms of integrated gradients. Its role is to provide a position-indexed sensitivity measure compatible with a forward–backward depth system and with exact regional energy identities.

1.1.0.2 Adjoint sensitivity and continuous-depth training.

The use of adjoints for neural ordinary differential equations is standard [5]. The present contribution is not the adjoint method alone, but its combination with a causal Volterra derivative, a normalized positional measure, and a decomposition that distinguishes residual transmission from nonlocal causal amplification. The continuous-depth viewpoint is related to control formulations of deep networks [6], [7]; unlike a mean-field state equation, however, the vector field studied here has no independent law-valued coefficient.

1.1.0.3 Jacobian, stability, and observability regularization.

Jacobian penalties have been used to improve robustness or stabilize implicit-depth models [8], [9]. Our observability construction is closer to a task-conditioned finite-horizon Gramian: it asks whether perturbations at different token positions remain visible through a fixed family of observation maps. Equalization is therefore relative to a specified task readout, not an intrinsic property of the network.

1.1.0.4 Long-context position bias and attention sinks.

Lost-in-the-Middle behavior, positional-attention calibration, and attention sinks have been documented empirically [1], [2], [10]. Chowdhury’s exact birth-time theory [11] derives a closed-form initialized U-shape under a Cesàro causal-decoder abstraction. Our framework is complementary: it tracks a training-time adjoint density for a controlled continuous-depth model and gives conditional, not universal, boundary conclusions.

1.1.0.5 Continuum and operator limits of attention.

Function-space formulations of attention establish continuum operators and discretization consistency in neural-operator settings [12]. Our Volterra limit specializes this perspective to normalized causal attention on a one-dimensional context interval and couples it to residual-depth and adjoint dynamics.

1.1.0.6 Novelty statement.

Relative to the preceding strands of work, the contribution is an integrated forward–backward framework for analyzing positional sensitivity, rather than an unconditional theory asserting that causal Transformers must exhibit a particular boundary shape. On the analytic side, we establish well-posedness and discretization estimates for a continuous-depth causal residual flow with a Volterra attention operator, and couple this framework to an adjoint-based positional influence measure whose normalized evolution along deterministic gradient flow is explicit. On the mechanism side, we separate residual transmission, nonlocal causal-cone amplification, and local channels; retain their cross-covariances in exact regional energy identities; and derive independently checkable sufficient conditions for primacy, persistence of a right-boundary bias, and a middle deficit. On the methodological side, we formulate finite-token diagnostics and candidate regularizers—including influence balancing, positional reweighting, singular-channel tests, and task-conditioned observability balancing—while stating their differentiation requirements, estimator limitations, computational costs, and present validation scope. Accordingly, the approximation and sensitivity identities are unconditional under their stated regularity hypotheses, whereas the conclusions about primacy, recency, and a U-shaped Lost-in-the-Middle profile remain conditional.

1.2 Mathematical viewpoint↩︎

The framework has three components.

  1. A causal residual Transformer is treated as an Euler discretization of a continuous-depth dynamical system. The causal mask is retained in the continuum limit as a Volterra support condition on the attention kernel.

  2. Full-batch gradient descent is studied through its continuous-time gradient-flow limit. Thus training time is a deterministic variable \(s\), and the parameter state follows a single trajectory rather than a stochastic law.

  3. The induced distribution of positional influence is defined directly by the adjoint sensitivity density ?? . Primacy, recency, and lost-in-the-middle are then diagnosed from this density along the gradient-flow path.

All claims are formulated at the level of deterministic continuous training, adjoint calculus, and structural properties of causal residual dynamics. The central message is simple: lost-in-the-middle is a positional sensitivity imbalance; the most direct theoretical cure is to control the training objective or architecture so that the normalized influence density matches a desired target, typically the uniform density. Equivalently, the positional bias studied here is not treated as a static summary of attention weights. The relevant object is the normalized adjoint-energy density, and deterministic gradient flow changes this density by increasing the relative influence of positions whose adjoint energy decays more slowly than the positional average. Thus training determines which positions gain or lose sensitivity, while the residual identity path supplies a separate structural channel through which localized terminal or right-readout sensitivity can persist backward through depth and appear as recency.

1.3 Contributions↩︎

The contributions are separated into proved approximation/sensitivity results, conditional mechanism statements, and proposed diagnostics.

  1. Analytic well-posedness and discretization. We formulate the causal residual Transformer as a controlled depth flow on \(L^2\), prove regularity of the masked attention–FNN vector field on bounded trajectories, quantify residual-to-ODE shadowing for convergent controls, and establish finite-token-to-Volterra and readout-adjoint consistency estimates.

  2. Exact adjoint and influence identities. We define the positional influence measure from input-adjoint energy, derive its exact normalized gradient-flow evolution, and give an exact regional residual/cone/local decomposition including all cross terms.

  3. Conditional sufficient conditions for boundary advantages. We prove a cone-channel primacy criterion, a pathwise/moment-based residual-persistence estimate, and an independently checkable double-deficit theorem based on channel lower bounds, middle and local-channel upper bounds, and Cauchy–Schwarz correlation constants.

  4. Finite-token diagnostics and proposed regularizers. We formulate influence balancing, positional reweighting, singular-channel diagnostics, and task-aligned observability balancing, and state their second-order differentiation requirements, estimator bias/variance, and computational complexity.

Theorem 1 (Core rigorous and conditional conclusions).

Assume Assumptions 1 and 2, together with the spatial regularity and differentiability hypotheses stated in the corresponding results below.

  1. If the interpolated layer controls converge to a limiting control in \(L^1(0,T)\), the finite-depth residual recursion, including the split-layer defect, converges to the corresponding depth flow with the error stated in Theorem 7.

  2. On bounded \(C^1\) position fields and away from the left endpoint, finite-token masked attention converges to the continuum Volterra operator at order \(O(L^{-1})\) as stated in Theorem 6; the endpoint is defined through the trace of the averaged operator.

  3. The input-adjoint energy defines the influence measure of Definition 1. Under Proposition 11, its \(L^1\) density is differentiable along gradient flow and satisfies 37 in \(L^1(0,1)\).

  4. The Volterra identity 56 supplies a possible left-boundary amplification channel. Under the joint random-kernel, cross-depth coercivity and middle upper-bound conditions of Proposition 14, the depth-integrated cone channel has a quantitative left advantage. The residual identity formula 68 preserves any sufficiently strong right-localized terminal sensitivity, and Proposition 15 quantifies in regional squared energy when that right bias survives to depth zero.

  5. Lost-in-the-middle at scale \(\delta\) is equivalent to positivity of the lost-in-the-middle index 42 . Theorem 16 gives sufficient channel and cross-covariance conditions for this inequality. Under the extra \(L^2\)-valued differentiability hypothesis of Proposition 12, the regularizers in 97 , 102 , and 113 are proposed mechanisms for reducing the measured imbalance; they are not asserted to solve it without the additional optimization conditions stated in the algorithmic discussion.

1.3.0.1 Organization of the paper.

Section 2 gives the model, parameter topology, well-posedness, joint discretization results, and adjoint influence definition. Section 3 derives influence-density dynamics, the exact channel decomposition, independently verifiable sufficient conditions and countervailing terms, and the proposed algorithms with complexity estimates. Section 4 reports controlled finite-token experiments. Section 5 separates proved statements, conditional conclusions, and open problems.

2 Continuous-Depth Residual Transformers↩︎

This section fixes the notation and turns a finite-depth causal residual Transformer into a continuous-depth controlled flow. The construction is intentionally explicit about the causal mask, since the triangular support of masked attention is one architectural source of boundary-sensitive positional dynamics.

2.1 Notation convention↩︎

At finite sequence length, we use the one-based convention \[\begin{align} X=(x_1,\ldots,x_L)^\top\in\mathbb{R}^{L\times d_x}, \end{align}\] where \(d_x\) is the hidden dimension and \(L\) is the context length. The vocabulary dimension is \(d_v\), the token label at position \(l\in\{1,\ldots,L\}\) is denoted by \(Y_l(X_0)\in\{e_1,\ldots,e_{d_v}\}\), and \(H_L(X)\in\mathbb{R}^{L\times d_v}\) denotes the discrete logit readout. We set \(p_l=l/L\) and use the right-endpoint cells \(I_l=((l-1)/L,l/L]\). The token-averaged cross-entropy loss is denoted by \(\mathcal{\ell}(Y,H_L(X))\); see 27 and 29 . The subscript in the input state \(X_0\) denotes depth time and is unrelated to token indexing.

To avoid overloading notation, we distinguish the following variables throughout the paper:

  • \(t\in[0,T]\) is depth time. A finite-depth Transformer recursion with step size \(\varepsilon=T/M\) is regarded as an explicit Euler discretization of a controlled flow.

  • \(\theta_t\in\Theta\) is the depth-dependent layer-parameter control, and \(\zeta(t)\) is a scalar depth profile as in [7]. For a fixed control path, \(\rho_t^\theta=\operatorname{Law}(X_t^\theta)\) denotes the law of hidden states along depth.

  • \(s\ge0\) is training time, the continuous limit of full-batch gradient descent iterations. The deterministic parameter state is denoted by \(\vartheta_s\) and evolves according to the gradient flow 35 .

  • \(p\in[0,1]\) is normalized context position. The main object of this paper is the positional influence density \(m_s^{\rm GD}(p)\) along gradient flow, which is distinct from the hidden-state law \(\rho_t^\theta\).

2.2 Context continuum and hidden states↩︎

Let \(\Omega=(0,1)\) be the normalized context interval. The base state space for the depth evolution is \[\mathcal{X}:=L^2(\Omega;\mathbb{R}^{d_x}),\] while sampling and pointwise quadrature statements are made on the regular class \[\mathcal{X}_{\rm reg}:=H^1(\Omega;\mathbb{R}^{d_x})\cap L^\infty(\Omega;\mathbb{R}^{d_x}).\] In one spatial dimension, \(H^1\) functions have continuous representatives, so \(X_t(p)\) is meaningful for \(X_t\in\mathcal{X}_{\rm reg}\). The continuous-depth well-posedness results are formulated in \(\mathcal{X}\); whenever a theorem uses pointwise samples or Lipschitz quadrature, the additional spatial regularity is stated explicitly. A “perturbation at position \(p\)” is interpreted through perturbations supported on shrinking cells around \(p\), not through a Dirac mass, which is not an element of \(L^2\).

Following the explicit-Euler notation of [7], a residual block with \(M\) layers, depth step size \(\varepsilon=T/M\), and depth profile \(\zeta\) has the form \[\label{eq:discrete-residual} X_{k+1}^{\varepsilon}(p)=X_k^{\varepsilon}(p)+\varepsilon\,\zeta(k\varepsilon)\mathscr F_{\theta_k}\bigl(X_k^{\varepsilon}\bigr)(p), \qquad k=0,\ldots,M-1.\tag{1}\] Here \(\theta_k\) denotes the trainable layer parameter and \(\mathscr F_{\theta_k}\) is the complete masked-attention plus tokenwise feed-forward vector field defined in 7 . No independent empirical-law or mean-field coefficient is introduced in the modeled dynamics.

For clarity, we now spell out the masked Transformer block that defines the induced vector field \(\mathscr F_\theta\). At finite sequence length, let \(\mathrm{LN}\) denote the regularized smooth layer-normalization map defined tokenwise as follows: for \(z\in\mathbb{R}^{d_x}\), set \[\label{eq:smooth-layernorm} \bar z:=\frac{1}{d_x}\sum_{m=1}^{d_x}z_m, \qquad \mathrm{LN}(z) :=\gamma_{\rm LN}\odot \frac{z-\bar z\mathbf{1}}{\left(d_x^{-1}\sum_{m=1}^{d_x}(z_m-\bar z)^2+\eta_{\rm LN}\right)^{1/2}} +\beta_{\rm LN}, \qquad \eta_{\rm LN}>0,\tag{2}\] where \(\gamma_{\rm LN},\beta_{\rm LN}\in\mathbb{R}^{d_x}\) are bounded affine parameters and \(\eta_{\rm LN}\) removes the variance singularity. For a sequence \(X\), \(\mathrm{LN}(X)_i\) means \(\mathrm{LN}(X_i)\). Let \(N_h\) be the number of attention heads, and for head \(h\in\{1,\ldots,N_h\}\), write \[\begin{align} q_i^h=W_{Q,h}\mathrm{LN}(X)_i,\qquad k_j^h=W_{K,h}\mathrm{LN}(X)_j,\qquad v_j^h=W_{V,h}\mathrm{LN}(X)_j. \end{align}\] The causal mask enters through the lower-triangular attention weights \[\label{eq:causal-attention-weights} A_{ij}^h(X) =\frac{\one_{\{1\le j\le i\}}\exp\!\left(\langle q_i^h,k_j^h\rangle/\sqrt{d_k}\right)}{\sum_{r=1}^{i}\exp\!\left(\langle q_i^h,k_r^h\rangle/\sqrt{d_k}\right)}, \qquad 1\le i,j\le L,\tag{3}\] where \(d_k\) is the head dimension. The masked multi-head attention and feed-forward maps are \[\begin{align} \mathrm{Attn}_{\theta}^{\rm c}(X)_i &=W_O\bigl[\sum_{j=1}^{i}A_{ij}^1(X)v_j^1,\ldots, \sum_{j=1}^{i}A_{ij}^{N_h}(X)v_j^{N_h}\bigr],\tag{4}\\ \mathrm{FNN}_{\theta}(Z)_i &=W_2\,\sigma\bigl(W_1\mathrm{LN}(Z)_i+b_1\bigr)+b_2, \qquad 1\le i\le L.\tag{5} \end{align}\] A pre-normalized residual Transformer layer may therefore be written as the split Euler step \[\label{eq:complete-transformer-layer} \widetilde{X}_{k+1}^{\varepsilon}=X_k^{\varepsilon} +\varepsilon\zeta(k\varepsilon)\mathrm{Attn}_{\theta_k}^{\rm c}(X_k^{\varepsilon}), \qquad X_{k+1}^{\varepsilon}=\widetilde{X}_{k+1}^{\varepsilon} +\varepsilon\zeta(k\varepsilon)\mathrm{FNN}_{\theta_k}(\widetilde{X}_{k+1}^{\varepsilon}).\tag{6}\] Equivalently, up to an \(O(\varepsilon^2)\) splitting error, 6 has the one-field form 1 with \[\label{eq:complete-vector-field-finite} \mathscr F_{\theta_k}(X_k^{\varepsilon}) =\mathrm{Attn}_{\theta_k}^{\rm c}(X_k^{\varepsilon})+ \mathrm{FNN}_{\theta_k}(X_k^{\varepsilon}).\tag{7}\] In the continuum-position notation, the corresponding masked attention component has Volterra form \[\begin{align} \mathrm{Attn}_{\theta}^{\rm c}(X)(p) &=W_O\left[\int_{0}^{p}A_\theta^1(p,q;X)W_{V,1}\mathrm{LN}(X(q))\,\mathrm{d}q,\ldots, \int_{0}^{p}A_\theta^{N_h}(p,q;X)W_{V,N_h}\mathrm{LN}(X(q))\,\mathrm{d}q\right],\tag{8}\\ A_\theta^h(p,q;X) &=\frac{\one_{\{0\le q\le p\}}\exp\!\left(\langle W_{Q,h}\mathrm{LN}(X(p)),W_{K,h}\mathrm{LN}(X(q))\rangle/\sqrt{d_k}\right)}{\int_{0}^{p}\exp\!\left(\langle W_{Q,h}\mathrm{LN}(X(p)),W_{K,h}\mathrm{LN}(X(r))\rangle/\sqrt{d_k}\right)\,\mathrm{d}r}, \qquad p>0. \tag{9} \end{align}\] For \(p>0\), it is sometimes convenient to encode the normalized kernel as the headwise prefix-attention probability measure \[\label{eq:attention-induced-measure} \mu_{\theta,X}^{h,p}(\,\mathrm{d}q) :=A_\theta^h(p,q;X)\,\mathrm{d}q, \qquad \mu_{\theta,X}^{h,p}([0,p])=1, \qquad \operatorname{supp}\mu_{\theta,X}^{h,p}\subset[0,p].\tag{10}\] Thus the notation \(\mu_X\) used below, when it is useful as shorthand, refers only to the collection \(\{\mu_{\theta,X}^{h,p}\}_{h,p}\) induced by 89 ; it is not an additional law-valued state variable. The kernel itself is not assigned a finite value at \(p=0\): even for a constant score it equals \(p^{-1}\one_{[0,p]}(q)\). What has a trace is the averaged attention operator. For \(X\in\mathcal{X}_{\rm reg}\) we define \[\label{eq:continuum-attention-left-trace} \mathrm{Attn}_{\theta}^{\rm c}(X)(0) :=W_O\bigl[W_{V,1}\mathrm{LN}(X(0)),\ldots, W_{V,N_h}\mathrm{LN}(X(0))\bigr],\tag{11}\] which is the limit of 8 as \(p\downarrow0\). Thus the first-token rule is represented by the trace of the operator, not by the pointwise limit of the normalized kernel. Globally in \(L^2\), the model operator has the same endpoint structure as the Hardy averaging operator \(p^{-1}\int_0^p f(q)\,\mathrm{d}q\), whose boundedness is used in the regularity argument.

The essential point is the support condition \(A_\theta^h(p,q;X)=0\) for \(q>p\). Causal masking therefore makes the depth dynamics triangular in the position variable: position \(p\) can aggregate only positions \(q\le p\). This triangular structure is one of the architectural sources that can create boundary wells in the effective positional potential \(V\) and links the present training-time influence theory to the birth-time position-bias mechanism of [11]. Throughout the regularity, depth-flow, adjoint, and primacy analyses below, \(\mathscr F_\theta\) denotes only the causal masked-attention plus tokenwise feed-forward block in 79 . Equivalently, the only retained \(X\)-dependent measure is the prefix-attention family 10 . Every quantity at output position \(p\) therefore depends only on \(X|_{[0,p]}\); no independent law of \(X\), global empirical statistic, or other noncausal lifting is included.

All matrix, bias, layer-normalization, and readout parameters are collected in a finite-dimensional product space \[\Theta\cong\mathbb{R}^{d_\theta},\] equipped with the Euclidean product norm (equivalently, with the Frobenius norm on matrix blocks and the Euclidean norm on vector blocks). Control paths are measured in the Bochner norms \(L^1(0,T;\Theta)\) and \(L^\infty(0,T;\Theta)\).

Assumption 1 (Parameter topology and boundedness). The admissible set \(\Theta_{\rm ad}\subset\Theta\) is closed and bounded, hence compact since \(\Theta\) is finite dimensional. All uniform parameter constants below depend only on explicit bounds for the blocks in \(\Theta_{\rm ad}\). The regularized layer-normalization map 2 and the activation \(\sigma\) have bounded first derivatives on bounded sets.

Assumption 2 (Depth profile and depth-control regularity). The scalar depth profile is nonnegative and satisfies \[\begin{align} 0\le \zeta(t)\le \|\zeta\|_{L^\infty(0,T)}<\infty \quad\text{for a.e. }t, \qquad \mathrm{Lip}(\zeta)<\infty. \end{align}\] The depth-control path \(t\mapsto\theta_t\) is Lipschitz, or piecewise Lipschitz with finitely many pieces and uniform Lipschitz constants, with values in a bounded admissible parameter set \(\Theta_{\rm ad}\).

Remark 2. The assumption for the limiting depth-control path \(t\mapsto\theta_t\) is a continuous-depth idealization of the finite-layer situation. An \(M\)-layer network defines a layerwise interpolation \(\theta^\varepsilon\), but bounded weights alone do not force these interpolants to converge to one common control as \(M\to\infty\). Theorem 7 therefore distinguishes two statements: an \(O(\varepsilon)\) Euler shadowing estimate for each interpolated control, and convergence to a single continuous-depth flow only when \(\theta^\varepsilon\to\theta\) in \(L^1(0,T)\). Smooth or bounded-variation depth parameterizations are standard sufficient ways to obtain this convergence.

Proposition 3 (Regularity of the masked Transformer vector field). Assume that Assumptions 1 and 2 hold. Let \[\label{eq:induced-transformer-field} \mathscr F_\theta(X) :=\mathrm{Attn}_{\theta}^{\rm c}(X)+\mathrm{FNN}_{\theta}(X),\qquad{(1)}\] where the attention term is given by 89 , or equivalently by integration against the prefix-attention measures 10 . Assume also that the matrix and affine parameterization of \(\mathscr F_\theta\) is locally Lipschitz in \(\theta\), uniformly on bounded state sets. Then \(\mathscr F_\theta(X)\in\mathcal{X}\), and the induced field has the following properties on bounded subsets of \(\mathcal{X}\). For every \(R>0\) there exists \(C_{F,R}>0\) such that, for all admissible \(\theta\in\Theta_{\rm ad}\) and all \(X,Y\in\mathcal{X}\) with \(\|X\|_{\mathcal{X}},\|Y\|_{\mathcal{X}}\le R\), \[\label{eq:transformer-field-state-lipschitz} \|\mathscr F_\theta(X)-\mathscr F_\theta(Y)\|_{\mathcal{X}} \le C_{F,R}\|X-Y\|_{\mathcal{X}}.\qquad{(2)}\] Moreover, \(\mathscr F_\theta(X)\) is uniformly bounded on bounded subsets of \(\mathcal{X}\), uniformly over admissible controls: for every \(R>0\) there exists \(M_R<\infty\) such that \[\label{eq:transformer-field-uniform-bound} \sup_{\theta\in\Theta_{\rm ad}}\sup_{\|X\|_{\mathcal{X}}\le R} \|\mathscr F_\theta(X)\|_{\mathcal{X}}\le M_R.\qquad{(3)}\] In addition, with \(G(t,X):=\zeta(t)\mathscr F_{\theta_t}(X)\) and \(t_k=k\varepsilon\), there exists \(C_G>0\) such that, for \(X\) in a bounded set and whenever \(t\) and \(t_k\) lie in the same Lipschitz piece of the control path, \[\label{eq:depth-time-consistency} \|G(t,X)-G(t_k,X)\|_{\mathcal{X}}\le C_G|t-t_k|.\qquad{(4)}\]

Remark 4 (Bounded state sets under pre-layer normalization). The bounded-state restriction in Proposition 3 is compatible with the pre-layer-normalized architecture in 6 . In each attention and feed-forward sublayer, the trainable maps act on \(\mathrm{LN}(X_i)\), not directly on the unnormalized token state. By 2 , for bounded affine layer-normalization parameters one has a uniform bound \[\label{eq:preln-uniform-normalized-bound} \|\mathrm{LN}(z)\|\le C_{\rm LN}, \qquad z\in\mathbb{R}^{d_x},\qquad{(5)}\] where the constant depends only on \(d_x\), \(\eta_{\rm LN}\), and the bounds on \(\gamma_{\rm LN}\) and \(\beta_{\rm LN}\). Hence the attention values, queries, keys, and feed-forward inputs are evaluated on a bounded normalized set, uniformly over the unnormalized residual stream. The residual update still adds increments to \(X\), but on a finite depth horizon these increments are bounded by ?? and Gronwall’s inequality. Thus, if the input set is bounded in \(\mathcal{X}\), the entire pre-layer-normalized residual trajectory remains in a bounded state set depending only on the input radius, the admissible parameter bounds, and \(T\).

Remark 5 (Scope of the attention-induced measure). The family \(\mu_{\theta,X}^{h,p}\) in 10 is only a notation for the normalized causal attention kernel. It is neither the probability law \(\mathcal{L}(X(U))\) of a randomly sampled token state nor an independent measure-valued argument of the vector field. Consequently, the derivative \(D\mathscr F_\theta(X)\) is obtained by differentiating the explicit normalized attention formula and the tokenwise feed-forward map; no separate Lipschitz assumption in a measure metric and no Lions derivative are required. A genuinely global statistic, such as an integral over the full context that enters every output position, is excluded from the present model since its derivative would add a non-Volterra channel and would invalidate the purely lower-triangular causal decomposition used in the primacy analysis.

Proof. We give the endpoint estimate explicitly. Fix one head and abbreviate \[u_X(p):=\mathrm{LN}(X(p)),\qquad V_X(p):=W_Vu_X(p),\qquad h(p):=\|X(p)-Y(p)\|.\] Regularization by \(\eta_{\rm LN}>0\) and bounded affine parameters make \(u_X\) globally bounded and globally Lipschitz. Hence the scores \(s_X(p,q)\) are uniformly bounded, say \(|s_X|\le B\), and \[\label{eq:score-l2-difference} |s_X(p,q)-s_Y(p,q)|\le C\bigl(h(p)+h(q)\bigr).\tag{12}\] Write \[a_X(p,q):=\frac{e^{s_X(p,q)}}{Z_X(p)},\qquad Z_X(p):=\int_0^p e^{s_X(p,r)}\,\mathrm{d}r.\] For \(p>0\), \[\label{eq:attention-density-hardy-bound} pe^{-B}\le Z_X(p)\le pe^B,\qquad 0\le a_X(p,q)\le \frac{e^{2B}}{p}.\tag{13}\] Using the quotient identity for \(a_X-a_Y\), the Lipschitz continuity of the exponential on \([-B,B]\), and 12 , we obtain \[\label{eq:normalized-kernel-l1-difference} \int_0^p|a_X(p,q)-a_Y(p,q)|\,\mathrm{d}q \le C\left(h(p)+\frac{1}{p}\int_0^ph(q)\,\mathrm{d}q\right).\tag{14}\] Let \((Hh)(p):=p^{-1}\int_0^ph(q)\,\mathrm{d}q\) be the Hardy averaging operator. For the head output \(\mathcal{T}(X)(p)=\int_0^pa_X(p,q)V_X(q)\,\mathrm{d}q\), add and subtract \(\int_0^pa_X(p,q)V_Y(q)\,\mathrm{d}q\). The global bound on \(V_Y\), the Lipschitz bound on \(V_X-V_Y\), and 1314 give \[\label{eq:attention-pointwise-hardy-stability} \|\mathcal{T}(X)(p)-\mathcal{T}(Y)(p)\| \le C\bigl(h(p)+(Hh)(p)\bigr).\tag{15}\] Hardy’s inequality on \((0,1)\), \[\|Hh\|_{L^2(0,1)}\le2\|h\|_{L^2(0,1)},\] therefore yields \[\label{eq:global-l2-attention-stability} \|\mathcal{T}(X)-\mathcal{T}(Y)\|_{L^2} \le C\|X-Y\|_{\mathcal{X}}.\tag{16}\] This controls the whole neighborhood of \(p=0\); the trace 11 is needed only to identify the continuous representative at the single endpoint, not to prove the \(L^2\) bound.

Concatenating finitely many heads and applying \(W_O\) preserves 16 . The feed-forward map is pointwise Lipschitz since its layer-normalized input is bounded. Thus the explicit dependence of the attention measures \(\mu_{\theta,X}^{h,p}\) on \(X\), already controlled by 14 , gives ?? directly; no additional law-valued coefficient is present. Normalization of the attention weights and boundedness of the normalized values and feed-forward inputs give ?? .

Finally, local Lipschitz dependence on \(\theta\), Lipschitz continuity of \(\zeta\), and Assumption 2 imply, for \(t,r\) in the same Lipschitz control piece, \[\|\zeta(t)\mathscr F_{\theta_t}(X)-\zeta(r)\mathscr F_{\theta_r}(X)\|_{\mathcal{X}} \le C_R|t-r|,\] which proves ?? on each such piece. ◻

The previous proposition justifies treating the discrete masked Transformer block as a regular vector field on bounded state sets. We now record the complementary consistency estimate showing that the finite-token causal attention formula is approximated by the continuum Volterra attention operator away from the singular left endpoint.

Theorem 6 (Finite-token attention to Volterra attention). Let \(p_i=i/L\), \(1\le i\le L\), and let \(X_L=(X(p_1),\ldots,X(p_L))\) be samples of a function \(X\in C^1([0,1];\mathbb{R}^{d_x})\). Assume that \(X\in C^1([0,1];\mathbb{R}^{d_x})\) with \(\|X\|_{C^1}\le R\) and Assumption 1 holds. Then, for every \(\eta\in(0,1)\), there exists \(C_{R,\eta}<\infty\)1 independent of \(L\), such that for every \(i\) with \(p_i\in[\eta,1]\), \[\label{eq:attention-continuum-pointwise-error} \left\| \mathrm{Attn}_{\theta}^{\rm c}(X_L)_i - \mathrm{Attn}_{\theta}^{\rm c}(X)(p_i) \right\| \le \frac{C_{R,\eta}}{L} .\qquad{(6)}\] Consequently, in the discrete normalized \(\ell^2\) norm away from the left boundary, \[\label{eq:attention-continuum-l2-error} \left( \frac{1}{L}\sum_{i:\,p_i\in[\eta,1]} \left\| \mathrm{Attn}_{\theta}^{\rm c}(X_L)_i - \mathrm{Attn}_{\theta}^{\rm c}(X)(p_i) \right\|^2 \right)^{1/2} \le \frac{C_{R,\eta}}{L} .\qquad{(7)}\] At the left boundary, the comparison is understood through the operator trace 11 , which agrees with the first-token finite attention rule. Moreover, there is a constant \(C_R\), independent of \(L\) and \(i\), such that the endpoint-aware estimate \[\label{eq:attention-global-cell-error} \left\| \mathrm{Attn}_{\theta}^{\rm c}(X_L)_i- \mathrm{Attn}_{\theta}^{\rm c}(X)(p_i) \right\| \le \frac{C_R}{i},\qquad 1\le i\le L,\qquad{(8)}\] holds. Consequently, if \(n_L=\lfloor\delta L\rfloor\ge1\), then the average error over the first \(n_L\) cells satisfies \[\label{eq:attention-left-regional-error} \frac{1}{n_L}\sum_{i=1}^{n_L} \left\| \mathrm{Attn}_{\theta}^{\rm c}(X_L)_i- \mathrm{Attn}_{\theta}^{\rm c}(X)(p_i) \right\| \le \frac{C_R(1+\log n_L)}{n_L}.\qquad{(9)}\] Thus regional primacy averages may include the first cells.

Proof. It is enough to prove the estimate head by head, since concatenation over finitely many heads and multiplication by the bounded matrix \(W_O\) only change the constant. Fix one head. Since \(X\in C^1([0,1];\mathbb{R}^{d_x})\) with \(\|X\|_{C^1}\le R\), Assumption 1 implies that the regularized layer-normalization map and the linear maps \(W_{Q,h},W_{K,h},W_{V,h}\) are uniformly bounded and uniformly Lipschitz on the range of \(X\), uniformly over \(\theta\in\Theta_{\rm ad}\). Hence the score \[\label{eq:continuum-attention-score} S_h(p,q;X) :=\frac{\left\langle W_{Q,h}\mathrm{LN}(X(p)),W_{K,h}\mathrm{LN}(X(q))\right\rangle}{\sqrt{d_k}}\tag{17}\] is uniformly bounded and Lipschitz in \((p,q)\) on \([0,1]^2\). In particular, there are constants \(B_R,L_R<\infty\), depending only on the bounded state set and the admissible parameter set, such that \[\label{eq:score-bounded-lipschitz-from-bounded-assumption} |S_h(p,q;X)|\le B_R, \qquad |S_h(p,q;X)-S_h(p',q';X)|\le L_R(|p-p'|+|q-q'|).\tag{18}\] The map \(q\mapsto W_{V,h}\mathrm{LN}(X(q))\) is also uniformly Lipschitz and bounded. Since the exponential is Lipschitz on \([-B_R,B_R]\), the two integrands \[q\mapsto \exp(S_h(p_i,q;X))W_{V,h}\mathrm{LN}(X(q)), \qquad q\mapsto \exp(S_h(p_i,q;X))\] are uniformly Lipschitz in \(q\), for \(p_i\in[\eta,1]\). Define \[\begin{align} N_i^L&:=\frac{1}{L}\sum_{j=1}^i \exp(S_h(p_i,p_j;X))\,W_{V,h}\mathrm{LN}(X(p_j)),\tag{19}\\ D_i^L&:=\frac{1}{L}\sum_{j=1}^i \exp(S_h(p_i,p_j;X)),\tag{20}\\ N(p_i)&:=\int_0^{p_i}\exp(S_h(p_i,q;X))W_{V,h}\mathrm{LN}(X(q))\,\mathrm{d}q,\tag{21}\\ D(p_i)&:=\int_0^{p_i}\exp(S_h(p_i,q;X))\,\mathrm{d}q.\tag{22} \end{align}\] The factor \(1/L\) cancels in the quotient \(N_i^L/D_i^L\), so this quotient is exactly the finite attention average in 4 . The right-endpoint Riemann-sum estimate for uniformly Lipschitz integrands gives \[\label{eq:riemann-attention-error} \|N_i^L-N(p_i)\|+|D_i^L-D(p_i)|\le \frac{C_R}{L},\tag{23}\] with \(C_R\) independent of \(L\), \(i\), and the admissible parameter. Moreover, the score bound gives direct lower bounds for both denominators, for every \(L\ge1\) and every \(1\le i\le L\): \[\label{eq:attention-two-denominator-lower-bounds} D(p_i)\ge p_i e^{-B_R}, \qquad D_i^L=\frac{1}{L}\sum_{j=1}^i e^{S_h(p_i,p_j;X)}\ge p_i e^{-B_R}.\tag{24}\] Applying the quotient identity for the difference of the two normalized numerators, the Riemann bound 23 , and \(\|N(p_i)\|\le C_Rp_i\) gives directly \[\left\|\frac{N_i^L}{D_i^L}-\frac{N(p_i)}{D(p_i)}\right\| \le \frac{C_R}{Lp_i}=\frac{C_R}{i}.\] This proves ?? . If \(p_i\ge\eta\), then \(i\ge\eta L\) and the same estimate gives ?? ; the normalized \(\ell^2\) bound follows by squaring and summing. Finally, \[\frac{1}{n_L}\sum_{i=1}^{n_L}\frac{C_R}{i} \le \frac{C_R(1+\log n_L)}{n_L},\] which proves ?? and explicitly treats the first cells in a left-boundary regional average. ◻

2.3 The continuous-depth limit↩︎

The continuous-depth model is the neural ODE generated by the explicit masked Transformer field, \[\label{eq:continuous-depth} \partial_t X_t(p)=\zeta(t)\mathscr F_{\theta_t}(X_t)(p), \qquad X_0=x,\tag{25}\] where \(t\mapsto\theta_t\) is a measurable control path and the only state-dependent measures inside \(\mathscr F_{\theta_t}\) are the prefix-attention measures 10 . In this formulation the network depth is a control horizon, and \(\rho_t^\theta=\operatorname{Law}(X_t^\theta)\) denotes only the induced law of hidden states across the data distribution; it is not an argument of the vector field.

Theorem 7 (Residual-to-ODE limit for convergent controls). Assume Assumptions 1 and 2, and the local Lipschitz dependence on \(\theta\) stated in Proposition 3. Let \(\theta\in W^{1,\infty}(0,T;\Theta_{\rm ad})\) be a fixed limiting control. For \(\varepsilon=T/M\), let \[\theta^\varepsilon(t)=\theta_k^\varepsilon, \qquad t\in[t_k,t_{k+1}), \qquad t_k=k\varepsilon,\] be a uniformly bounded layerwise control satisfying \[\label{eq:control-l1-convergence} \|\theta^\varepsilon-\theta\|_{L^1(0,T)}\longrightarrow0.\qquad{(10)}\] Suppose the finite-depth update has the form \[\label{eq:discrete-residual-with-defect} X_{k+1}^{\varepsilon} =X_k^{\varepsilon} +\varepsilon\zeta(t_k)\mathscr F_{\theta_k^\varepsilon} (X_k^{\varepsilon}) +d_k^\varepsilon, \qquad \|d_k^\varepsilon\|_{\mathcal{X}}\le C_{\rm split}\varepsilon^2.\qquad{(11)}\] The one-field recursion 1 has \(d_k^\varepsilon=0\), while the split attention–FNN step 6 satisfies ?? on bounded trajectories under the stated smoothness assumptions. Let \(X^\theta\) solve 25 with the fixed control \(\theta\). Then there is \(C_T\), independent of \(\varepsilon\), such that the piecewise constant interpolation satisfies \[\label{eq:resnet-ode-convergent-control-bound} \sup_{0\le t\le T} \|X_t^{\varepsilon}-X_t^\theta\|_{\mathcal{X}} \le C_T\left( \varepsilon+ \|\theta^\varepsilon-\theta\|_{L^1(0,T)} \right).\qquad{(12)}\] In particular, the discrete trajectories converge to the single depth flow \(X^\theta\). If no limiting control is assumed, the same Euler calculation gives only an \(O(\varepsilon)\) shadowing estimate relative to the \(\varepsilon\)-dependent ODE driven by \(\theta^\varepsilon\), not convergence to a common flow.

The proof is deferred to Appendix 6.

2.4 Adjoint sensitivities and positional influence↩︎

We now define positional influence. For an input-label pair \((X_0,Y)\) and a depth-control path \(\theta=(\theta_t)_{0\le t\le T}\), let \(X_t^\theta\) be the solution of 25 . The single-example terminal loss is \[\label{eq:single-example-loss} \mathscr L(\theta;X_0,Y):=\mathcal{\ell}\bigl(Y,H(X_T^\theta)\bigr),\tag{26}\] where \(H:\mathcal{X}\to L^2(\Omega;\mathbb{R}^{d_v})\) is the logit readout and \(\mathcal{\ell}(Y,H(X))\) is the token-averaged cross-entropy. In finite length, this means \[\begin{align} \label{eq:discrete-cross-entropy} \mathcal{\ell}(Y,H_L(X))=-\frac{1}{L}\sum_{l=1}^{L}Y_l^\top\log\operatorname{softmax}(H_L(X)_l). \end{align}\tag{27}\] The finite-token terminal adjoint is therefore \[\label{eq:terminal-adjoint-softmax} P_{T,L}^{\theta,X_0,Y} =D_X\{\mathcal{\ell}(Y,H_L(X))\}\big|_{X=X_{T,L}^\theta} =DH_L(X_{T,L}^\theta)^*\Bigl(\frac{1}{L}\bigl(\operatorname{softmax}(H_L(X_{T,L}^\theta))-Y\bigr)\Bigr).\tag{28}\] Here the subscript \(L\) records the finite-token discretization rather than the network depth. More precisely, \[H_L:\mathbb{R}^{L\times d_x}\longrightarrow\mathbb{R}^{L\times d_v}\] is the discrete readout acting on an \(L\)-token hidden-state matrix, and \(H_L(X)_l\in\mathbb{R}^{d_v}\) is the logit vector at token \(l\). The state \(X_{T,L}^{\theta}\in\mathbb{R}^{L\times d_x}\) is the terminal hidden matrix obtained after evolving the \(L\)-token residual recursion to depth time \(T\) under the control path \(\theta\); equivalently, if \(T=M\varepsilon\), then \(X_{T,L}^{\theta}=X_{M,L}^{\varepsilon,\theta}\). Thus \(P_{T,L}^{\theta,X_0,Y}\) is the Euclidean gradient of the token-averaged loss with respect to this terminal finite-token state. The continuum objects \(H\) and \(X_T^\theta\) below are the corresponding position-field readout and terminal depth-flow state, and Theorem 8 compares the cellwise embedding of the finite-token adjoint with the continuum adjoint.

In the continuum-position normalization, \(Y:[0,1]\to\mathbb{R}^{d_v}\) and \(H(X):[0,1]\to\mathbb{R}^{d_v}\), and token averaging is replaced by normalized Lebesgue integration: \[\begin{align} \mathcal{\ell}(Y,H(X)) &=-\int_0^1 Y(p)^\top\log\operatorname{softmax}(H(X)(p))\,\mathrm{d}p,\tag{29}\\ \left.\frac{\,\mathrm{d}}{\,\mathrm{d}\epsilon}\mathcal{\ell}(Y,H(X)+\epsilon U)\right|_{\epsilon=0} &=\int_0^1\bigl(\operatorname{softmax}(H(X)(p))-Y(p)\bigr)^\top U(p)\,\mathrm{d}p.\tag{30} \end{align}\] Hence, with respect to the \(L^2(\Omega;\mathbb{R}^{d_v})\) pairing, the continuum logit-gradient is the pointwise residual \(\operatorname{softmax}(H(X)(\cdot))-Y(\cdot)\), and the continuum terminal adjoint is \[\label{eq:continuum-terminal-adjoint} P_T^{\theta,X_0,Y} =DH(X_T^\theta)^*\Bigl(\operatorname{softmax}(H(X_T^\theta)(\cdot))-Y(\cdot)\Bigr).\tag{31}\]

Theorem 8 (Discrete-to-continuum loss and terminal adjoint). Let \(p_l=l/L\) for \(l=1,\ldots,L\), and suppose that \(Z(p):=H(X)(p)\) and \(Y(p)\) are Lipschitz on \([0,1]\), with \(Y_l=Y(p_l)\) and \(H_L(X_L)_l=Z(p_l)\). Assume also that the logit values remain in a bounded set, so that \[g(p):=-Y(p)^\top\log\operatorname{softmax}(Z(p))\] is Lipschitz with constant \(L_g\). Then \[\label{eq:cross-entropy-discrete-continuum-error} \left| -\frac{1}{L}\sum_{l=1}^{L}Y_l^\top\log\operatorname{softmax}(H_L(X_L)_l) + \int_0^1Y(p)^\top\log\operatorname{softmax}(H(X)(p))\,\mathrm{d}p \right| \le \frac{L_g}{L} .\qquad{(13)}\] Moreover, define the continuum logit residual \[\Delta(p):=\operatorname{softmax}(Z(p))-Y(p)\] and its sampled vector \(\Delta_L=(\Delta(p_1),\ldots,\Delta(p_L))\). Assume that the continuum readout adjoint is bounded on the relevant state set, \[\label{eq:readout-adjoint-boundedness} \|DH(X)^*U\|_{\mathcal{X}}\le M_R\|U\|_{L^2(0,1;\mathbb{R}^{d_v})},\qquad \|X\|\le R,\qquad{(14)}\] and that the discrete readout adjoint is first-order consistent after the finite-token gradient is interpreted as a cellwise density.2 Then there exists \(C_R<\infty\), independent of \(L\), such that \[\label{eq:terminal-adjoint-discrete-continuum-error} \left\|\mathcal{E}_L\!\bigl(LP_{T,L}^{\theta,X_0,Y}\bigr)-P_T^{\theta,X_0,Y}\right\|_{\mathcal{X}}\le \frac{C_R}{L},\qquad{(17)}\] where \(P_T^{\theta,X_0,Y}\) is the continuum terminal adjoint in 31 .

Remark 9. In practice this Lipschitz sampling assumption is the continuum version of the usual finite-token regularity imposed before passing from a sequence to a position field. The hidden path \(p\mapsto X(p)\) is taken from a bounded smooth or piecewise smooth interpolation of the discrete hidden states, and the readout is local or uniformly Lipschitz on bounded sets. Hence \[\|Z(p)-Z(q)\|=\|H(X)(p)-H(X)(q)\|\le L_H\|X(p)-X(q)\|\le L_H L_X |p-q|.\] For labels, the statement is exact when the continuum label field is obtained by interpolation or piecewise smoothing of the token labels and then sampled at \(p_l\). For discontinuous one-hot labels, the same estimate can be read after replacing the Lipschitz assumption by bounded variation or by applying the argument on intervals between finitely many label jumps; the resulting Riemann error remains first order away from those jumps, with constants depending on the total variation. Thus the hypothesis is not an additional modeling mechanism, but the regularity needed to compare the token average in 27 with the Lebesgue integral in 29 .

Remark 10 (Linear readout consistency). For the common local linear readout \[\label{eq:linear-readout-example} H(X)(p)=X(p)W_{\rm out},\qquad{(18)}\] where \(W_{\rm out}\in\mathbb{R}^{d_x\times d_v}\) is bounded, the consistency assumption in ?? is immediate. The continuum readout derivative is \[\label{eq:linear-readout-continuum-adjoint} DH(X)U(p)=U(p)W_{\rm out}, \qquad DH(X)^*\Delta(p)=\Delta(p)W_{\rm out}^{\!*},\qquad{(19)}\] where \(W_{\rm out}^{\!*}\) denotes the Euclidean adjoint of \(W_{\rm out}\). At finite length, if \(H_L(X_L)_l=X_lW_{\rm out}\), then \[\label{eq:linear-readout-discrete-adjoint} \bigl[DH_L(X_L)^*\Delta_L\bigr]_l=\Delta(p_l)W_{\rm out}^{\!*}.\qquad{(20)}\] Consequently the cellwise density embedding satisfies \[\label{eq:linear-readout-cellwise-consistency} \mathcal{E}_L\!\bigl(DH_L(X_L)^*\Delta_L\bigr)(p) =\Delta(p_l)W_{\rm out}^{\!*}, \qquad p\in I_l.\qquad{(21)}\] If \(\Delta\) is Lipschitz, then \[\begin{align} \left\| \mathcal{E}_L\!\bigl(DH_L(X_L)^*\Delta_L\bigr)-DH(X)^*\Delta \right\|_{\mathcal{X}}^2 &=\sum_{l=1}^{L}\int_{(l-1)/L}^{p_l} \|\bigl(\Delta(p_l)-\Delta(p)\bigr)W_{\rm out}^{\!*}\|^2\,\mathrm{d}p\notag\\ &\le \|W_{\rm out}\|_{\rm op}^2\,\mathrm{Lip}(\Delta)^2 \sum_{l=1}^{L}\int_{(l-1)/L}^{p_l}|p-p_l|^2\,\mathrm{d}p\notag\\ &\le \frac{\|W_{\rm out}\|_{\rm op}^2\mathrm{Lip}(\Delta)^2}{L^2}.\label{eq:linear-readout-first-order-bound} \end{align}\qquad{(22)}\] Taking the square root gives the first-order bound required in ?? . Thus, for a tokenwise linear readout, the only discrepancy between the finite-token adjoint density and the continuum adjoint is the standard cellwise sampling error of the logit residual.

The same conclusion holds for tokenwise \(C^1\) nonlinear readouts with uniformly Lipschitz derivatives on the bounded state set, and for uniformly banded local readouts whose discrete kernels are consistent quadratures of a local continuum kernel. A genuinely nonlocal readout is permitted in the loss and adjoint definitions, but it can spread terminal sensitivity across the whole context. It therefore preserves the causal interpretation of the forward* masked dynamics while invalidating any recency argument that relies on a localized terminal adjoint. Such a readout must be analyzed through its actual adjoint kernel rather than through the local-readout examples used below.*

Proof. The loss estimate is the right-endpoint Riemann-sum error for the Lipschitz function \(g\): \[\left|\frac{1}{L}\sum_{l=1}^{L}g(p_l)-\int_0^1 g(p)\,\mathrm{d}p\right| \le \sum_{l=1}^{L}\int_{(l-1)/L}^{p_l}|g(p_l)-g(p)|\,\mathrm{d}p \le \frac{L_g}{L}.\] This proves ?? . For the adjoint, the finite Euclidean gradient in 28 contains the factor \(1/L\). After multiplying by \(L\) and embedding it as a cellwise density, it represents the continuum logit residual \(\Delta\). The softmax map is smooth and Lipschitz on bounded logit sets, so \(\Delta(p_l)\) is a first-order Riemann sample of \(\Delta(p)\). Applying the assumed first-order consistency ?? and boundedness of \(DH(X)^*\) gives ?? . ◻

Let \(\mathscr F_\theta\) be the explicit masked-attention plus tokenwise feed-forward field in ?? . The adjoint variable \(P_t^{\theta,X_0,Y}\in\mathcal{X}\) is defined backward in depth time by \[\label{eq:adjoint} -\partial_t P_t^{\theta,X_0,Y} =\zeta(t)D\mathscr F_{\theta_t}\!(X_t^\theta)^*P_t^{\theta,X_0,Y}, \qquad \begin{gather} P_T^{\theta,X_0,Y}\text{ given by \eqref{eq:terminal-adjoint-softmax} in the finite-token convention,}\\ \text{or by \eqref{eq:continuum-terminal-adjoint} in the continuum-position convention.} \end{gather}\tag{32}\] Here \(D\mathscr F_{\theta_t}(X_t^\theta)^*\) is the Hilbert-space adjoint of the Fréchet derivative of the complete induced Transformer vector field. Its nonlocal part includes the derivative of the normalized prefix-attention measures 10 , and there is no additional derivative through an independent law-valued coefficient. If the input representation is perturbed as \(X_0+\epsilon\xi\), then the first variation is \[\label{eq:first-variation} \left.\frac{\,\mathrm{d}}{\,\mathrm{d}\epsilon}\mathscr L(\theta;X_0+\epsilon\xi,Y)\right|_{\epsilon=0} =\langle P_0^{\theta,X_0,Y},\xi\rangle_{\mathcal{X}}.\tag{33}\] Thus \(\|P_0^{\theta,X_0,Y}(p)\|\) is the first-order loss sensitivity of position \(p\) under the controlled continuous-depth Transformer.

Definition 1 (Positional influence measure and density). For a parameter path \(\theta_\cdot\) and a data law \(\mathcal{D}\), define a finite positive measure on Borel sets \(B\subset(0,1)\) by \[\label{eq:influence-measure} \mathcal{I}_{\theta_\cdot}(B) :=\mathbb{E}_{(X_0,Y)\sim\mathcal{D}} \|\Pi_B P_0^{\theta,X_0,Y}\|_{\mathcal{X}}^2,\qquad{(23)}\] where \(\Pi_B\) is multiplication by \(\one_B\). Since \(P_0\in L^2\), this measure is absolutely continuous with respect to Lebesgue measure. Its Radon–Nikodym density is denoted by \[\label{eq:unnormalized-influence-density} I_{\theta_\cdot}(p) :=\frac{\,\mathrm{d}\mathcal{I}_{\theta_\cdot}}{\,\mathrm{d}p}(p) =\mathbb{E}\|P_0^{\theta,X_0,Y}(p)\|^2 \quad\text{for a.e. }p,\qquad{(24)}\] where the final equality uses any jointly measurable representative. If \(\mathcal{I}_{\theta_\cdot}((0,1))>0\), define \[\label{eq:influence-density} \mathfrak m_{\theta_\cdot}(p) =\frac{I_{\theta_\cdot}(p)}{\int_0^1 I_{\theta_\cdot}(q)\,\mathrm{d}q}.\qquad{(25)}\] Influence at a single position is therefore understood as the density obtained from shrinking cells: for almost every \(p\), \[I_{\theta_\cdot}(p) =\lim_{h\downarrow0}\frac{1}{|B_h(p)|}\mathcal{I}_{\theta_\cdot}(B_h(p)),\] whenever the Lebesgue differentiation theorem applies.

This definition is deliberately analytic. It does not define influence by an experimental retrieval score, but by the adjoint sensitivity of the continuous-depth predictor. It also differs from the closed-form Cesàro influence density of [11]: that density is computed directly from an initialized causal decoder abstraction, whereas ?? is a normalized expected adjoint energy for a controlled continuous-depth predictor evaluated along the deterministic gradient-flow path.

3 Deterministic Gradient Flow and Influence-Density Dynamics↩︎

This section gives a deterministic continuous-time analysis of primacy, recency, and Lost-in-the-Middle. The only training dynamics in this section are full-batch gradient descent in the small-step limit.

3.1 Gradient-flow training and the normalized influence density↩︎

Let \(\vartheta\) denote the finite-dimensional training state encoding the depth-dependent controls \(\theta_t^\vartheta\), the readout \(H_\vartheta\), and all other trainable parameters. For a data law \(\mathcal{D}\) on \((X_0,Y)\), define \[\label{eq:gd-risk} R(\vartheta) := \mathbb{E}_{(X_0,Y)\sim\mathcal{D}} \left[ \mathcal{\ell}\bigl(Y,H_\vartheta(X_T^{\theta^\vartheta})\bigr) \right],\tag{34}\] where \(X_T^{\theta^\vartheta}\) is obtained from the continuous-depth equation 25 . The full-batch gradient-flow limit of gradient descent is \[\label{eq:gd-flow} \dot{\vartheta}_s=-\nabla R(\vartheta_s), \qquad \vartheta_{s=0}=\vartheta_0.\tag{35}\] For a deterministic initialization, the positional influence density along training time is \[\label{eq:gd-influence-density} m_s^{\rm GD}(p) = \frac{I_{\vartheta_s}(p)}{\displaystyle\int_0^1 I_{\vartheta_s}(q)\,\mathrm{d}q},\tag{36}\] where \(I_{\vartheta_s}\) is the unnormalized adjoint influence from Definition 1, computed using the depth-control path \(\theta^{\vartheta_s}_\cdot\) and the readout \(H_{\vartheta_s}\). Thus the object being tracked is the same positional measure as in ?? , but evaluated along the single deterministic gradient-flow trajectory.

In practice this follows on any smooth training region where the Transformer maps, the readout, and the layer-normalization surrogate are differentiable in \(\vartheta\), since \(P_0^{\theta^{\vartheta},X_0,Y}\) depends differentiably on \(\vartheta\) through the forward ODE and the backward adjoint equation. Hence \(I_\vartheta(p)=\mathbb{E}\|P_0^{\theta^{\vartheta},X_0,Y}(p)\|^2\) is differentiable for almost every \(p\) by differentiation under the expectation.

The preceding differentiability statement is the only place where a domination estimate is needed. We record the estimate explicitly so that the subsequent derivative of the normalized influence density is justified by standard sensitivity bounds for the forward–backward system.

Regularized layer normalization and softmax attention make the explicit finite-token block smooth in states and finite-dimensional parameters on bounded sets. Proposition 3, however, proves only the state-space regularity needed for well-posedness. The twice Fréchet differentiable lifted forward–backward system assumed in the next proposition is an additional sensitivity hypothesis, not a consequence already proved for every admissible control parameterization. Likewise, the \(L^4\) condition in Proposition 12 is separately imposed and is stronger than the \(L^2\) well-posedness theory.

Proposition 11 (\(L^1\) differentiability of the influence density). Assume Assumptions 1 and 2. Let \(U\) be a compact parameter neighborhood intersecting the gradient-flow trajectory. Assume that the lifted vector field \(G_\vartheta(t,X)=\zeta(t)\mathscr F_{\theta_t^\vartheta}(X)\) is twice continuously Fréchet differentiable in \((X,\vartheta)\) on the bounded trajectory set, with its first and second derivatives uniformly bounded for \(\vartheta\in U\). Assume a tokenwise linear readout \(H_\vartheta(X)=XW_{\rm out}\) with \(W_{\rm out}\) in a compact set and \[\label{eq:data-fourth-moment} \mathbb{E}\bigl[\|X_0\|_{\mathcal{X}}^4+\|Y\|_{L^2}^4\bigr]<\infty.\qquad{(26)}\] Then the maps \(\vartheta\mapsto P_0^\vartheta\) and \(\vartheta\mapsto D_\vartheta P_0^\vartheta\) satisfy \[\label{eq:adjoint-parameter-uniform-bound} \sup_{\vartheta\in U} \left( \|P_0^\vartheta\|_{\mathcal{X}}^2+ \|D_\vartheta P_0^\vartheta\|_{\mathcal{L}(\mathbb{R}^{d_\vartheta},\mathcal{X})}^2 \right) \le C_{U,T}\bigl(1+\|X_0\|_{\mathcal{X}}^2+\|Y\|_{L^2}^2\bigr).\qquad{(27)}\] Consequently the map \(\vartheta\mapsto I_\vartheta\) is Fréchet differentiable from \(U\) into \(L^1(0,1)\), with \[\label{eq:influence-parameter-gradient} D_\vartheta I_\vartheta[\eta](p) =2\mathbb{E}\left\langle D_\vartheta P_0^\vartheta[\eta](p),P_0^\vartheta(p) \right\rangle \quad\text{in }L^1(0,1).\qquad{(28)}\] In particular, the normalization \(Z(\vartheta)=\int_0^1I_\vartheta(p)\,\mathrm{d}p\) is differentiable whenever finite.

Proof. On the bounded trajectory set, the integral equation for the forward flow and Gronwall’s inequality give \[\sup_{0\le t\le T}\|X_t^\vartheta\|_{\mathcal{X}} \le C_{U,T}(1+\|X_0\|_{\mathcal{X}}).\] Differentiating the flow in \(\vartheta\) gives a linear inhomogeneous equation with uniformly bounded coefficients and forcing, hence \[\sup_{0\le t\le T}\|D_\vartheta X_t^\vartheta\| \le C_{U,T}.\] For the linear readout, the softmax residual is bounded by \(C(1+\|Y\|_{L^2})\), while differentiation with respect to \(W_{\rm out}\) introduces at most one factor \(X_T^\vartheta\). Therefore \[\|P_T^\vartheta\|_{\mathcal{X}} +\|D_\vartheta P_T^\vartheta\| \le C_{U,T}(1+\|X_0\|_{\mathcal{X}}+\|Y\|_{L^2}).\] The backward adjoint and its parameter derivative are linear final-value equations. The second-derivative assumption controls the forcing term \(D_\vartheta[D_XG_\vartheta(t,X_t^\vartheta)]^*P_t^\vartheta\). Variation of constants and backward Gronwall therefore yield the same linear growth bound at \(t=0\); squaring proves ?? .

For \(\eta\in\mathbb{R}^{d_\vartheta}\), the pointwise algebraic identity \[\|P_0^{\vartheta+h\eta}\|^2-\|P_0^\vartheta\|^2 =2\langle P_0^\vartheta,P_0^{\vartheta+h\eta}-P_0^\vartheta\rangle +\|P_0^{\vartheta+h\eta}-P_0^\vartheta\|^2\] holds for measurable representatives. Integrating in \(p\), taking expectations, and applying Cauchy–Schwarz in \(L^2(0,1)\) shows that the difference quotient converges in \(L^1\) to the right-hand side of ?? . Indeed, \[\|\langle D_\vartheta P_0^\vartheta[\eta],P_0^\vartheta\rangle\|_{L^1} \le \|D_\vartheta P_0^\vartheta[\eta]\|_{L^2} \|P_0^\vartheta\|_{L^2},\] and the expectation of this product is finite by ?? , ?? , and Cauchy–Schwarz in probability. This proves \(L^1\) differentiability and justifies differentiating all regional masses and the normalizing integral. ◻

Proposition 12 (\(L^2\) differentiability for the squared-density penalty).

In addition to Proposition 11, assume that3 \[\label{eq:adjoint-l4-sensitivity-assumption} \vartheta\longmapsto P_0^\vartheta \quad\text{is continuously Fr\'echet differentiable from U into}\quad L^4\bigl(\mathcal{D}\otimes\mathrm{Leb};\mathbb{R}^{d_x}\bigr),\qquad{(29)}\] with \[\sup_{\vartheta\in U} \left( \|P_0^\vartheta\|_{L^4(\mathcal{D}\otimes\mathrm{Leb})} +\|D_\vartheta P_0^\vartheta\|_{\mathcal{L}(\mathbb{R}^{d_\vartheta},L^4)} \right)<\infty.\] Assume also that \[\label{eq:influence-normalization-away-from-zero} Z(\vartheta)=\int_0^1I_\vartheta(p)\,\mathrm{d}p\ge z_0>0, \qquad \vartheta\in U.\qquad{(30)}\] Then \(\vartheta\mapsto I_\vartheta\) and \(\vartheta\mapsto m_\vartheta=I_\vartheta/Z(\vartheta)\) are continuously Fréchet differentiable as maps into \(L^2(0,1)\). In particular, for every target density \(\nu\in L^2(0,1)\), \[\label{eq:l2-influence-discrepancy} \Phi_\nu(\vartheta):=\frac{1}{2}\|m_\vartheta-\nu\|_{L^2(0,1)}^2\qquad{(31)}\] is differentiable and \[\label{eq:l2-influence-discrepancy-derivative} D\Phi_\nu(\vartheta)[\eta] =\left\langle m_\vartheta-\nu, Dm_\vartheta[\eta]\right\rangle_{L^2}.\qquad{(32)}\]

Remark 13 (Practical interpretation of the \(L^4\) sensitivity hypothesis). At a fixed finite token length, ?? is a natural local regularity condition for a smooth, regularized network: all norms on the finite-dimensional token state are equivalent, and bounded data, bounded parameter sets, regularized layer normalization, and bounded first and second derivatives give local fourth-moment bounds for the adjoint and its parameter sensitivity. The continuum statement is stronger and is not automatic from the \(L^2\) well-posedness theory. It requires uniform integrability of fourth powers jointly over data and position, so it is reasonable when inputs and readout derivatives have controlled fourth moments and the linearized forward and backward propagators remain uniformly bounded. It may fail for heavy-tailed data, unbounded readouts or activations, nearly singular normalization, or a sequence-length limit in which fourth moments are not uniform. In such regimes, the \(L^1\) differentiability result of Proposition 11 and regional-mass penalties remain valid under weaker assumptions, whereas the continuum squared-density penalty should be treated as a finite-token objective or used only after empirical fourth-moment and Jacobian-sensitivity diagnostics support the stronger hypothesis.

Proof. The map \[\mathcal{Q}:L^4(\mathcal{D}\otimes\mathrm{Leb})\to L^2(0,1), \qquad \mathcal{Q}(P)(p):=\mathbb{E}\|P(p)\|^2,\] is continuously differentiable. Indeed, Jensen’s inequality gives \[\|\mathcal{Q}(P)\|_{L^2}^2 =\int_0^1\bigl(\mathbb{E}\|P(p)\|^2\bigr)^2\,\mathrm{d}p \le \mathbb{E}\int_0^1\|P(p)\|^4\,\mathrm{d}p,\] and Hölder’s inequality gives the bounded derivative \[D\mathcal{Q}(P)[Q](p)=2\mathbb{E}\langle P(p),Q(p)\rangle, \qquad \|D\mathcal{Q}(P)[Q]\|_{L^2} \le2\|P\|_{L^4}\|Q\|_{L^4}.\] The quadratic remainder satisfies \(\|\mathcal{Q}(P+Q)-\mathcal{Q}(P)-D\mathcal{Q}(P)[Q]\|_{L^2} \le\|Q\|_{L^4}^2\). Composing \(\mathcal{Q}\) with ?? proves the \(L^2\) differentiability of \(I_\vartheta\). Integration is a bounded functional on \(L^2(0,1)\), so \(Z\) is differentiable; ?? and the Banach-space quotient rule give the \(L^2\) differentiability of \(m_\vartheta\). Finally, the squared \(L^2\) norm is continuously differentiable, which proves ?? . ◻

With Proposition 11 in hand, the unnormalized density is differentiable as an \(L^1\) function. Hence the quotient rule below holds in \(L^1(0,1)\), and therefore almost everywhere after choosing a representative. The normalization \(Z_s=\int_0^1 I_{\vartheta_s}(q)\,\mathrm{d}q\) is positive whenever the terminal loss has a nonzero adjoint on a set of positive probability; if the model has zero adjoint everywhere, the influence density is degenerate and the positional-bias diagnostic is unnecessary. Then differentiation of 36 along 35 gives \[\label{eq:gd-mdot} \partial_s m_s^{\rm GD}(p) = -\frac{\nabla_\vartheta I_{\vartheta_s}(p)\cdot\nabla R(\vartheta_s)}{Z_s} +\frac{m_s^{\rm GD}(p)}{Z_s} \int_0^1 \nabla_\vartheta I_{\vartheta_s}(q)\cdot\nabla R(\vartheta_s)\,\mathrm{d}q .\tag{37}\] The identity 37 is exact but not closed. Indeed, set \[r_s(p):= \begin{cases} \dfrac{\nabla_\vartheta I_{\vartheta_s}(p)\cdot\nabla R(\vartheta_s)}{I_{\vartheta_s}(p)},& I_{\vartheta_s}(p)>0,\\[0.8em] 0,& I_{\vartheta_s}(p)=0, \end{cases} \qquad \overline{r}_s:=\int_0^1 m_s^{\rm GD}(q)r_s(q)\,\mathrm{d}q .\] On the positive set, \(\partial_s I=-r_sI\). On \(\{I=0\}\) the original quotient-rule identity 37 is used; since \(I_\vartheta(p)\ge0\) and is differentiable in the parameter, its parameter gradient vanishes at an interior zero. Thus the convention above makes \(m_sr_s=0\) on the zero set and the replicator form \[\partial_s m_s^{\rm GD}(p)=m_s^{\rm GD}(p)\bigl(\overline{r}_s-r_s(p)\bigr)\] holds almost everywhere. Thus the normalized mass at position \(p\) increases exactly when its relative decay rate \(r_s(p)\) is smaller than the density-weighted average \(\overline{r}_s\); equivalently, other positions are being suppressed faster than position \(p\). This is the training-time form of positional bias: the bias is determined not only by the current attention weights, but also by the way the full gradient-flow vector field redistributes normalized adjoint energy across positions.

3.2 Quantifying primacy, recency, and the middle deficit↩︎

For \(\delta\in(0,1/2)\), use the half-open convention \[E_L=[0,\delta),\qquad E_M=[\delta,1-\delta),\qquad E_R=[1-\delta,1].\] The continuum boundary points have zero measure, while the convention gives an unambiguous finite-token assignment. Define \[\begin{align} \overline{m}_L^{\rm GD}(s,\delta) &:=\frac{1}{\delta}\int_{E_L} m_s^{\rm GD}(p)\,\mathrm{d}p, \tag{38}\\ \overline{m}_M^{\rm GD}(s,\delta) &:=\frac{1}{1-2\delta}\int_{E_M}m_s^{\rm GD}(p)\,\mathrm{d}p,\tag{39}\\ \overline{m}_R^{\rm GD}(s,\delta) &:=\frac{1}{\delta}\int_{E_R}m_s^{\rm GD}(p)\,\mathrm{d}p.\tag{40} \end{align}\] since \(m_s^{\rm GD}\) integrates to one over an interval of length one, these are regional densities, not regional probability masses; an average may therefore exceed one when influence is concentrated in a short region.

Primacy means \(\overline{m}_L^{\rm GD}>\overline{m}_M^{\rm GD}\), recency means \(\overline{m}_R^{\rm GD}>\overline{m}_M^{\rm GD}\), and Lost-in-the-Middle means both. Define the unregularized boundary-minus-middle gap \[\label{eq:gd-lim-gap} \Delta_\delta(s):= \min\{\overline{m}_L^{\rm GD},\overline{m}_R^{\rm GD}\}-\overline{m}_M^{\rm GD},\tag{41}\] the stabilized index \[\label{eq:gd-lim-index} \mathsf{LIM}_{\delta}(s) :=\frac{\Delta_\delta(s)}{\min\{\overline{m}_L^{\rm GD},\overline{m}_R^{\rm GD}\}+\varepsilon_0}, \qquad \varepsilon_0>0,\tag{42}\] and the symmetric scale-free contrast \[\label{eq:gd-lim-symmetric-contrast} \mathsf{SC}_{\delta}(s) :=\frac{\Delta_\delta(s)}{\min\{\overline{m}_L^{\rm GD},\overline{m}_R^{\rm GD}\}+\overline{m}_M^{\rm GD}+\varepsilon_0}.\tag{43}\] All three quantities have the same sign. The magnitude of \(\mathsf{LIM}_\delta\) depends on \(\varepsilon_0\), so numerical reports in Section 4 state its value and also report \(\Delta_\delta\) and \(\mathsf{SC}_\delta\).

3.3 Primacy from the Volterra adjoint cone↩︎

The causal support condition in 9 implies \[\label{eq:causal-support-recalled} A_\theta^h(p,q;X)=0 \qquad\text{whenever }q>p.\tag{44}\] Equations 8 and 9 show that every quantity retained at output position \(p\) depends only on the prefix \(X|_{[0,p]}\). The feed-forward sublayer is tokenwise, and any retained positional statistic is likewise prefix-causal. Differentiation therefore preserves this lower-triangular dependence, so no additional causal-derivative-closure assumption is needed. Consequently, the Fréchet derivative of causal attention is lower triangular in the position variable. More precisely, a linearized causal component acting on a perturbation \(\xi\) has the Volterra–local form \[\label{eq:volterra-linearization} (\mathcal{K}_t\xi)(p) =\int_0^p K_t(p,q)\xi(q)\,\mathrm{d}q +\mathcal{B}_t(p)\xi(p),\tag{45}\] where the integral term transports key and value perturbations from positions \(q\le p\), while \(\mathcal{B}_t(p)\) collects terms depending only on the perturbation at the output position \(p\). We now derive this decomposition directly from the normalized attention map.

For one head, set \[\label{eq:one-head-causal-map} \mathcal{T}_h(X)(p) :=\int_0^p A_h(p,q;X)\mathsf V_h(X(q))\,\mathrm{d}q, \qquad p>0,\tag{46}\] where \[\mathsf Q_h(z):=W_{Q,h}\mathrm{LN}(z),\qquad \mathsf K_h(z):=W_{K,h}\mathrm{LN}(z),\qquad \mathsf V_h(z):=W_{V,h}\mathrm{LN}(z).\] Writing \[\label{eq:one-head-score-and-kernel} s_h(p,q;X) :=\frac{\langle\mathsf Q_h(X(p)),\mathsf K_h(X(q))\rangle}{\sqrt{d_k}}, \qquad A_h(p,q;X) :=\frac{e^{s_h(p,q;X)}}{\int_0^p e^{s_h(p,r;X)}\,\mathrm{d}r},\tag{47}\] the first variation of \(\mathcal{T}_h\) is \[\begin{align} D\mathcal{T}_h(X)[\xi](p) &=\int_0^p A_h(p,q;X) D\mathsf V_h(X(q))[\xi(q)]\,\mathrm{d}q \notag\\ &\quad+\int_0^p DA_h(p,q;X)[\xi]\, \mathsf V_h(X(q))\,\mathrm{d}q . \label{eq:attention-first-variation-schematic} \end{align}\tag{48}\] The score variation separates into a query contribution, depending on \(\xi(p)\), and a key contribution, depending on \(\xi(q)\): \[\begin{align} \delta s_{h,Q}(p,q)[\eta] &:=\frac{\langle D\mathsf Q_h(X(p))[\eta],\mathsf K_h(X(q))\rangle}{\sqrt{d_k}},\notag\\ \delta s_{h,K}(p,q)[\eta] &:=\frac{\langle\mathsf Q_h(X(p)),D\mathsf K_h(X(q))[\eta]\rangle}{\sqrt{d_k}}. \label{eq:head-score-variation} \end{align}\tag{49}\] Thus \[Ds_h(p,q;X)[\xi] =\delta s_{h,Q}(p,q)[\xi(p)] +\delta s_{h,K}(p,q)[\xi(q)].\] Differentiating both the numerator and the normalizing denominator in 47 gives \[\label{eq:normalized-kernel-variation} DA_h(p,q;X)[\xi] =A_h(p,q;X) \left( Ds_h(p,q;X)[\xi] -\int_0^p A_h(p,r;X)Ds_h(p,r;X)[\xi]\,\mathrm{d}r \right).\tag{50}\] Let \[\label{eq:one-head-attention-mean-value} \overline{V}_h(p):=\mathcal{T}_h(X)(p) =\int_0^p A_h(p,r;X)\mathsf V_h(X(r))\,\mathrm{d}r.\tag{51}\] Substituting 50 into 48 and using the normalization \(\int_0^p A_h(p,q;X)\,\mathrm{d}q=1\) yields the exact identity \[\label{eq:attention-linearization-exact} D\mathcal{T}_h(X)[\xi](p) =\mathsf B_{h,Q}(p)\xi(p) +\int_0^p\mathsf K_{h,X}(p,q)\xi(q)\,\mathrm{d}q,\tag{52}\] where the nonlocal causal kernel is the linear map \[\begin{align} \mathsf K_{h,X}(p,q)\eta :=A_h(p,q;X)\Bigl(&D\mathsf V_h(X(q))[\eta]\notag\\ &+\delta s_{h,K}(p,q)[\eta] \bigl(\mathsf V_h(X(q))-\overline{V}_h(p)\bigr)\Bigr), \label{eq:attention-nonlocal-kernel} \end{align}\tag{53}\] and the local query operator is \[\label{eq:attention-local-operator} \mathsf B_{h,Q}(p)\eta :=\int_0^p A_h(p,q;X)\, \delta s_{h,Q}(p,q)[\eta] \bigl(\mathsf V_h(X(q))-\overline{V}_h(p)\bigr)\,\mathrm{d}q.\tag{54}\] All layer-normalization derivatives are contained in \(D\mathsf Q_h\), \(D\mathsf K_h\), and \(D\mathsf V_h\). Equation 53 is supported only on \(q\le p\), whereas 54 multiplies \(\xi(p)\) and is therefore local in position. The formulas are stated for \(p>0\); at the left endpoint the derivative is understood through the operator trace in 11 , consistently with the treatment of the nonlinear attention map.

After concatenating the headwise derivatives, applying the output projection \(W_O\), adding the tokenwise feed-forward derivative, and including the prefix-causal auxiliary contributions retained in the modeled architecture, one obtains \[\label{eq:multihead-volterra-derivative} D\mathscr F_{\theta_t}(X_t)[\xi](p) =\int_0^p K_t(p,q)\xi(q)\,\mathrm{d}q +\mathcal{B}_t(p)\xi(p).\tag{55}\] Here \(K_t\) contains the projected headwise kernels 53 and any other causal nonlocal derivative terms, while \(\mathcal{B}_t\) contains the projected query operators 54 , the feed-forward derivative, and other tokenwise local terms. The residual identity is not part of \(D\mathscr F_{\theta_t}(X_t)\) and is therefore not included in \(\mathcal{B}_t\). It enters the derivative of a discrete residual step as the separate identity factor and appears in continuous depth as the direct terminal term \(P_T\) in the Duhamel formula 68 . This separation is necessary for distinguishing the Volterra primacy channel from residual transmission.

The Hilbert-space adjoint of 45 satisfies \[\label{eq:volterra-adjoint} (\mathcal{K}_t^*a)(q) = \int_q^1 K_t(p,q)^*a(p)\,\mathrm{d}p+\mathcal{B}_t(q)^*a(q).\tag{56}\] Thus an early perturbation \(q\) has a larger future causal cone \(\{p:p\ge q\}\) than a middle perturbation, and the adjoint formula shows this transport explicitly: a terminal covector value \(a(p)\) at a later position \(p\ge q\) contributes to the earlier position \(q\) through the term \(K_t(p,q)^*a(p)\), and the total contribution is obtained by integrating over all future positions \(p\in[q,1]\). Thus loss sensitivity located later in the context is pushed backward through the causal edges to every earlier position that could have influenced it. This is the deterministic adjoint mechanism behind primacy.

One can make the cone effect visible through the cone mass \[\label{eq:cone-mass} \mathcal{C}(q) := \int_q^1 \|K_t(p,q)\|_{\rm op}^2\,\mathrm{d}p .\tag{57}\] In practice one can estimate \(\mathcal{C}(q)\) from the finite-token Jacobian of the attention block. Let \(\widehat K_t(i,j)\) denote the off-diagonal block of the Jacobian mapping a perturbation at token \(j\) into the attention output at token \(i\), with \(j\le i\). Then the discrete cone-mass estimator is \[\label{eq:finite-cone-mass-estimator} \widehat{\mathcal{C}}_t(j):=\sum_{i=j}^{L}\Delta p\,\|\widehat K_t(i,j)\|_{\rm op}^2,\qquad \Delta p:=L^{-1}.\tag{58}\] When forming the full Jacobian is too expensive, the operator norms can be replaced by Hutchinson-type probe estimates, for independent unit-variance probes \(\xi^{(r)}\) supported near token \(j\): \[\label{eq:probe-cone-mass-estimator} \widehat{\mathcal{C}}_t(j)\approx \frac{1}{N_{\rm probe}}\sum_{r=1}^{N_{\rm probe}}\sum_{i=j}^{L}\Delta p\,\|(D\mathrm{Attn}_t(X)\xi^{(r)})(i)\|^2.\tag{59}\] This gives a measurable finite-token proxy for 57 .

The cone length alone gives only an upper bound by Cauchy–Schwarz and does not rule out vector cancellation. Moreover, both the kernel and adjoint are data-dependent. Write \(\omega=(X_0,Y)\) and \[K_{t,\omega}(p,q):=K_t(p,q;X_{t,\omega}), \qquad P_{t,\omega}:=P_t^{\theta,X_0,Y}.\] For a fixed depth time the random-kernel identity is \[\begin{align} Q_t(q) &:=\mathbb{E}\left\|\int_q^1K_{t,\omega}(p,q)^*P_{t,\omega}(p)\,\mathrm{d}p\right\|^2\notag\\ &=\int_q^1\!\int_q^1 \operatorname{Re}\mathbb{E}\operatorname{tr}\!\left[ K_{t,\omega}(p,q)^*P_{t,\omega}(p) P_{t,\omega}(r)^*K_{t,\omega}(r,q) \right]\,\mathrm{d}p\,\mathrm{d}r. \label{eq:exact-cone-covariance-energy} \end{align}\tag{60}\] The kernel cannot be pulled outside the expectation unless it is deterministic or an appropriate conditional covariance is used. In the present setting the depth-control path \(\theta_t\) is deterministic once the full-batch training trajectory is fixed, but the kernel is evaluated at the example-dependent hidden state \(X_{t,\omega}\). Hence \(K_{t,\omega}(p,q)=K_t(p,q;X_{t,\omega})\) is generally random under \(\omega\sim\mathcal{D}\) (through \(X_0\); the label \(Y\) enters the adjoint rather than the forward state). The kernel is deterministic only in special cases, such as a state-independent linearized attention operator, a data distribution concentrated on a single forward trajectory, or an analysis conditioned on a fixed input trajectory. Therefore 60 , with the kernel retained inside the expectation, is the appropriate identity for the modeled data-averaged influence.

The actual cone channel in the input adjoint is integrated over depth: \[\label{eq:depth-integrated-cone-channel} P_{\rm cone,\omega}(q) :=\int_0^T\zeta(t)\int_q^1 K_{t,\omega}(p,q)^*P_{t,\omega}(p)\,\mathrm{d}p\,\mathrm{d}t.\tag{61}\] Define the joint cross-depth integrand \[\begin{align} \Gamma(t,u,p,r;q) :={}&\zeta(t)\zeta(u)\operatorname{Re}\mathbb{E}\operatorname{tr}\!\left[ K_{t,\omega}(p,q)^*P_{t,\omega}(p)\right.\notag\\[-0.3em] &\left.{}\times P_{u,\omega}(r)^*K_{u,\omega}(r,q) \right]. \label{eq:joint-cross-depth-cone-integrand} \end{align}\tag{62}\] Whenever this integrand is absolutely integrable, Fubini’s theorem gives the exact identity \[\begin{align} Q_{\rm cone}(q) &:=\mathbb{E}\|P_{\rm cone,\omega}(q)\|^2\notag\\ &=\int_0^T\!\int_0^T\!\int_q^1\!\int_q^1 \Gamma(t,u,p,r;q)\,\mathrm{d}p\,\mathrm{d}r\,\mathrm{d}t\,\mathrm{d}u. \label{eq:full-depth-cone-energy} \end{align}\tag{63}\]

Proposition 14 (Conditional primacy under joint non-cancellation). Let \(E_L=[0,\delta)\) and \(E_M=[\delta,1-\delta)\), and define \[\label{eq:cone-geometric-region-factor} A_B:=\frac{1}{|B|}\int_B(1-q)^2\,\mathrm{d}q.\qquad{(33)}\] Assume absolute integrability in 63 and constants \(\kappa_L,K_M>0\) such that \[\begin{align} \Gamma(t,u,p,r;q)&\ge\kappa_L &&\text{for a.e. }q\in E_L, (t,u)\in(0,T)^2, (p,r)\in(q,1)^2, \label{eq:cone-left-joint-coercivity}\\ |\Gamma(t,u,p,r;q)|&\le K_M &&\text{for a.e. }q\in E_M, (t,u)\in(0,T)^2, (p,r)\in(q,1)^2. \label{eq:cone-middle-joint-upper-bound} \end{align}\] {#eq: sublabel=eq:eq:cone-left-joint-coercivity,eq:eq:cone-middle-joint-upper-bound} Then \[\begin{align} \frac{1}{|E_L|}\int_{E_L}Q_{\rm cone}(q)\,\mathrm{d}q &\ge \kappa_LT^2A_{E_L}, \label{eq:cone-left-actual-lower-bound}\\ \frac{1}{|E_M|}\int_{E_M}Q_{\rm cone}(q)\,\mathrm{d}q &\le K_MT^2A_{E_M}. \label{eq:cone-middle-actual-upper-bound} \end{align}\] {#eq: sublabel=eq:eq:cone-left-actual-lower-bound,eq:eq:cone-middle-actual-upper-bound} Consequently, if \[\label{eq:cone-actual-primacy-condition} \kappa_LA_{E_L}>K_MA_{E_M},\qquad{(34)}\] the actual depth-integrated cone energy, not merely its lower bound, has a strict left–middle regional advantage.

Proof. For \(q\in E_L\), integrate ?? over \((0,T)^2\times(q,1)^2\) in 63 ; this gives \(Q_{\rm cone}(q)\ge\kappa_LT^2(1-q)^2\). For \(q\in E_M\), absolute integrability and ?? give \(Q_{\rm cone}(q)\le K_MT^2(1-q)^2\). Regional averaging proves ?? –?? , and ?? gives the strict comparison. ◻

Without joint cross-data and cross-depth non-cancellation, and without an upper control on the middle channel, the Volterra formula identifies a possible primacy channel but does not by itself prove primacy.

3.4 Recency from the residual identity channel↩︎

Residual connections provide a distinct transmission channel: they preserve terminal sensitivity through depth, but they generate a right-boundary advantage only when the terminal readout or task protocol is already right biased.

For the split residual layer, write its exact state Jacobian on a bounded trajectory as \[\label{eq:residual-jacobian-remainder} D_{X_k}X_{k+1}=I+\varepsilon A_k+E_k, \qquad A_k:=\zeta(t_k)D\mathscr F_{\theta_k}(X_k), \qquad \|E_k\|_{\rm op}\le C\varepsilon^2.\tag{64}\] The remainder bound is the differentiated form of the split-layer defect. Duality gives \[\label{eq:residual-adjoint-step} P_k=(I+B_k)P_{k+1}, \qquad B_k:=\varepsilon A_k^*+E_k^*.\tag{65}\] Hence the exact backward propagator is \[\label{eq:residual-adjoint-product} P_0=(I+B_0)(I+B_1)\cdots(I+B_{M-1})P_M,\tag{66}\] and its finite algebraic expansion is \[\label{eq:identity-plus-interaction-expansion} P_0=P_M+ \sum_{r=1}^{M}\;\sum_{0\le k_1<\cdots<k_r\le M-1} B_{k_1}\cdots B_{k_r}P_M.\tag{67}\] The first term \(P_M\) is present even if the attention and feed-forward derivative terms are set to zero. It is therefore the purely residual identity channel. In the continuous-depth notation, the same decomposition follows from Duhamel’s formula: \[\label{eq:continuous-adjoint-duhamel} P_0(p) =P_T(p)+\int_0^T \zeta(t)\bigl(D\mathscr F_{\theta_t}(X_t)^*P_t\bigr)(p)\,\mathrm{d}t .\tag{68}\] Thus terminal sensitivity at position \(p\) is copied to depth zero at zeroth order, while the Volterra and feed-forward interactions only add the integral correction.

Let \(E_R=[1-\delta,1]\), \(E_M=[\delta,1-\delta)\), and let \(\Pi_R\) and \(\Pi_M\) denote multiplication by the indicators of these two sets. Recency is not imposed as an additional assumption; in the present framework it is the part of the measured influence density that is explained by the residual transmission of terminal adjoint energy. The residual contribution is therefore defined by \[\label{eq:residual-influence-component} I_{\rm res}(s,p) :=\mathbb{E}_{(X_0,Y)\sim\mathcal{D}} \left[\left\|P_T^{\theta^{\vartheta_s},X_0,Y}(p)\right\|^2\right],\tag{69}\] and its right-middle contrast is the observable quantity \[\label{eq:residual-recency-contrast} \mathsf R_{\rm res}(s,\delta) := \frac{1}{\delta}\int_{1-\delta}^1 I_{\rm res}(s,p)\,\mathrm{d}p - \frac{1}{1-2\delta}\int_{\delta}^{1-\delta} I_{\rm res}(s,p)\,\mathrm{d}p .\tag{70}\] A positive value of 70 means exactly that the average terminal adjoint energy density on \(E_R\) exceeds the corresponding average on \(E_M\). Since the residual identity part in 68 transmits \(P_T\) directly into \(P_0\), any right-heavy terminal energy produced by the loss or readout is already a right-heavy input-sensitivity component before the Volterra correction is added.

In an autoregressive decoder this right-boundary contribution can be expressed directly from the next-token loss without assuming a priori that the query locations are terminal. Let \[\alpha^L=(\alpha_1^L,\ldots,\alpha_L^L), \qquad \alpha_i^L\ge0, \qquad \sum_{i=1}^{L}\alpha_i^L=1,\] be the empirical weight with which the training or evaluation objective reads out the next-token loss at query location \(i\). These weights are useful in practice since different protocols use different query sets: full teacher-forced training averages over many positions, prompt evaluation may use only the last query, and benchmark losses may average over a selected subset. Mathematically, full teacher-forced training corresponds to \(\alpha_i^L=1/L\) for all available query positions, because the loss averages the next-token prediction error at every token. Last-token prompt evaluation corresponds to \(\alpha_L^L=1\) and \(\alpha_i^L=0\) for \(i<L\), because the benchmark conditions on the whole prompt and scores only the prediction after the final prompt token. More generally, if a benchmark scores a subset \(S_L\subset\{1,\ldots,L\}\), then \(\alpha_i^L=|S_L|^{-1}\one_{\{i\in S_L\}}\). The finite-context autoregressive loss can therefore be written as the weighted objective \[\label{eq:autoregressive-terminal-loss} \mathcal{\ell}_{\rm AR}(Y,H(X)) =-\sum_{i=1}^{L}\alpha_i^L Y_{i+1}^{\top}\log\operatorname{softmax}(H(X)_i),\tag{71}\] Equation 71 uses an observed context of \(L\) tokens together with an additional continuation target \(Y_{L+1}\), so it has \(L\) scored query positions. Ordinary teacher forcing on a sequence containing exactly \(L\) observed tokens instead scores only \(i=1,\ldots,L-1\) with targets \(Y_2,\ldots,Y_L\); in that convention set \(\alpha_L^L=0\) and renormalize the weights over the available queries. The continuation convention is used below only when the task supplies or evaluates the token following the full context. The logit residual entering the terminal adjoint is therefore \[\label{eq:autoregressive-logit-residual} G_i^{\rm AR} = \alpha_i^L \bigl(\operatorname{softmax}(H(X_T)_i)-Y_{i+1}\bigr), \qquad i=1,\ldots,L,\tag{72}\] and hence \[\label{eq:autoregressive-terminal-adjoint} P_T =DH(X_T)^*G^{\rm AR}.\tag{73}\] Thus recency is tied to the measured distribution of the query weights \(\alpha^L\), not to an extra structural assumption on their support. Define the discrete right and middle query masses \[\label{eq:discrete-right-middle-sets} E_R^L(\delta):=\{i: i/L\in[1-\delta,1]\}, \qquad E_M^L(\delta):=\{i: i/L\in[\delta,1-\delta)\}.\tag{74}\] \[\label{eq:query-mass-discrete} \omega_R^L(\delta):=\sum_{i\in E_R^L(\delta)}\alpha_i^L, \qquad \omega_M^L(\delta):=\sum_{i\in E_M^L(\delta)}\alpha_i^L.\tag{75}\] A right-biased query/readout geometry means \[\label{eq:right-biased-query-mass} \omega_R^L(\delta)>\omega_M^L(\delta),\tag{76}\] or, more sharply for squared adjoint energy, that \[\label{eq:right-biased-squared-query-mass} \frac{1}{|E_R^L(\delta)|}\sum_{i\in E_R^L(\delta)}(\alpha_i^L)^2 > \frac{1}{|E_M^L(\delta)|}\sum_{i\in E_M^L(\delta)}(\alpha_i^L)^2.\tag{77}\] This condition is an observable property of the loss/readout placement. For example, in last-token prompt evaluation, namely the application scenario in which a model is given the full prompt or retrieved context and the metric scores only the next-token or answer prediction made from the final query position, one has \(\alpha_{L}^L=1\) and \(\alpha_i^L=0\) for \(i<L\), so 76 holds for every fixed \(\delta>0\) once \(L\) is large.

In fully teacher-forced training with uniform token averaging, each query contributes equally, so \(\alpha_i^L=1/L\). Then \(\omega_R^L(\delta)=|E_R^L(\delta)|/L\approx\delta\) and \(\omega_M^L(\delta)=|E_M^L(\delta)|/L\approx1-2\delta\). Moreover, the per-position squared averages satisfy \(|E_R^L|^{-1}\sum_{i\in E_R^L}(\alpha_i^L)^2=|E_M^L|^{-1}\sum_{i\in E_M^L}(\alpha_i^L)^2=L^{-2}\), up to endpoint discretization. Hence the terminal-loss channel by itself does not select the right boundary. This is why the present formulation separates the residual identity mechanism from the task protocol: recency appears only when the measured query/readout weights or residual factors actually make 70 positive.

In the simple tokenwise readout case \(H(X)_i=C_iX(i)\), one obtains \[\label{eq:tokenwise-terminal-adjoint} P_T(i) =C_i^*G_i^{\rm AR} =\alpha_i^L C_i^* \bigl(\operatorname{softmax}(H(X_T)_i)-Y_{i+1}\bigr).\tag{78}\] Consequently the right-minus-middle residual energy satisfies \[\begin{align} &\frac{1}{|E_R^L|}\sum_{i\in E_R^L}\mathbb{E}\|P_T(i)\|^2 - \frac{1}{|E_M^L|}\sum_{i\in E_M^L}\mathbb{E}\|P_T(i)\|^2 \notag\\ &\quad = \frac{1}{|E_R^L|}\sum_{i\in E_R^L}(\alpha_i^L)^2 \mathbb{E}\left\|C_i^*\bigl(\operatorname{softmax}(H(X_T)_i)-Y_{i+1}\bigr)\right\|^2 \notag\\ &\qquad - \frac{1}{|E_M^L|}\sum_{i\in E_M^L}(\alpha_i^L)^2 \mathbb{E}\left\|C_i^*\bigl(\operatorname{softmax}(H(X_T)_i)-Y_{i+1}\bigr)\right\|^2 .\label{eq:finite-recency-energy-contrast} \end{align}\tag{79}\] To make this comparison explicit, define \[\rho_i:=\mathbb{E}\left\|C_i^*\bigl(\operatorname{softmax}(H(X_T)_i)-Y_{i+1}\bigr)\right\|^2.\] Then 79 is \[\frac{1}{|E_R^L|}\sum_{i\in E_R^L}(\alpha_i^L)^2\rho_i - \frac{1}{|E_M^L|}\sum_{i\in E_M^L}(\alpha_i^L)^2\rho_i.\] If the readout residual energies are comparable, say \(0<\rho_-\le \rho_i\le \rho_+<\infty\) and \(\rho_+/\rho_-\) is close to one across the two regions, then the sign is governed mainly by the difference of the squared query-weight averages in 77 . In this notation, the informal statement that the last observed tokens are closest to the terminal query/readout locations should be read as a coefficient-level condition: after the query weights and readout residual energies are combined, the effective coefficients \((\alpha_i^L)^2\rho_i\) have a larger average over \(E_R^L(\delta)\) than over \(E_M^L(\delta)\).

For a local readout \[\label{eq:local-readout-finite} H(X)_i=\sum_{j=1}^{L}C_{ij}X(j), \qquad C_{ij}=0\quad\text{if } |i-j|>r_H,\tag{80}\] Here the condition \(C_{ij}=0\) for \(|i-j|>r_H\) models a readout whose logit at query \(i\) depends only on hidden states within a window of radius \(r_H\) around \(i\). This includes tokenwise readout as the case \(r_H=0\), and it also covers local smoothing, local pooling, or finite-window output heads in practice: these are readouts in which the logit at position \(i\) is formed from a short neighborhood of hidden states rather than from exactly one hidden state. It is a locality assumption on the readout operator, not an additional assumption on the causal attention mask. The terminal adjoint is \[\label{eq:local-readout-adjoint} P_T(j)=\sum_{i=1}^{L}C_{ij}^*G_i^{\rm AR},\tag{81}\] and hence \[\label{eq:local-readout-support} \operatorname{supp} P_T \subset \{j:\operatorname{dist}(j,\operatorname{supp}\alpha^L)\le r_H\}.\tag{82}\] Thus, if \(\alpha^L\) is concentrated in \(E_R^L(\delta)\), then \(G_i^{\rm AR}\) is concentrated at right-side query indices. The formula \(P_T(j)=\sum_i C_{ij}^*G_i^{\rm AR}\) can move this support only to indices \(j\) with \(|i-j|\le r_H\). Therefore the terminal adjoint remains right-biased, with at most an \(r_H\)-neighborhood of spatial spreading.

In continuum notation, let \(\chi\ge0\) be a continuum query/readout density, namely the continuum limit of the discrete weights \(\alpha_i^L\) so that \(\chi(p)\,\mathrm{d}p\) gives the fraction of loss/readout mass assigned to query positions in a small interval around \(p\) with \(\int_0^1\chi(p)\,\mathrm{d}p=1\). The autoregressive readout loss is \[\label{eq:continuum-ar-loss} \mathcal{\ell}_{\rm AR}(Y,H(X)) =-\int_0^1 \chi(p)Y(p)^\top \log\operatorname{softmax}(H(X)(p))\,\mathrm{d}p,\tag{83}\] and the continuum terminal adjoint is \[\label{eq:continuum-ar-terminal-adjoint} P_T =DH(X_T)^*\Bigl(\chi(\cdot) \bigl(\operatorname{softmax}(H(X_T)(\cdot))-Y(\cdot)\bigr)\Bigr).\tag{84}\] The right-bias condition is now the measurable inequality \[\label{eq:continuum-query-right-bias} \frac{1}{\delta}\int_{1-\delta}^1\chi(p)^2\,\mathrm{d}p > \frac{1}{1-2\delta}\int_{\delta}^{1-\delta}\chi(p)^2\,\mathrm{d}p,\tag{85}\] possibly after multiplying by the local readout-residual energy. Under 85 , the residual component 69 is naturally larger near the right boundary, up to the spatial spread of \(DH(X_T)^*\). The theory only requires the measurable contrast 70 ; the autoregressive formulas 7185 explain one standard mechanism by which that contrast can become positive.

For a possibly unbounded data distribution, define the pathwise operator size \[\label{eq:adjoint-operator-bound} \Lambda_s(\omega):= \sup_{0\le t\le T}|\zeta(t)|\, \|D\mathscr F_{\theta_t^{\vartheta_s}}(X_{t,\omega})^*\|_{\mathcal{L}(\mathcal{X})}, \qquad c_s(\omega):=e^{T\Lambda_s(\omega)}-1.\tag{86}\] Whenever \(\Lambda_s(\omega)<\infty\), backward Gronwall gives the pathwise estimate \[\label{eq:residual-stability-bound} \|P_0(\omega)-P_T(\omega)\|_{\mathcal{X}} \le c_s(\omega)\|P_T(\omega)\|_{\mathcal{X}}.\tag{87}\]

Proposition 15 (Regional squared-energy persistence of recency). Let \[\mathcal{A}_B(P):=\frac{1}{|B|}\mathbb{E}\|\Pi_BP\|_{\mathcal{X}}^2\] and assume \[\label{eq:recency-moment-condition} \mathbb{E}\bigl[(2c_s+c_s^2)\|P_T\|_{\mathcal{X}}^2\bigr]<\infty.\qquad{(35)}\] Then, for every measurable region \(B\) of positive length, \[\label{eq:regional-squared-energy-stability} |\mathcal{A}_B(P_0)-\mathcal{A}_B(P_T)| \le \frac{1}{|B|} \mathbb{E}\bigl[(2c_s+c_s^2)\|P_T\|_{\mathcal{X}}^2\bigr].\qquad{(36)}\] Consequently the terminal right bias survives whenever \[\label{eq:regional-recency-survival-condition} \mathsf R_{\rm res}(s,\delta) > \left(\frac{1}{\delta}+\frac{1}{1-2\delta}\right) \mathbb{E}\bigl[(2c_s+c_s^2)\|P_T\|_{\mathcal{X}}^2\bigr].\qquad{(37)}\] Under this condition, \(\mathcal{A}_{E_R}(P_0)>\mathcal{A}_{E_M}(P_0)\).

Proof. Set \(d=P_0-P_T\). The identity \[\|\Pi_B(P_T+d)\|^2-\|\Pi_BP_T\|^2 =2\operatorname{Re}\langle\Pi_BP_T,\Pi_Bd\rangle+\|\Pi_Bd\|^2\] and 87 give pathwise \[\left|\|\Pi_BP_0\|^2-\|\Pi_BP_T\|^2\right| \le(2c_s+c_s^2)\|P_T\|^2.\] Taking expectations proves ?? ; applying it to \(E_R\) and \(E_M\) proves the survival statement. ◻

Thus the residual mechanism preserves a right-heavy terminal component only when its squared-energy gap dominates a quantitatively matched interaction error.

The primacy and recency channels are therefore mathematically distinct. Primacy is generated by the depth-integrated nonlocal Volterra adjoint term \[\label{eq:primacy-channel-term} P_{\rm cone}(q) = \int_0^T\zeta(t)\int_q^1 K_t(p,q)^*P_t(p)\,\mathrm{d}p\,\mathrm{d}t,\tag{88}\] whose integration domain is largest for early \(q\). Recency is transmitted by the residual identity term \[\label{eq:recency-channel-term} P_{\rm res}(p)\sim P_T(p),\tag{89}\] whose size is controlled by the terminal readout and loss geometry. Hence one may define the boundary channel strengths \[\begin{align} \mathsf P_L(s,\delta)&:=\int_0^\delta \mathbb{E}\|P_{\rm cone}(p)\|^2\,\mathrm{d}p,\tag{90}\\ \mathsf P_R(s,\delta)&:=\int_{1-\delta}^1 \mathbb{E}\|P_{\rm res}(p)\|^2\,\mathrm{d}p.\tag{91} \end{align}\] A model can have strong primacy with weak recency when \(\mathsf P_L(s,\delta)\) dominates \(\mathsf P_R(s,\delta)\), strong recency with weak primacy in the opposite regime, or both when the Volterra cone and the residual terminal channel are simultaneously large. Lost-in-the-Middle is most pronounced when both boundary strengths dominate the corresponding middle influence components.

3.5 Lost-in-the-Middle as an exact channel-energy comparison↩︎

Use Duhamel’s formula and the local/nonlocal splitting of the adjoint generator to write \[\label{eq:adjoint-channel-vector-decomposition} P_0=P_{\rm res}+P_{\rm cone}+P_{\rm loc},\tag{92}\] where \(P_{\rm res}:=P_T\), while \(P_{\rm cone}\) and \(P_{\rm loc}\) are the time integrals of the nonlocal Volterra and local generator terms. This is a decomposition by generator terms, not an evolution of three independent signals: both integral components are evaluated using the full adjoint \(P_t\).

For a region \(B\) of positive length define \[\begin{align} J_\alpha(B)&:=\frac{1}{|B|}\mathbb{E}\|\Pi_BP_\alpha\|_{\mathcal{X}}^2, \qquad \alpha\in\{\mathrm{res},\mathrm{cone},\mathrm{loc}\},\tag{93}\\ \Xi(B)&:=\frac{2}{|B|}\sum_{\alpha<\beta} \operatorname{Re}\mathbb{E}\langle\Pi_BP_\alpha,\Pi_BP_\beta\rangle_{\mathcal{X}}. \tag{94} \end{align}\] Then \[\label{eq:exact-regional-influence-decomposition} \mathcal{A}_s(B):=\frac{1}{|B|}\int_B I_{\vartheta_s}(p)\,\mathrm{d}p =J_{\rm res}(B)+J_{\rm cone}(B)+J_{\rm loc}(B)+\Xi(B).\tag{95}\]

Corollary 1 (Exact regional characterization). For \(E_L=[0,\delta)\), \(E_M=[\delta,1-\delta)\), and \(E_R=[1-\delta,1]\), Lost-in-the-Middle holds if and only if \[\begin{align} \sum_\alpha[J_\alpha(E_L)-J_\alpha(E_M)]&>\Xi(E_M)-\Xi(E_L),\label{eq:exact-left-middle-channel-condition}\\ \sum_\alpha[J_\alpha(E_R)-J_\alpha(E_M)]&>\Xi(E_M)-\Xi(E_R).\label{eq:exact-right-middle-channel-condition} \end{align}\] {#eq: sublabel=eq:eq:exact-left-middle-channel-condition,eq:eq:exact-right-middle-channel-condition} This is an exact algebraic characterization, not an independently verifiable mechanism theorem.

For estimable sufficient conditions, suppose that for each region \(B\) and pair \(\alpha<\beta\) one has a correlation bound \[\label{eq:channel-correlation-bound} \left|\frac{1}{|B|}\operatorname{Re}\mathbb{E}\langle\Pi_BP_\alpha,\Pi_BP_\beta\rangle\right| \le \rho_{\alpha\beta,B}\sqrt{J_\alpha(B)J_\beta(B)}, \qquad 0\le\rho_{\alpha\beta,B}\le1.\tag{96}\] Cauchy–Schwarz always permits \(\rho_{\alpha\beta,B}=1\); smaller empirical constants record weak channel alignment.

Theorem 16 (Independently checkable channel-energy criterion). Assume nonnegative upper bounds \[\label{eq:channel-upper-bounds} J_\alpha(B)\le u_{\alpha,B}, \qquad \alpha\in\{\mathrm{res},\mathrm{cone},\mathrm{loc}\},\quad B\in\{E_L,E_M,E_R\},\qquad{(38)}\] and lower bounds \[\label{eq:boundary-channel-lower-bounds} J_{\rm cone}(E_L)\ge \ell_L, \qquad J_{\rm res}(E_R)\ge \ell_R.\qquad{(39)}\] Let \[\label{eq:channel-cross-bound-gamma} \Gamma_B:=2\sum_{\alpha<\beta} \rho_{\alpha\beta,B}\sqrt{u_{\alpha,B}u_{\beta,B}}, \qquad U_M:=u_{\rm res,E_M}+u_{\rm cone,E_M}+u_{\rm loc,E_M}.\qquad{(40)}\] The quantities \(u_{\rm cone,E_M}\) and \(u_{\rm res,E_M}\) are middle-channel upper bounds, while \(u_{\rm loc,B}\) supplies an explicit local-channel smallness control. If \[\label{eq:checkable-double-deficit-conditions} \ell_L>U_M+\Gamma_M+\Gamma_L, \qquad \ell_R>U_M+\Gamma_M+\Gamma_R,\qquad{(41)}\] then \[\overline{m}_L^{\rm GD}>\overline{m}_M^{\rm GD}, \qquad \overline{m}_R^{\rm GD}>\overline{m}_M^{\rm GD},\] and hence \(\Delta_\delta>0\), \(\mathsf{LIM}_\delta>0\), and \(\mathsf{SC}_\delta>0\).

Proof. Equation 96 and the upper bounds imply \(|\Xi(B)|\le\Gamma_B\). Therefore \[\mathcal{A}_s(E_L)\ge \ell_L-\Gamma_L, \qquad \mathcal{A}_s(E_R)\ge \ell_R-\Gamma_R,\] whereas \[\mathcal{A}_s(E_M)\le U_M+\Gamma_M.\] The two inequalities in ?? give strict left–middle and right–middle separation. Division by the common positive normalizer \(Z_s\) preserves these inequalities. ◻

The constants in Theorem 16 can be estimated from finite-token channel vectors or bounded analytically. The criterion can be conservative when only the universal choice \(\rho=1\) is available, but unlike Corollary 1, it does not assume the desired regional differences themselves.

3.6 Explicit counterexamples to unconditional shape claims↩︎

The following finite-dimensional examples isolate why each structural slogan needs additional hypotheses.

  1. Causal masking without primacy. Take a causally masked attention block with \(W_V=0\) and a tokenwise identity residual path. The causal mask is present, but \(D\mathrm{Attn}=0\). For a spatially uniform terminal adjoint, \(P_0=P_T\) is uniform, so no left advantage occurs.

  2. Residual connections without recency. Take the identity residual network \(X_{k+1}=X_k\) with a uniformly distributed teacher-forced loss. Then \(P_0=P_T\) and the influence profile is uniform. The residual identity transmits terminal geometry but does not create a right bias.

  3. Equal observability without equal influence. With two positions, identity dynamics, and identical observation maps, \(G_1=G_2=I\). If the task covectors are \(P_T(1)=2e_1\) and \(P_T(2)=e_1\), then the influence energies are \(4\) and \(1\) despite exactly equal observability Gramians.

  4. Negative cross terms can destroy a U-shape. On three representative regions let \(P_{\rm res}=(1,\tfrac12,1)\) and \(P_{\rm cone}=(-1,\tfrac12,-1)\) in a scalar feature space. Each channel separately has energies \((1,\tfrac14,1)\), but their sum has energy \((0,1,0)\). Boundary anti-correlation reverses the channel-energy conclusion, demonstrating why the cross terms in 95 cannot be discarded.

3.7 Gradient-flow remedies↩︎

The deterministic formulation suggests remedies that act directly on the influence density rather than indirectly on retrieval scores. Let \(\nu\) be a target positional density selected from the intended deployment query distribution, task semantics, or a constrained robustness objective. The uniform choice \(\nu\equiv1\) is a neutral diagnostic default. For the continuum squared-density penalty, assume \(\nu\in L^2(0,1)\), the \(L^4\) adjoint-sensitivity hypothesis ?? , and the nondegeneracy condition ?? ; Proposition 12 then makes the following objective finite and differentiable. Without these extra hypotheses, the squared penalty is asserted only at finite token length (or after a stated spatial smoothing). A direct penalty is \[\label{eq:gd-target-penalty} R_{\rm bal}(\vartheta) = R(\vartheta) + \frac{\beta}{2} \int_0^1 \left( \frac{I_\vartheta(p)}{\int_0^1 I_\vartheta(q)\,\mathrm{d}q} -\nu(p) \right)^2 \,\mathrm{d}p .\tag{97}\] A finite-token implementation is the following influence-balancing procedure. Algorithm 1: target influence-density balancing.

  1. Choose a target density \(\nu_i\approx\nu(p_i)\), a regularization strength \(\beta>0\), and a batch \(\mathcal{B}\).

  2. For each \((X_0,Y)\in\mathcal{B}\), run the forward residual Transformer and compute the usual task loss \(\mathcal{L}_{\mathcal{B}}(\vartheta)\).

  3. Backpropagate the terminal adjoint through depth and record the tokenwise input-adjoint energies \[E_b(p_i):=\|P_{0,b}^{\vartheta}(p_i)\|^2, \qquad (X_0,Y)=b\in\mathcal{B} .\]

  4. Estimate and normalize the positional influence by \[\widehat I_\vartheta(p_i):=\frac{1}{|\mathcal{B}|}\sum_{b\in\mathcal{B}}E_b(p_i), \qquad \widehat m_\vartheta(p_i):= \frac{\widehat I_\vartheta(p_i)}{\sum_j\Delta p_j\widehat I_\vartheta(p_j)+\epsilon_{\rm inf}}, \qquad \epsilon_{\rm inf}>0.\]

  5. Minimize the empirical balanced objective \[\label{eq:discrete-target-influence-penalty} \widehat R_{\rm bal}(\vartheta) :=\mathcal{L}_{\mathcal{B}}(\vartheta) +\frac{\beta}{2}\sum_i\Delta p_i \bigl(\widehat m_\vartheta(p_i)-\nu_i\bigr)^2 .\tag{98}\]

Differentiating the penalty through \(\widehat m_\vartheta\) requires derivatives of an input gradient with respect to model parameters, hence second-order automatic differentiation or Hessian–vector products. If the gradient through \(\widehat m_\vartheta\) is stopped, the penalty contributes no gradient to the same optimization objective. A stop-gradient version must therefore be implemented as an outer loop: first measure \(\widehat m\), then update separate positional weights as in Algorithm 2, and only then optimize the task loss. The stabilizer \(\epsilon_{\rm inf}\) prevents division by nearly vanishing total adjoint energy.

Remark 17. The additional term in 97 generally introduces an intentional regularization bias* relative to the minimizer of the original task risk. To make this precise, define \[\label{eq:influence-discrepancy-functional} \Phi_\nu(\vartheta) :=\frac{1}{2}\int_0^1\bigl(m_\vartheta(p)-\nu(p)\bigr)^2\,\mathrm{d}p, \qquad m_\vartheta(p):=\frac{I_\vartheta(p)}{\int_0^1I_\vartheta(q)\,\mathrm{d}q},\tag{99}\] so that \(R_{\rm bal}=R+\beta\Phi_\nu\). Let \(\vartheta_0\) be an isolated local minimizer of \(R\), assume that \(H_0:=\nabla^2R(\vartheta_0)\) is nonsingular, and let \(\vartheta_\beta\) denote the nearby stationary point of \(R_{\rm bal}\). The implicit-function theorem gives \[\label{eq:influence-regularization-bias-expansion} \vartheta_\beta-\vartheta_0 =-\beta H_0^{-1}\nabla\Phi_\nu(\vartheta_0) +O(\beta^2).\tag{100}\] Hence the balanced objective changes the fitted parameter at first order unless \(\nabla\Phi_\nu(\vartheta_0)=0\), which occurs, for example, when the original optimum already matches the target influence profile locally. Under local \(\mu\)-strong convexity of \(R\) and a bound \(\|\nabla\Phi_\nu\|\le G\), one obtains the stability estimate \[\label{eq:influence-regularization-parameter-shift-bound} \|\vartheta_\beta-\vartheta_0\| \le \frac{\beta G}{\mu},\tag{101}\] showing explicitly how the bias is controlled by \(\beta\).*

This is objective bias, not bias of the population gradient: exact differentiation of 97 gives the true gradient of the modified risk. A separate finite-batch estimation bias can arise in 98 , because normalization is nonlinear and generally \[\mathbb{E}\!\left[ \frac{\widehat I_\vartheta(p_i)}{\sum_j\Delta p_j\widehat I_\vartheta(p_j)+\epsilon_{\rm inf}} \right] \ne \frac{I_\vartheta(p_i)}{\sum_j\Delta p_j I_\vartheta(p_j)+\epsilon_{\rm inf}}.\] Large batches, moving-average estimates, and an independently estimated denominator reduce this ratio-estimation effect. In applications, \(\beta\) should therefore be selected by a task-performance-versus-balance tradeoff, annealed toward zero when asymptotic task-risk consistency is required, or replaced by a constrained formulation \(\Phi_\nu(\vartheta)\le\tau\) with a tolerable imbalance level \(\tau\).

A mask-aware remedy is to compensate the causal-cone asymmetry by a position weight \(w(p)>0\) in the loss: \[\label{eq:weighted-cross-entropy} \mathcal{\ell}_w(Y,H(X)) = -\int_0^1 w(p)Y(p)^\top\log\operatorname{softmax}(H(X)(p))\,\mathrm{d}p .\tag{102}\] The idealized choice is to increase \(w\) on positions whose influence density is too small. A practical finite-token update can be written as follows. Algorithm 2: mask-aware positional reweighting.

  1. Fix a target density \(\nu_i\) and initialize positive weights \(w_i^{(0)}>0\), for example \(w_i^{(0)}=1\).

  2. At training epoch \(n\), estimate \(\widehat m^{(n)}(p_i)\) using the adjoint-energy procedure of Algorithm 1.

  3. Update the weights by a damped multiplicative rule \[\label{eq:mask-aware-weight-update} \widetilde{w}_i^{(n+1)} =\operatorname{clip}_{[w_{\min},w_{\max}]} \left( w_i^{(n)}\exp\{\eta_w(\nu_i-\widehat m^{(n)}(p_i))\} \right), \qquad w_i^{(n+1)}= \frac{\widetilde{w}_i^{(n+1)}}{\sum_j\Delta p_j\widetilde{w}_j^{(n+1)}}.\tag{103}\]

  4. Train the next epoch using the discrete weighted loss \[\label{eq:discrete-weighted-cross-entropy} \widehat{\mathcal{\ell}}_w =-\sum_i\Delta p_i\,w_i^{(n+1)}Y_i^\top \log\operatorname{softmax}(H(X)_i).\tag{104}\]

The update 103 increases the loss weight where the measured influence \(\widehat m^{(n)}\) is below the target and decreases it where influence is excessive. The renormalization in 103 fixes the total positional mass of the loss. The update is an outer-loop response rule rather than a gradient step on \(\mathsf{LIM}_\delta\). To express this distinction, let \(\vartheta^+(w)\) denote the parameter obtained after the next training stage using weights \(w\), and define the induced influence-response map \[\label{eq:influence-response-map} \mathcal{M}(w):=m_{\vartheta^+(w)}^{\rm GD}, \qquad \mathcal{L}_\delta(w):=\mathsf{LIM}_\delta(\vartheta^+(w)).\tag{105}\] Ignoring clipping for a first-order calculation and writing \(d_i^{(n)}:=\nu_i-\widehat m^{(n)}(p_i)\), normalization of the multiplicative step gives \[\label{eq:normalized-multiplicative-first-order-step} w_i^{(n+1)}-w_i^{(n)} =\eta_w w_i^{(n)} \left( d_i^{(n)}- \sum_j\Delta p_j w_j^{(n)}d_j^{(n)} \right) +O(\eta_w^2).\tag{106}\] Even though this direction increases relative weight where the current influence is below target, the next influence is \(\mathcal{M}(w^{(n+1)})\), not the current \(\widehat m^{(n)}\). Consequently, \[\label{eq:lim-outer-loop-first-order-change} \mathcal{L}_\delta(w^{(n+1)})-\mathcal{L}_\delta(w^{(n)}) =\nabla_w\mathcal{L}_\delta(w^{(n)})^\top \bigl(w^{(n+1)}-w^{(n)}\bigr) +O\!\left(\|w^{(n+1)}-w^{(n)}\|^2\right),\tag{107}\] and the sign of the leading term is not determined by 103 . A monotone decrease would require an additional response-alignment condition, for example \[\label{eq:lim-response-alignment-condition} \nabla_w\mathcal{L}_\delta(w^{(n)})^\top \bigl(w^{(n+1)}-w^{(n)}\bigr) \le -c\|w^{(n+1)}-w^{(n)}\|^2\tag{108}\] for some \(c>0\), or a suitable contraction property of \(\mathcal{M}\). Neither convexity of the original task risk nor such an alignment property is assumed here. Therefore Algorithm 2 is a diagnostic-driven heuristic: clipping, damping, and validation can stabilize it, but no monotone decrease of \(\mathsf{LIM}_\delta\) is claimed.

Remark 18. The weighting in 102 changes the positional measure under which prediction error is minimized. If \[\label{eq:positionwise-risk-density} r_\vartheta(p) :=\mathbb{E}\!\left[-Y(p)^\top \log\operatorname{softmax}(H_\vartheta(X)(p))\right],\qquad{(42)}\] then the unweighted and weighted population risks are \[\label{eq:weighted-risk-change-of-measure} R(\vartheta)=\int_0^1r_\vartheta(p)\,\mathrm{d}p, \qquad R_w(\vartheta)=\int_0^1w(p)r_\vartheta(p)\,\mathrm{d}p.\qquad{(43)}\] Unless \(w\) is constant, the two objectives need not have the same minimizer when parameters are shared across positions or the model class is misspecified. Thus positional reweighting generally introduces an intentional bias toward the positions receiving larger weight. This effect is absent in the idealized unrestricted, well-specified setting in which the model can realize the Bayes conditional distribution separately at every position: since cross-entropy is a strictly proper scoring rule, multiplication by any positive \(w(p)\) leaves the pointwise Bayes predictor unchanged and changes only the relative statistical efficiency across positions.

The weighting is not biased relative to a desired deployment risk when it is used as importance weighting. If training positions are distributed according to a density \(\pi(p)>0\) and the target deployment density is \(\tau(p)\), choosing \[\label{eq:positional-importance-weight} w(p)=\frac{\tau(p)}{\pi(p)}\qquad{(44)}\] makes \(R_w\) equal to the target-position risk, provided the usual support and integrability conditions hold. For the original uniform-position risk, however, any nonconstant adaptive weight changes the estimand. Clipping in 103 controls variance and extreme optimization pressure, while normalization controls scale but does not remove this change of estimand. Accordingly, the weighted objective should be evaluated against the original unweighted validation risk as well as the positional-balance diagnostic.

A residual-aware diagnostic can be formulated through a finite-token observability Gramian. Stack the hidden states at layer \(k\) as a vector in \(\mathbb{R}^{Ld_x}\), and let \[\label{eq:finite-token-position-injection} E_i:\mathbb{R}^{d_x}\longrightarrow\mathbb{R}^{Ld_x}\tag{109}\] be the canonical injection that places a feature perturbation in token block \(i\) and sets all other token blocks to zero. If \(\Phi_{k\leftarrow0}:=D_{X_0}X_k\) denotes the tangent propagator of the finite residual network from the input to layer \(k\), define \[\label{eq:finite-token-position-jacobian} J_{k\leftarrow0,i}:=\Phi_{k\leftarrow0}E_i \in\mathcal{L}(\mathbb{R}^{d_x},\mathbb{R}^{Ld_x}).\tag{110}\] Thus \(J_{k\leftarrow0,i}\xi\) is the full layer-\(k\) hidden-state variation generated by an input perturbation \(\xi\) applied only at token \(i\).

Let \(C_k:\mathbb{R}^{Ld_x}\to\mathbb{R}^{r_k}\) be a chosen linear observation map. It may, for example, select a collection of token coordinates, linearize an intermediate readout, or retain only the final-token representation. The observation protocol must be fixed before positions are compared. A task-aligned choice is \(C_k=DH_k(X_k)\) for an auxiliary readout \(H_k\), or a fixed final-query selector when the evaluation task reads only the last token. Different \(C_k\) produce different Gramians; observability equalization is therefore relative to the selected task output and is not an intrinsic model property. A non-task-aligned coordinate selector remains a legitimate structural diagnostic, but it should not be interpreted as equalization of predictive influence. The positionwise finite-horizon observability Gramian is \[\label{eq:finite-token-observability-gramian} G_i:=\sum_{k=0}^{M-1}\varepsilon\, J_{k\leftarrow0,i}^*C_k^*C_kJ_{k\leftarrow0,i} \in\mathbb{R}^{d_x\times d_x}.\tag{111}\] It is symmetric positive semidefinite and satisfies, for every \(\xi\in\mathbb{R}^{d_x}\), \[\label{eq:finite-token-observability-quadratic-form} \xi^*G_i\xi =\sum_{k=0}^{M-1}\varepsilon\, \|C_kJ_{k\leftarrow0,i}\xi\|^2.\tag{112}\] Hence \(G_i\) measures how strongly perturbations inserted at token \(i\) remain visible through the selected observations over depth. Equalizing the entire matrices leads to \[\label{eq:gd-observability-penalty} \Omega_{\rm obs} :=\sum_{i=1}^{L}\Delta p_i\|G_i-\overline{G}\|_{\rm HS}^2, \qquad \overline{G}:=\sum_{i=1}^{L}\Delta p_iG_i, \qquad \Delta p_i=\frac{1}{L}.\tag{113}\] The matrix penalty controls directional observability, since \(\|G_i-\overline{G}\|_{\rm HS}^2\) compares all eigen-directions, not only total energy. Forming every \(G_i\), however, generally requires \(d_x\) tangent propagations per position and is therefore practical only for small hidden dimension or restricted token subsets.

For large models, use randomized trace estimation. Let \(\xi_1,\ldots,\xi_{N_{\rm probe}}\) be independent isotropic probes satisfying \[\label{eq:isotropic-observability-probes} \mathbb{E}\xi_r=0, \qquad \mathbb{E}[\xi_r\xi_r^*]=I_{d_x};\tag{114}\] standard Gaussian or Rademacher vectors are admissible. Define \[\label{eq:discrete-observability-trace-estimator} \widehat g_i :=\frac{1}{N_{\rm probe}}\sum_{r=1}^{N_{\rm probe}} \sum_{k=0}^{M-1}\varepsilon\, \|C_kJ_{k\leftarrow0,i}\xi_r\|^2.\tag{115}\] Using 112 and the trace identity \(\mathbb{E}[\xi_r^*G_i\xi_r]=\operatorname{tr}(G_i\mathbb{E}[\xi_r\xi_r^*])\), one obtains \[\label{eq:observability-trace-unbiasedness} \mathbb{E}\widehat g_i=\operatorname{tr}G_i.\tag{116}\] Thus \(g_i:=\operatorname{tr}G_i\) measures the total observability of an isotropic perturbation at token \(i\). A scalar trace-balancing target is \[\label{eq:true-trace-observability-penalty} \Omega_{\rm obs}^{\rm tr} :=\sum_{i=1}^{L}\Delta p_i(g_i-\overline{g}_{\rm true})^2, \qquad \overline{g}_{\rm true}:=\sum_{i=1}^{L}\Delta p_i g_i,\tag{117}\] and its computational surrogate is \[\label{eq:discrete-observability-penalty} \widehat\Omega_{\rm obs}^{\rm tr} :=\sum_{i=1}^{L}\Delta p_i(\widehat g_i-\overline{g})^2, \qquad \overline{g}:=\sum_{i=1}^{L}\Delta p_i\widehat g_i.\tag{118}\] Although each \(\widehat g_i\) is unbiased, the squared empirical penalty is generally upward biased because of probe variance. This does not prevent its use as a stochastic objective; using common probes across positions, increasing \(N_{\rm probe}\), or averaging estimates across iterations reduces its variance. The trace penalty equalizes total local observability but cannot detect anisotropy between feature directions; the Hilbert–Schmidt penalty 113 should therefore be used only when matrix-valued estimates are actually available.

Algorithm 3: randomized observability balancing.

  1. Choose observation maps \(C_k\), a set of monitored layers \(\mathcal{K}\subset\{0,\ldots,M-1\}\), monitored token positions \(\mathcal{S}\subset\{1,\ldots,L\}\), a probe count \(N_{\rm probe}\), and a regularization strength \(\lambda_{\rm obs}\ge0\).

  2. Run the usual forward pass and retain the checkpoints needed to evaluate Jacobian–vector products through the residual layers.

  3. Draw isotropic probes \(\xi_r\). For each \(i\in\mathcal{S}\) and each probe, inject \(E_i\xi_r\) at the input and propagate it with forward-mode automatic differentiation to compute \(J_{k\leftarrow0,i}\xi_r\) for \(k\in\mathcal{K}\).

  4. Accumulate \[\widehat g_i =\frac{1}{N_{\rm probe}}\sum_{r=1}^{N_{\rm probe}} \sum_{k\in\mathcal{K}}\varepsilon_k \|C_kJ_{k\leftarrow0,i}\xi_r\|^2,\] form the weighted mean \(\overline{g}\), and evaluate \(\widehat\Omega_{\rm obs}^{\rm tr}\) from 118 . If only a diagnostic is required, report the profile \(i\mapsto\widehat g_i\) and stop here.

  5. For regularized training, minimize \[\label{eq:observability-balanced-training-objective} \widehat R_{\rm obs}(\vartheta) :=\mathcal{L}_{\mathcal{B}}(\vartheta) +\frac{\lambda_{\rm obs}}{2}\widehat\Omega_{\rm obs}^{\rm tr}(\vartheta).\tag{119}\] Since \(J_{k\leftarrow0,i}\) itself depends on \(\vartheta\), differentiating the observability term requires differentiating through Jacobian–vector products, i.e. mixed second-order derivatives. A stop-gradient implementation remains a diagnostic and does not regularize the current parameter update.

  6. To control cost, subsample \(\mathcal{S}\) and \(\mathcal{K}\), reuse common probes across positions, and update the observability estimate less frequently than the task-loss gradient. Validate both the original task risk and the positional influence profile, since equal observability is a structural surrogate rather than a guaranteed decrease of \(\mathsf{LIM}_\delta\).

Remark 19 (Bias induced by Algorithm 3). Algorithm 3 can introduce two distinct notions of bias, and they should not be conflated. The first is objective bias. If the randomized observability term is used only as a diagnostic, the fitted parameter is unchanged. If it is included in the training objective 119 , however, the population target becomes \[R_{\lambda_{\rm obs}}(\vartheta) :=R(\vartheta)+\frac{\lambda_{\rm obs}}{2} \Omega_{\rm obs}^{\rm tr}(\vartheta),\] which generally has a different minimizer from the original task risk. More precisely, suppose that \(\vartheta_0\) is an isolated minimizer of \(R\), that \(H_0:=\nabla^2R(\vartheta_0)\) is positive definite, and that \(\Omega_{\rm obs}^{\rm tr}\) is twice continuously differentiable near \(\vartheta_0\). The implicit-function theorem then gives, for sufficiently small \(\lambda_{\rm obs}\), \[\label{eq:algorithm3-objective-bias-expansion} \vartheta_{\lambda_{\rm obs}}-\vartheta_0 =-\frac{\lambda_{\rm obs}}{2} H_0^{-1}\nabla\Omega_{\rm obs}^{\rm tr}(\vartheta_0) +O(\lambda_{\rm obs}^2).\qquad{(45)}\] Thus the leading displacement vanishes only when the unregularized solution is already stationary for the observability penalty, or when \(\lambda_{\rm obs}=0\). This is an intentional structural bias, analogous to the regularization bias discussed for Algorithm 1.

The second issue is estimation bias. Let \(d_i:=\Delta p_i\), \(D:=\operatorname{diag}(d_1,\ldots,d_L)\), \(d=(d_1,\ldots,d_L)^\top\), and \(Q:=D-dd^\top\succeq0\). Writing \(g=(g_1,\ldots,g_L)^\top\) and \(\widehat g=(\widehat g_1,\ldots,\widehat g_L)^\top\), the true and empirical trace penalties are \[\Omega_{\rm obs}^{\rm tr}=g^\top Qg, \qquad \widehat\Omega_{\rm obs}^{\rm tr}=\widehat g^\top Q\widehat g.\] If \(\mathbb{E}\widehat g=g\) and \(\Sigma:=\operatorname{Cov}(\widehat g)\), then \[\label{eq:algorithm3-penalty-estimator-bias} \mathbb{E}\widehat\Omega_{\rm obs}^{\rm tr} =\Omega_{\rm obs}^{\rm tr}+\operatorname{tr}(Q\Sigma) \ge \Omega_{\rm obs}^{\rm tr}.\qquad{(46)}\] Hence unbiased trace estimates do not produce an unbiased squared balancing penalty. Common probes across positions can reduce \(\operatorname{tr}(Q\Sigma)\), but do not in general remove it. An unbiased cross-fitted alternative uses two independent probe batches: \[\label{eq:algorithm3-crossfit-unbiased-penalty} \widetilde{\Omega}_{\rm obs}^{\rm tr} :=(\widehat g^{(a)})^\top Q\widehat g^{(b)}, \qquad \mathbb{E}\widetilde{\Omega}_{\rm obs}^{\rm tr} =\Omega_{\rm obs}^{\rm tr}.\qquad{(47)}\] The cross-fitted estimator need not be nonnegative on an individual batch and can have substantially larger gradient variance because two independent noisy profiles are multiplied. It is therefore intended primarily as an unbiased diagnostic or theoretical comparison. For routine training, the nonnegative biased estimator 118 , with common probes and temporal averaging, is usually the safer objective; using the cross-fitted form as a training loss requires explicit variance control and validation.

Subsampling monitored positions or layers creates a further approximation bias unless the sampled terms are corrected by their inclusion probabilities. Inverse-probability weighting preserves the desired trace sum when the sampling probabilities are known and positive; uncorrected subsampling instead optimizes the observability of the monitored subset. Finally, stopping the gradient through the Jacobian–vector products does not create a biased gradient of the same objective; it removes the observability contribution altogether and leaves Algorithm 3 as a diagnostic. Consequently, the regularized model should be evaluated using the original unweighted task risk, the true or high-probe observability profile, and the positional influence profile.

3.8 Computational complexity of Algorithms 1 and 3↩︎

Let \(L\) be sequence length, \(M\) depth, \(d_x\) hidden dimension, \(S=|\mathcal{S}|\) monitored positions, \(K=|\mathcal{K}|\) monitored layers, and \(R=N_{\rm probe}\). Write \(C_{\rm layer}(L,d_x)\) for the cost of one residual layer; for dense attention it is \(O(L^2d_x+Ld_x^2)\) up to head constants.

Algorithm 1 obtains all positional input-adjoint energies from one ordinary reverse pass, so diagnostic evaluation costs \(O(MC_{\rm layer})\) time and standard activation storage \(O(MLd_x)\) without checkpointing. Training through the normalized influence penalty requires mixed second-order differentiation or Hessian–vector products. Its asymptotic order remains \(O(MC_{\rm layer})\) with a larger constant and additional graph retention; checkpointing can reduce memory toward \(O(Ld_x\log M)\) at the cost of recomputation. Importantly, computing the complete input-gradient profile does not require \(L\) separate backward passes.

The forward-mode implementation of Algorithm 3 propagates \(SR\) input tangent directions. Its naive time is \[O\!\left(SR\,M C_{\rm layer}(L,d_x)\right),\] with accumulation at the \(K\) monitored layers costing an additional \(O(SRKd_x)\) and with probes/positions batchable when memory permits. Exact matrix Gramians replace \(R\) by \(d_x\). Reverse mode is preferable when the total number of independent observation cotangents, approximately \(R\sum_{k\in\mathcal{K}}r_k\), is smaller than the number \(SR\) of input directions; forward mode is preferable in the opposite regime, especially for a small monitored position set. Subsampling positions or layers lowers cost in proportion to \(S\) or \(K\) but changes the estimand unless inclusion-probability correction is used.

In summary, the three proposed remedies are: a second-order influence penalty, an outer-loop position reweighting rule, and an observability diagnostic or regularizer. Their optimization effectiveness requires separate analysis and is not implied solely by the structural identities above.

4 Controlled Finite-Token Simulations↩︎

This section reports reproducible numerical experiments for the three remedies in 3.7. The experiments are deliberately small and mechanistic: they test whether each algorithm changes the mathematical quantity that it is designed to control. They are not presented as a language-model benchmark or as evidence that the same hyperparameters transfer to a production Transformer. In particular, Algorithm 2 and Algorithm 3 are evaluated together with the limitations already identified in 108 and Remark 19.

4.1 Simulation model and metrics↩︎

We use \(L=48\) token positions, feature dimension \(d_x=4\), and \(M=12\) residual steps with \(T=1\) and \(\varepsilon=T/M\). Let \[\label{eq:simulation-causal-matrix} C_{ij}:=\frac{\one_{\{j\le i\}}}{i}, \qquad 1\le i,j\le L,\tag{120}\] be the finite Cesàro causal-averaging matrix and let \[\label{eq:simulation-feature-mixing} S= \begin{pmatrix} 1 & 0.25 & 0 & 0\\ 0.15 & 0.8 & 0.2 & 0\\ 0 & 0.15 & 0.6 & 0.2\\ 0 & 0 & 0.2 & 0.45 \end{pmatrix}.\tag{121}\] The linear residual propagator is \[\label{eq:simulation-residual-propagator} B:=I_{Ld_x}+\varepsilon\alpha(C\otimes S), \qquad \alpha=2, \qquad \Phi:=B^M.\tag{122}\] Positive trainable positional gates are represented by \[\label{eq:simulation-input-gate} D(u):=\operatorname{diag}(e^{u_1},\ldots,e^{u_L})\otimes I_{d_x}.\tag{123}\] They may be viewed as a simplified input-embedding or early residual scaling mechanism. The terminal covector combines a final-token query with weak uniform supervision, \[\label{eq:simulation-terminal-covector} r:=\left(\beta e_L+\frac{1-\beta}{L}\mathbf{1}\right)\otimes v, \qquad \beta=0.85, \qquad v=(1,0.5,-0.3,0.2)^\top.\tag{124}\] The input adjoint and normalized positional influence masses are \[\label{eq:simulation-influence-profile} p_0(u)=D(u)\Phi^*r, \qquad m_i(u)= \frac{\|E_i^*p_0(u)\|^2}{\sum_{j=1}^L\|E_j^*p_0(u)\|^2}.\tag{125}\] For Algorithm 2, the uniform part of the terminal positional covector is replaced by \((1-\beta)L^{-1}w\), and the weights are updated by 103 .

For Algorithm 3, every observation map selects the final token and retains all four features. The exact finite-token Gramian is therefore \[\label{eq:simulation-observability-gramian} G_i(u)=e^{2u_i} \sum_{k=0}^{M-1}\varepsilon E_i^*(B^k)^*C_{\rm out}^*C_{\rm out}B^kE_i,\tag{126}\] where \(C_{\rm out}\) extracts the final-token block. We normalize its trace by \[\label{eq:simulation-normalized-observability} q_i(u):= \frac{\operatorname{tr}G_i(u)}{\sum_j\operatorname{tr}G_j(u)}.\tag{127}\] The influence and observability imbalances are \[\label{eq:simulation-imbalance-metrics} \mathcal{E}_{\rm inf}:=\frac{1}{L}\sum_{i=1}^L(Lm_i-1)^2, \qquad \mathcal{E}_{\rm obs}:=\frac{1}{L}\sum_{i=1}^L(Lq_i-1)^2.\tag{128}\] Regional averages use the discrete half-open sets corresponding to \(E_L,E_M,E_R\). We set \(\varepsilon_0=10^{-8}\) and report the discrete analogues of \(\Delta_{0.2}\), \(\mathsf{SC}_{0.2}\), and \(\mathsf{LIM}_{0.2}\). All runs use NumPy and PyTorch in double precision with random seed \(20260717\).

Since \(G_i(u)=e^{2u_i}G_i(0)\), every eigenvalue is rescaled by the same positionwise factor and the condition number is exactly invariant: \[\label{eq:simulation-condition-number-invariance} \kappa_0(G_i(u))=\kappa_0(G_i(0)).\tag{129}\] Thus this parameterization can balance traces while leaving directional anisotropy untouched. For the stated \(d_x=4\) matrices, the baseline regional means of the ordinary condition number are \(171.1\), \(36.7\), and \(16.6\) on the left, middle, and right regions, respectively; the corresponding ranges are \([76.5,482.8]\), \([20.8,70.2]\), and \([1.03,20.2]\). These values are unchanged after Algorithm 3. This spectral diagnostic is more informative than the trace plot alone and confirms that the toy objective does not establish matrix observability balance.

4.2 Optimization protocols↩︎

Algorithm 1 minimizes \[\label{eq:simulation-algorithm1-objective} J_1(u)=\frac{\gamma_1}{2L}\|u\|^2 +\frac{\lambda_1}{2}\mathcal{E}_{\rm inf}(u), \qquad \gamma_1=0.12, \quad \lambda_1=15,\tag{130}\] for 400 Adam steps with learning rate \(0.04\). The quadratic term is a task-fidelity proxy preventing unrestricted gate changes. Algorithm 2 uses 160 outer-loop updates with \(\eta_w=0.5\), clipping interval \([0.15,8]\), and unit-mean renormalization. Algorithm 3 minimizes \[\label{eq:simulation-algorithm3-objective} J_3(u)=\frac{\gamma_3}{2L}\|u\|^2 +\frac{\lambda_3}{2}\mathcal{E}_{\rm obs}(u), \qquad \gamma_3=1, \quad \lambda_3=0.5,\tag{131}\] for 400 Adam steps with learning rate \(0.03\). since \(d_x=4\) is small, this run uses the exact traces rather than stochastic probes; probe error is examined separately below.

4.3 Running results↩︎

The baseline exhibits both boundary channels: the left, middle, and right influence-density averages are respectively \(3.2269\), \(0.1417\), and \(1.3479\). The corresponding gap is \(\Delta_{0.2}=1.2062\), the symmetric contrast is \(\mathsf{SC}_{0.2}=0.8097\), and \(\mathsf{LIM}_{0.2}=0.8949\) for \(\varepsilon_0=10^{-8}\). The table also reports the raw total adjoint energy \(Z\), which is not determined by the normalized profile.

Table 1: Algebraic finite-token results under the normalized discrete convention. The raw energy is \(Z=\sum_i\|E_i^*p_0\|^2\); the stabilizer is \(\varepsilon_0=10^{-8}\).
Method Left Middle Right \(\Delta_{0.2}\) \(\mathsf{SC}_{0.2}\) \(\mathsf{LIM}_{0.2}\) \(\mathcal E_{\rm inf}\) \(Z\)
Baseline 3.2269 0.1417 1.3479 1.2062 0.8097 0.8949 9.2503 4.1487
Algorithm 1 1.0163 0.9960 0.9958 -0.0002 -0.0001 -0.0002 0.0002 9.7981
Algorithm 2 3.0303 0.1631 1.4804 1.3173 0.8015 0.8898 8.5348 3.8028
Algorithm 3 2.8192 0.6671 0.1794 -0.4877 -0.5761 -2.7181 1.3952 1.7226

Algorithm 1 nearly reaches the uniform target: \(\mathcal{E}_{\rm inf}\) falls from \(9.2503\) to \(0.0002\), and the three regional averages become approximately equal. Because a separate \(u_i\) directly rescales each influence energy, this result only confirms the implementation on a directly controllable toy model; it does not verify removal of imbalance by shared Transformer parameters.

The raw normalization values reveal information hidden by the normalized profiles. Algorithm 1 more than doubles \(Z\) while flattening the shape, Algorithm 2 changes \(Z\) only modestly, and Algorithm 3 reduces \(Z\) to about \(41.5\%\) of baseline. A method can therefore appear balanced after normalization while substantially amplifying or suppressing total task sensitivity; both quantities must be interpreted together.

Algorithm 2 produces only a modest change. Its influence imbalance decreases by approximately \(7.7\%\), while \(\mathsf{LIM}_{0.2}\) changes from \(0.8949\) to \(0.8898\). The reason is visible in the response map: the adjustable uniformly supervised component has mass only \(1-\beta=0.15\), whereas the fixed final-token component and the causal propagator dominate the adjoint. Thus the multiplicative rule points in the intended current-profile direction but encounters a poorly aligned model response. This run numerically illustrates why 108 , rather than the sign of the weight discrepancy alone, is needed for a monotonic guarantee.

Algorithm 3 strongly reduces its scalar trace target in this directly rescalable parameterization: the observability imbalance falls from \(39.4921\) to \(0.3392\). The normalized trace-density regional averages move from \((0.2905,0.0351,4.6041)\) to \((1.6292,0.8788,0.7343)\). This is an algebraic check of trace rescaling, not evidence that training a shared-parameter Transformer equalizes observability matrices. Equation 129 shows that the directional condition numbers do not improve at all. Its influence imbalance becomes \(1.3952\), while the right average is pushed below the middle average and \(\mathsf{LIM}_{0.2}=-2.7181\). This is not a contradiction: Algorithm 3 equalizes a scalar structural surrogate, not the influence density or the full Gramian spectrum. The result confirms the implementation of the scalar trace objective and the warning in Step 7 of Algorithm 3 that trace equalization need not monotonically improve \(\mathsf{LIM}_\delta\).

None

Figure 1: Algebraic influence-density profiles. Algorithm 1 directly rescales the profile toward uniformity, Algorithm 2 changes it only weakly, and Algorithm 3 over-corrects the right boundary. The vertical scale is logarithmic..

None

Figure 2: Normalized trace-observability profile before and after the algebraic Algorithm 3 update. Scalar trace balance does not imply eigenvalue or condition-number balance..

None

Figure 3: Directly controlled surrogate imbalance divided by its initial value in the algebraic examples..

4.4 Probe bias and variance↩︎

To check the stochastic implementation in Algorithm 3, we estimate the baseline trace profile with common Gaussian probes and repeat the experiment 1000 times. For each probe count \(N_{\rm probe}\in\{1,2,4,8,16,32\}\), we evaluate the empirical squared penalty in 118 . The true unnormalized penalty is \(0.3438\). In the supplied reproducible run, the relative upward bias decreases from \(44.0\%\) with one probe to \(1.0\%\) with 32 probes, while the single-batch relative standard deviation decreases from \(216.3\%\) to \(24.4\%\). The plotted dispersion is not the standard error of the 1000-repetition mean; that standard error is the displayed standard deviation divided by \(\sqrt{1000}\). These results agree with ?? : increasing the probe count reduces both the covariance term and stochastic variability, at the cost of more Jacobian–vector products.

None

Figure 4: Monte Carlo bias and single-batch standard deviation of the empirical squared observability penalty, expressed as percentages of the true penalty. The plotted dispersion is not the standard error of the 1000-repetition mean..

4.5 Interpretation, reproducibility, and required Transformer validation↩︎

The controlled examples support only three implementation-level statements: the directly differentiated influence penalty changes the directly rescaled influence profile; the outer-loop weighting rule can stall when its adjustable component is weak; and scalar trace balancing need not improve influence or directional observability. They do not support claims about optimization through an actual Transformer, task preservation, or generalization.

A future empirical validation of the proposed training procedures should use at least a small causal Transformer on a synthetic long-context retrieval task and report: shared parameters rather than free positionwise gains; train and validation cross-entropy or retrieval accuracy; multiple random seeds with confidence intervals or standard errors; ablations over regularization strength, sequence length, depth, target density, and probe count; measured wall-clock and memory overhead; position-shift evaluation of both influence and retrieval performance; and direct empirical tests of the channel-energy sufficient conditions. Until such an experiment is supplied, Algorithms 1–3 are proposals and the present section remains an algebraic controlled example.

The companion script remedy_simulations.py constructs the stated matrices and regenerates the scalar tables and plots. Reproducibility claims are limited to those algebraic outputs.

5 Conclusion and limitations↩︎

The paper proves an analytic foundation for positional adjoint sensitivity in causal residual Transformers: well-posed masked depth dynamics, residual-to-flow and finite-token-to-Volterra approximation, readout-adjoint consistency under stated regularity, differentiability of the normalized influence density, its exact gradient-flow evolution, and an exact regional generator-term decomposition retaining covariance cross terms.

The paper does not prove that causal masking and residual connections universally generate Lost-in-the-Middle. Primacy requires cross-data and cross-depth non-cancellation; recency requires a sufficiently right-biased terminal readout or task protocol and a moment-controlled interaction correction; a double deficit follows from Theorem 16 only when measurable channel-energy and correlation bounds separate both boundaries from the middle. The regularizers are likewise proposals: their optimization response is not implied by the structural identities.

Open problems include deriving sharp correlation constants from architectural assumptions, replacing sufficient energy gaps by necessary-and-sufficient dynamical conditions, extending endpoint-aware finite-token estimates to weaker spatial and BV label regularity, obtaining moment bounds for heavy-tailed data, and testing whether the proposed diagnostics predict retrieval accuracy in production-scale long-context models. A further practical question is how to choose task-aligned observation maps that remain computationally tractable while preserving the causal interpretation of the readout.

6 Proof of the Residual-to-ODE Limit↩︎

Proof. Write \[G_\theta(t,X):=\zeta(t)\mathscr F_{\theta(t)}(X), \qquad e_k:=\|X_k^\varepsilon-X_{t_k}^\theta\|_{\mathcal{X}}.\] The limiting flow and the discrete layer satisfy \[\begin{align} X_{t_{k+1}}^\theta &=X_{t_k}^\theta+ \int_{t_k}^{t_{k+1}}G_\theta(r,X_r^\theta)\,\mathrm{d}r,\\ X_{k+1}^\varepsilon &=X_k^\varepsilon+ \varepsilon\zeta(t_k)\mathscr F_{\theta_k^\varepsilon} (X_k^\varepsilon)+d_k^\varepsilon. \end{align}\] Using the state Lipschitz estimate from Proposition 3, subtracting these identities gives \[\label{eq:convergent-control-error-recursion-pre} e_{k+1} \le(1+C\varepsilon)e_k +\|d_k^\varepsilon\|_{\mathcal{X}} +\tau_k,\tag{132}\] where \[\tau_k:=\int_{t_k}^{t_{k+1}} \left\| \zeta(t_k)\mathscr F_{\theta_k^\varepsilon}(X_{t_k}^\theta) -\zeta(r)\mathscr F_{\theta(r)}(X_r^\theta) \right\|_{\mathcal{X}}\,\mathrm{d}r.\] Insert successively the terms with \((\zeta(r),\theta_k^\varepsilon,X_{t_k}^\theta)\) and \((\zeta(r),\theta(r),X_{t_k}^\theta)\). Lipschitz continuity of \(\zeta\), local Lipschitz dependence on \(\theta\), and state Lipschitz continuity yield \[\begin{align} \tau_k &\le C\varepsilon^2 +C\int_{t_k}^{t_{k+1}} \|\theta_k^\varepsilon-\theta(r)\|\,\mathrm{d}r +C\int_{t_k}^{t_{k+1}} \|X_r^\theta-X_{t_k}^\theta\|_{\mathcal{X}}\,\mathrm{d}r\notag\\ &\le C\varepsilon^2 +C\int_{t_k}^{t_{k+1}} \|\theta^\varepsilon(r)-\theta(r)\|\,\mathrm{d}r. \label{eq:convergent-control-local-defect} \end{align}\tag{133}\] The last inequality uses boundedness of the vector field, which gives \(\|X_r^\theta-X_{t_k}^\theta\|\le C|r-t_k|\). Combining 132 , 133 , and \(\|d_k^\varepsilon\|\le C_{\rm split}\varepsilon^2\) gives \[e_{k+1} \le(1+C\varepsilon)e_k+C\varepsilon^2 +C\int_{t_k}^{t_{k+1}} \|\theta^\varepsilon(r)-\theta(r)\|\,\mathrm{d}r.\] Discrete Gronwall and summation over \(k\) imply \[\label{eq:convergent-control-grid-bound} \max_{0\le k\le M}e_k \le C_T\left( \varepsilon+\|\theta^\varepsilon-\theta\|_{L^1(0,T)} \right).\tag{134}\] For \(t\in[t_k,t_{k+1})\), boundedness of the limiting vector field yields \(\|X_t^\theta-X_{t_k}^\theta\|\le C\varepsilon\). Combining this with 134 proves ?? . The same argument with \(\theta\) replaced by \(\theta^\varepsilon\) removes the \(L^1\) term and gives the stated shadowing interpretation, but the reference flow then depends on \(\varepsilon\). ◻

References↩︎

[1]
Cheng-Yu Hsieh, Yung-Sung Chuang, Chun-Liang Li, Zifeng Wang, Long Le, Abhishek Kumar, James Glass, Alexander Ratner, Chen-Yu Lee, Ranjay Krishna, and Tomas Pfister. Found in the middle: Calibrating positional attention bias improves long context utilization. In Findings of the Association for Computational Linguistics: ACL 2024, pages 14982–14995, 2024. .
[2]
Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12:157–173, 2024. .
[3]
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. Axiomatic attribution for deep networks. In Proceedings of the 34th International Conference on Machine Learning, volume 70 of PMLR, pages 3319–3328, 2017.
[4]
Samira Abnar and Willem Zuidema. Quantifying attention flow in Transformers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4190–4197, 2020. .
[5]
Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, and David Duvenaud. Neural ordinary differential equations. In Advances in Neural Information Processing Systems, 2018. arXiv:1806.07366. .
[6]
Weinan E, Jiequn Han, and Qianxiao Li. A mean-field optimal control formulation of deep learning. Research in the Mathematical Sciences, 6:10, 2019. .
[7]
Cheng Huan and Hongwei Yuan. A first-order mean field control analysis of Transformer layers under cross-entropy training. arXiv:2606.23235, 2026. .
[8]
Shaojie Bai, Vladlen Koltun, and J. Zico Kolter. Stabilizing equilibrium models by Jacobian regularization. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of PMLR, pages 554–565, 2021.
[9]
Judy Hoffman, Daniel A. Roberts, and Sho Yaida. Robust learning with Jacobian regularization. arXiv:1908.02729, 2019. .
[10]
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. Efficient streaming language models with attention sinks. In International Conference on Learning Representations, 2024. arXiv:2309.17453. .
[11]
Borun D. Chowdhury. Lost in the middle at birth: An exact theory of Transformer position bias. arXiv:2603.10123, 2026. .
[12]
Edoardo Calvello, Nikola B. Kovachki, Matthew E. Levine, and Andrew M. Stuart. Continuum attention for neural operators. arXiv:2406.06486, 2024. .

  1. The subscript \(R\) indicates that the constant is uniform for states in the bounded set \(\|X\|_{C^1}\le R\) and for admissible parameters in \(\Theta_{\rm ad}\), while the subscript \(\eta\) records the lower bound \(p_i\ge\eta\) away from the left endpoint.↩︎

  2. That is, if \(I_l=((l-1)/L,l/L]\), define the cellwise density embedding by \[\label{eq:terminal-adjoint-density-embedding} \mathcal{E}_L\!\bigl(LP_{T,L}^{\theta,X_0,Y}\bigr)(p) :=\bigl[DH_L(X_L)^*\Delta_L\bigr]_l, \qquad p\in I_l;\qquad{(15)}\] suppose that, for every Lipschitz residual field \(\Delta\), \[\label{eq:readout-adjoint-consistency} \left\|\mathcal{E}_L\!\bigl(DH_L(X_L)^*\Delta_L\bigr)-DH(X)^*\Delta\right\|_{\mathcal{X}}\le \frac{C_R}{L}.\qquad{(16)}\] ↩︎

  3. Here \(\mathcal{D}\otimes\mathrm{Leb}\) is the product of the data law \(\mathcal{D}\) on examples \(\omega=(X_0,Y)\) and Lebesgue measure on the position interval \((0,1)\). Thus, for a jointly measurable random field \(Q(\omega,p)\), \(\|Q\|_{L^4(\mathcal{D}\otimes\mathrm{Leb})}^4 =\mathbb{E}_{\omega\sim\mathcal{D}}\int_0^1\|Q(\omega,p)\|^4\,\mathrm{d}p\).↩︎