June 30, 2026
Residual-stream analysis asks how a language model’s computation evolves across depth, but this question requires a stable semantic coordinate system. Intermediate decoding is meaningful only when readout coordinates remain comparable across layers. In language models, this comparability is mediated by the token interface: if embedding-side anchors and unembedding-side readout disagree on the semantic span under analysis, apparent hidden-state motion may reflect measurement drift rather than computation. We introduce Semantic Reference Frames (SemRF), an anchor-based formalism that separates semantic measurement from residual-stream dynamics. A SemRF fixes the anchors used in the analysis and measures residual states against them. Pseudo-inverse tying supplies an exact anchor-synchronization case; under restricted bi-invertibility on the chosen span, SemRF gives stable semantic-basis coordinates, distortion bounds, and controlled near-identity frame changes. With the frame fixed, residual computation becomes a depthwise semantic trajectory. The same anchors induce a semantic Voronoi diagram: semantic distance, equivalently semantic evidence such as logits, assigns each layer to a coarse cell, while coordinates retain within-cell motion and margins. We define layerwise steps, target-direction contribution profiles, and imbalance diagnostics, then use the observed Voronoi trace to define a margin-relaxed semantic tube. The canonical trace is the minimum-action path inside this tube. In a nonempty tube with at least one positive quadratic weight, it is unique and satisfies a discrete spline equation away from active constraints. Excess action controls step mismatch, curvature mismatch, interior deviation, and target-direction contribution-profile mismatch. Finally, low curvature energy implies piecewise-linear compressibility. This yields a local notion of knowledge density: within a fixed admissible frame and tolerance, lower trace complexity means that the depth-by-coordinate trajectory can be summarized by fewer semantic knots. Through the residual network’s parameter-to-trajectory map, this also gives a conditional link to parameter efficiency: among admissible parameter settings that fit the same data, lower-action and lower-complexity traces use fewer effective semantic degrees of freedom. The guarantees require controlled interface error, small projection residual, and explicit tube constraints.
semantic reference frames, semantic dynamics, language models, mechanistic interpretability, knowledge density
Transformer language models update a shared residual state through layerwise writes. This suggests a trajectory view: each layer moves the same state, and intermediate residual vectors carry partial semantic evidence. LogitLens and Tuned Lens decode hidden states into vocabulary space across depth [1], [2], but such trajectories are meaningful only when readout coordinates remain comparable across layers. If embedding-side anchors and unembedding-side readout disagree on the semantic span under analysis, apparent semantic motion may reflect measurement drift rather than computation.
We introduce Semantic Reference Frames (SemRF), an anchor-based formalism for separating measurement from dynamics. A SemRF fixes the semantic anchors used for the analysis and measures residual states against them. The main structural assumption is restricted bi-invertibility of the token interface on this anchor span. Pseudo-inverse tying supplies an exact anchor-synchronization case; when the restricted interface condition holds, semantic coordinates are stably measurable and interface mismatch yields explicit distortion bounds. SemRF specifies when semantic readout, diagnosis, and comparison are well posed.
Once measurement is controlled, residual evolution becomes a semantic trajectory. The same semantic anchors induce a semantic Voronoi diagram: each state is assigned to the anchor with smallest semantic distance, equivalently largest semantic evidence such as a logit. This is the sense in which we use “Voronoi” throughout; ordinary Euclidean nearest-neighbor Voronoi is only the special case induced by a Euclidean distance score. The diagram gives cells, faces, margins, and a coarse Voronoi trace; continuous coordinates record within-cell motion, cell crossings, and reversals along a fixed target direction. We combine the two views through a margin-relaxed semantic tube: the observed Voronoi trace defines a coarse route, and a path action selects the minimum-action path inside the tube as the canonical trace. The terminology is close to Semantic Tube Prediction [3], but here the tube is an analysis constraint induced by measured Voronoi cells rather than a training objective or geodesic prior. When tube constraints are active, the comparison is no longer just endpoint interpolation.
The scope of our research is local and diagnostic: after choosing semantic anchors, checking interface and projection errors, and fixing tube constraints, depthwise semantic computation becomes a measurement-and-trajectory problem. Within this setting, we define layerwise semantic steps, target-direction contribution profiles, and imbalance measures; prove stability under nearby admissible frames; and show that excess action controls deviations from the canonical trace. Low-curvature traces are piecewise-linearly compressible, yielding a local notion of knowledge density: for a fixed admissible frame and tolerance, lower trace complexity means that the depth trajectory can be summarized by a few semantic knots rather than by all layer-by-coordinate states. Parameter-efficiency claims are read through the induced trajectory map \(\theta\mapsto z_{0:L}(x;\theta)\), not through a direct identification between hidden-state locations and unique parameter values: after interface alignment stabilizes the semantic basis, parameter settings can be compared by the complexity of the semantic transport they realize on the same data. Our contributions are as follows:
Controlled semantic measurement. We formalize semantic reference frames as anchor-based coordinate systems for residual-stream analysis, identify restricted interface conditions under which semantic coordinates are stably measurable, and derive distortion bounds on the chosen anchor span.
Stable semantic dynamics and canonical traces. We define layerwise steps, target-direction contribution profiles, and imbalance diagnostics; prove cross-frame stability; and use an anchor-induced semantic Voronoi diagram to define a margin-relaxed semantic tube whose minimum-action path is the canonical trace.
Deviation control and local knowledge density. We show that excess action controls specified deviations from the canonical trace, and that low curvature energy yields piecewise-linear compressibility, giving a frame-relative interpretation of knowledge density and a conditional bridge to parameter efficiency.
We position SemRF relative to work on intermediate readout, LM semantics, residual-stream analysis, geometric control, and efficiency.
LogitLens and Tuned Lens recover token-level information from intermediate residual states [1], [2]. Probing, contextual-geometry, and latent-visualization studies examine how semantic content is organized in hidden representations [4]–[6]. SemRF isolates the measurement issue behind these methods: trajectory-level claims require coordinates that remain comparable across depth, not a readout recalibrated independently at each layer.
Recent work formulates LM semantics through vocabulary-defined semantics, semantic transition analysis, latent semantic alignment, and cross-model semantic transfer [7]–[9]. Pseudo-inverse tying connects semantic stability to synchronization between the input embedding and output unembedding on the relevant span [10]. SemRF builds on this line by making the semantic anchors explicit, using them to form anchor-induced Voronoi cells, and turning restricted bi-invertibility and admissibility into quantitative conditions for comparing anchor-restricted coordinates across depth.
Mechanistic interpretability views Transformer computation as successive writes to a shared residual stream [11]–[13]. This perspective supports circuit analyses, causal mechanism identification, sparse autoencoders, superposition models, and residual-stream factorization [14]–[20]. Related work also studies representation stabilization and information transport through depth [21]–[24]. SemRF asks when depth-indexed semantic motion reflects computation rather than readout drift.
Work on linear features, representation steering, internal-update equivalence, isotropy, and embedding geometry studies when representations admit stable geometric interpretation [6], [25]–[29]. Semantic Tube Prediction uses a geometric trajectory prior as a JEPA-style learning signal [3]. SemRF uses the related idea as an analysis device: its semantic tube is the margin-relaxed set of paths that follow the trusted Voronoi trace induced by semantic anchors in a validated frame.
Scaling laws and capability-density analyses study efficiency at the level of parameters, data, and compute [30]–[32]. SemRF uses a narrower notion: after admissible measurement is fixed, parameter settings that solve the same data can be compared by the action and complexity of the semantic traces they induce. This trace-level ranking is complementary to global efficiency metrics, not a replacement for them.
We begin with the measurement problem. Let \(h_{\ell,t}\in\mathbb{R}^d\) denote the residual-stream state at depth \(\ell\in\{0,\dots,L\}\) and position \(t\), and write \[h_{\ell+1,t}=h_{\ell,t}+u_{\ell,t}(h_{\ell,t}),\qquad \Delta h_{\ell,t}:=u_{\ell,t}(h_{\ell,t}).\] We analyze a fixed position \(t\) and suppress it. Depth \(\ell\) is discrete time. A semantic reference frame (SemRF) consists of semantic anchors \(\mathcal{B}\) and a coordinate map \(\phi:\mathbb{R}^d\to\mathbb{R}^K\). It induces coordinates \(z_\ell:=\phi(h_\ell)\) and increments \(\Delta z_\ell:=z_{\ell+1}-z_\ell\). This section asks when these coordinates are stable measurements of the chosen anchor span.
A SemRF starts by fixing the semantic anchors that define the analysis. Let \(\mathcal{B}=\{b_k\}_{k=1}^K\subset \mathbb{R}^d\), \(K\ge2\), be these anchors, and let \[z = \phi(h)\in\mathbb{R}^K, \label{eq:phi95def}\tag{1}\] be the corresponding semantic coordinates of a residual state \(h\in\mathbb{R}^d\).
Definition 1 (SemRF coordinates). A semantic coordinate system is a pair \((\phi,\mathcal{B})\) consisting of anchors \(\mathcal{B}\) and a coordinate map \(\phi:\mathbb{R}^d\to\mathbb{R}^K\). The coordinate \(z_i=\phi_i(h)\) is interpreted as the semantic evidence of state \(h\) for anchor \(b_i\).
Throughout the paper, \(z=\phi(h)\) denotes the coordinate vector in the semantic basis fixed by \(\mathcal{B}\). Margins, Voronoi assignments, steps, action values, and compressibility quantities are all computed in this single reference frame.
\(\langle \cdot,\cdot\rangle\) is the standard Euclidean inner product in the relevant ambient coordinate space, and \(\lVert \cdot\rVert\) denotes the induced \(\ell_2\) norm on vectors and the corresponding operator norm on matrices.
Vocabulary readout is the motivating example. Let \(W_{\mathrm{in}}\in\mathbb{R}^{V\times d}\) be the input embedding and \(W_{\mathrm{U}}\in\mathbb{R}^{V\times d}\) the unembedding. Writing \(E:=W_{\mathrm{in}}^\top\in\mathbb{R}^{d\times V}\), the \(i\)-th column \(e_i\) is the input-side anchor for token \(i\). Output-side anchors can be obtained from \(\tilde{B}:=W_{\mathrm{U}}^{+}\in\mathbb{R}^{d\times V}\). A hidden state is measured by its evidence for the chosen anchors, for example through logit coordinates or cosine similarity [7], [8].
The same semantic anchors in \(\mathcal{B}\) induce a semantic Voronoi diagram through semantic distances, equivalently through monotone semantic evidence scores. In the logit case, larger logit evidence means smaller semantic distance; for cell assignment one may set \(d_i(h):=-\phi_i(h)\). Thus the cell of anchor \(b_i\) is \[\mathcal{V}_i:=\{h:\phi_i(h)\ge \phi_j(h)\;\text{for all }j\}.\] The induced Voronoi index is \[\kappa(h):=\arg\max_i \phi_i(h),\] with ties broken by a fixed rule. A Voronoi face is where two anchors tie, \(\phi_i(h)=\phi_j(h)\). The margin \[\operatorname{mar}(h):=\phi_{(1)}(h)-\phi_{(2)}(h),\] where \(\phi_{(1)}(h)\) and \(\phi_{(2)}(h)\) are the largest and second-largest coordinates, is the semantic-distance gap to the nearest competing face. Thus \(\kappa(h)\) gives a coarse semantic state, while \(\phi(h)\) and its margin retain within-cell information. Throughout the paper, “Voronoi” refers to this semantic-distance/evidence diagram. It reduces to ordinary Euclidean Voronoi only for special distance-derived scores; the results below use only the coordinate dominance inequalities \(z_i\ge z_j\), which are convex halfspace constraints.
We now state when vocabulary readout becomes a controlled measurement procedure on the chosen anchor span. The first two propositions record exact and approximate synchronization cases; the assumption that follows is the local condition used by the later bounds.
Suppose there exist matrices \(Z\in\mathbb{R}^{V\times d}\) and an invertible \(T\in\mathbb{R}^{d\times d}\) such that (i) the embedding and unembedding are parameterized as \(W_{\mathrm{in}}=ZT^{-1}\in\mathbb{R}^{V\times d}\) and \(W_{\mathrm{U}}=ZT^\top\in\mathbb{R}^{V\times d}\), and (ii) \(Z\) has orthonormal columns: \(Z^\top Z=I_d\). Then the unembedding pseudoinverse recovers the input-side anchor matrix: \[W_{\mathrm{U}}^{+} = T^{-\top}Z^\top = W_{\mathrm{in}}^\top.\] Equivalently, the output-side pseudoinverse anchors coincide with the embedding anchors, so semantic coordinates defined via the restricted unembedding pseudoinverse are consistent across the input and output interfaces.
Let \(E_{\mathcal{B}}\in\mathbb{R}^{d\times K}\) have full column rank and write \(E_{\mathcal{B}}^+=(E_{\mathcal{B}}^\top E_{\mathcal{B}})^{-1}E_{\mathcal{B}}^\top\). Let \(\eta:=\|U_{\mathcal{B}}^\top-E_{\mathcal{B}}^+\|\) measure the discrepancy between a restricted readout and the anchor pseudoinverse. Then \[\|U_{\mathcal{B}}^\top E_{\mathcal{B}}-I_K\|\le \eta\,\|E_{\mathcal{B}}\|.\] Moreover, if \(U_{\mathcal{B}}^\top=\tilde{E}_{\mathcal{B}}^+\) for a nearby full-column-rank synthesis matrix \(\tilde{E}_{\mathcal{B}}\) with \(\|E_{\mathcal{B}}-\tilde{E}_{\mathcal{B}}\|\le \delta\), then \[\|U_{\mathcal{B}}^\top E_{\mathcal{B}}-I_K\|\le \|\tilde{E}_{\mathcal{B}}^+\|\,\delta.\]
Assumption 1 (Token-interface bi-invertibility). Let \(\mathcal{B}=\{b_k\}_{k=1}^K\) be the anchor set used by \(\phi\). There exist an anchor synthesis matrix \(E_{\mathcal{B}}\in\mathbb{R}^{d\times K}\) and a restricted readout matrix \(U_{\mathcal{B}}\in\mathbb{R}^{d\times K}\) such that \[\lVert U_{\mathcal{B}}^\top E_{\mathcal{B}} - I_K\rVert \le \varepsilon,\] with \(E_{\mathcal{B}}\) full column rank and \(\|E_{\mathcal{B}}^+\|\le C_E\), \(\|U_{\mathcal{B}}\|\le C_U\) for fixed frame constants \(C_E,C_U\).
This local assumption rules out spurious coordinate drift created by the token interface. It requires stable synthesis and readout on the chosen anchor span, not a global semantic basis.
Definition 2 (Admissible SemRF regime). Fix tolerances \(\eta_{\mathrm{int}},\eta_{\mathrm{proj}}>0\). We call a prediction event \((\theta,x,y)\) admissible* for a frame \((\phi,\mathcal{B})\) when the interface mismatch satisfies \[\lVert U_{\mathcal{B}}^\top E_{\mathcal{B}}-I_K\rVert\le \eta_{\mathrm{int}},\] and the relative projection residual of each analyzed state satisfies \[\frac{\|h_\ell-E_{\mathcal{B}}E_{\mathcal{B}}^{+}h_\ell\|}{\|h_\ell\|}\le \eta_{\mathrm{proj}},\qquad \ell=0,\dots,L.\] The same requirement may also be imposed on residual writes \(\Delta h_\ell\) when step-level conclusions are needed.*
All quantitative results below are conditional on admissibility. Large interface mismatch means unstable readout; large projection residual means the chosen anchors miss active computation. Target-direction statements additionally require a fixed unit direction, either the target-token coordinate axis or a unit contrast direction.
The same diagnostics control how far measured coordinates can deviate from intended anchor-span coordinates.
The key observables are the interface mismatch \(\lVert U_{\mathcal{B}}^\top E_{\mathcal{B}}-I_K\rVert\) and the relative projection residual \[r(h):=h-E_{\mathcal{B}}E_{\mathcal{B}}^+h, \qquad \frac{\|r(h)\|}{\|h\|}.\] The first tests stability of synthesis and readout on the anchor span; the second tests whether the analyzed state lies mostly inside that span. Anchor choice therefore determines the semantic alternatives, interface conditioning, and reliability of later trajectory bounds.
Assume 1 and use the restricted linear readout coordinates \(\phi_{\mathrm{rd}}(h)=U_{\mathcal{B}}^\top h\). For any coordinate vector \(c\in\mathbb{R}^K\), let \(h_c:=E_{\mathcal{B}}c\) be its recomposition into the anchor subspace. Then \[\lVert \phi_{\mathrm{rd}}(h_c)-c\rVert=\lVert (U_{\mathcal{B}}^\top E_{\mathcal{B}}-I_K)c\rVert\le \varepsilon\,\lVert c\rVert.\] More generally, for any decomposition \(h=E_{\mathcal{B}}c+r\), \[\lVert \phi_{\mathrm{rd}}(h)-c\rVert\le \varepsilon\,\lVert c\rVert+\lVert U_{\mathcal{B}}\rVert\,\lVert r\rVert.\]
Corollary 1 (Trajectory distortion under bounded drift). Let \(\{h_\ell\}_{\ell=0}^L\) be any residual trajectory and suppose each state decomposes as \(h_\ell=E_{\mathcal{B}}c_\ell+r_\ell\). Let \(z_\ell=\phi_{\mathrm{rd}}(h_\ell)\) and \(\Delta c_\ell:=c_{\ell+1}-c_\ell\). Then \[\lVert z_\ell-c_\ell\rVert\le \varepsilon\lVert c_\ell\rVert+\lVert U_{\mathcal{B}}\rVert\lVert r_\ell\rVert, \qquad \lVert \Delta z_\ell-\Delta c_\ell\rVert\le \varepsilon(\lVert c_{\ell+1}\rVert+\lVert c_\ell\rVert)+\lVert U_{\mathcal{B}}\rVert(\lVert r_{\ell+1}\rVert+\lVert r_\ell\rVert).\]
These bounds identify the regime in which measured semantic motion is attributable to the residual trajectory rather than to interface error or off-span residuals. Section 4 then studies which dynamical features remain stable under admissible changes of frame.
With a checked SemRF, residual-stream computation induces a depthwise trajectory in the semantic-basis coordinates fixed in Section 3. For a prediction event \((\theta,x,y)\), write \(z_{0:L}\) for this trajectory. This section defines step observables, proves their stability under admissible frame changes, and introduces the tube-constrained action used in Section 5.
The semantic anchors induce Voronoi cells, so each layer has a coarse state \(\kappa_\ell:=\kappa(h_\ell)\) and margin \(\operatorname{mar}(h_\ell)\) to the nearest Voronoi face. The coordinates \(z_\ell\) retain the within-cell motion and the evidence changes that drive cell crossings. For a trajectory \(\{h_\ell\}_{\ell=0}^L\), define the induced semantic field at layer \(\ell\) by \[F_\ell(h):=\phi\!\bigl(h+u_\ell(h)\bigr)-\phi(h)\in\mathbb{R}^K. \label{eq:field95def}\tag{2}\] Along the realized trajectory, \(F_\ell(h_\ell)=\Delta z_\ell\).
Assumption 2 (Local linearization). Along trajectories of interest, \(\phi\) admits a first-order approximation with bounded Jacobian \(J_\phi\): \(\phi(h+\delta)=\phi(h)+J_\phi(h)\delta + r(h,\delta)\) with \(\lVert r(h,\delta)\rVert\le C_r\lVert \delta\rVert^2\).
Under 2, the induced semantic step satisfies \[\Delta z_\ell = J_\phi(h_\ell)\,\Delta h_\ell + \varepsilon_\ell, \qquad \lVert \varepsilon_\ell\rVert\le C_r \lVert \Delta h_\ell\rVert^2.\]
Thus each semantic step is the first-order image of the residual write, up to a controlled quadratic remainder.
For a fixed context \(x\), let \(e_y\) be a fixed unit direction in the chosen coordinates, either the target-token coordinate axis when included in the anchor set or a unit contrast direction. Define the signed target-direction contribution \[a_\ell := \langle \Delta z_\ell,e_y\rangle.\] The contribution profile \(\{a_\ell\}\) sums to the net target-direction displacement, \[\sum_{\ell=0}^{L-1} a_\ell = \langle z_L-z_0,e_y\rangle,\] and records where target-direction displacement is accumulated, delayed, or reversed.
We use semantic imbalance for trajectories in which later layers partially undo earlier progress before reaching the same endpoint. Three diagnostics make this precise: \[\text{motion energy } \sum_{\ell=0}^{L-1}\lVert \Delta z_\ell\rVert^2, \qquad \text{curvature energy } \sum_{\ell=1}^{L-1}\lVert \Delta^2 z_\ell\rVert^2,\] where \(\Delta^2 z_\ell:=z_{\ell+1}-2z_\ell+z_{\ell-1}\) is the discrete second difference, together with the cumulative target-direction backtracking penalty \[\sum_{\ell=0}^{L-1}\bigl[-\langle \Delta z_\ell,e_y\rangle\bigr]_+.\] The first tracks total semantic travel, the second tracks oscillation across depth, and the third isolates target-direction reversal. They flag detours, low-margin crossings, and oscillation between alternatives, without by themselves proving that a crossing is unnecessary.
We next show that admissible frame changes preserve the coordinate geometry and step structure used below.
Let \(E_{\mathcal{B}}\in\mathbb{R}^{d\times K}\) have full column rank and consider two restricted coordinate maps \(\phi(h)=U_{\mathcal{B}}^\top h\) and \(\tilde{\phi}(h)=\tilde{U}_{\mathcal{B}}^\top h\). Assume both satisfy the restricted interface condition with errors \(\varepsilon,\tilde{\varepsilon}<1\): \[\lVert U_{\mathcal{B}}^\top E_{\mathcal{B}}-I_K\rVert\le \varepsilon, \qquad \lVert \tilde{U}_{\mathcal{B}}^\top E_{\mathcal{B}}-I_K\rVert\le \tilde{\varepsilon}.\] Then on the anchor subspace \(\mathrm{span}(E_{\mathcal{B}})\) the two coordinate systems are related by a near-identity linear transform. For any \(h=E_{\mathcal{B}}c\), \[\tilde{\phi}(h)=\bigl(I_K + \Delta\bigr)\,\phi(h), \qquad \lVert \Delta\rVert\le\frac{\varepsilon+\tilde{\varepsilon}}{1-\varepsilon}.\]
Corollary 2 (Relational stability). Under the hypotheses of Proposition [prop:coord95equivalence], let \[\eta:=\frac{\varepsilon+\tilde{\varepsilon}}{1-\varepsilon}.\] For any \(u,v\in \mathrm{span}(E_{\mathcal{B}})\), \[\bigl|\langle \tilde{\phi}(u),\tilde{\phi}(v)\rangle-\langle \phi(u),\phi(v)\rangle\bigr| \le (2\eta+\eta^2)\,\lVert \phi(u)\rVert\,\lVert \phi(v)\rVert,\] and \[\lVert \tilde{\phi}(u)-\tilde{\phi}(v)\rVert \le (1+\eta)\,\lVert \phi(u)-\phi(v)\rVert.\] If \(\eta<1\), the matching lower bound also holds: \[(1-\eta)\,\lVert \phi(u)-\phi(v)\rVert \le \lVert \tilde{\phi}(u)-\tilde{\phi}(v)\rVert.\] Thus pairwise similarities and distances are stable when the interface errors are small.
Assumption 3 (Lipschitz stability). The coordinate map \(\phi\) is \(L_\phi\)-Lipschitz. For all \(h,h'\in\mathbb{R}^d\), \[\lVert \phi(h)-\phi(h')\rVert \le L_\phi \lVert h-h'\rVert.\]
Theorem 1 (Semantic-step stability bound). Under 3, for any residual update \(\Delta h_\ell\), \[\lVert \Delta z_\ell\rVert=\lVert \phi(h_\ell+\Delta h_\ell)-\phi(h_\ell)\rVert\le L_\phi \lVert \Delta h_\ell\rVert.\] In particular, if \(\lVert \Delta h_\ell\rVert\le U\) for all \(\ell\), then \(\lVert \Delta z_\ell\rVert\le L_\phi U\) for all \(\ell\).
Theorem 2 (Frame-to-step stability transfer). Consider two linear-readout coordinate maps \(\phi(h)=U_{\mathcal{B}}^\top h\) and \(\tilde{\phi}(h)=\tilde{U}_{\mathcal{B}}^\top h\) built on the same anchor synthesis \(E_{\mathcal{B}}\in\mathbb{R}^{d\times K}\), and assume both satisfy the restricted bi-invertibility condition (1) with errors \(\varepsilon,\tilde{\varepsilon}<1\). Define \[C_{\mathrm{frame}} := \frac{\varepsilon+\tilde{\varepsilon}}{\sigma_{\min}(E_{\mathcal{B}})}.\] Then on the anchor subspace \(\mathrm{span}(E_{\mathcal{B}})\) the two coordinate systems are uniformly close, and for any \(h=E_{\mathcal{B}}c\), \[\lVert \phi(h)-\tilde{\phi}(h)\rVert \le C_{\mathrm{frame}}\lVert h\rVert.\] Moreover, for any anchor-subspace update \(\Delta h=E_{\mathcal{B}}\Delta c\), \[\lVert (\phi(h+\Delta h)-\phi(h))-(\tilde{\phi}(h+\Delta h)-\tilde{\phi}(h))\rVert \le C_{\mathrm{frame}}\lVert \Delta h\rVert.\]
Together, these results show that, under admissibility, nearby frames preserve the anchor-span geometry and induced step profiles used below. Continuous semantic steps can then be compared across depth without being dominated by readout drift. Voronoi assignments are reliable when their margins dominate frame distortion; for example, if \(\|\tilde{\phi}(h)-\phi(h)\|_\infty\le \delta\) and \(\operatorname{mar}(h)>2\delta\), then \(\tilde{\kappa}(h)=\kappa(h)\).
We now introduce the comparison path. Endpoint matching makes pathwise deviation comparable. The semantic tube keeps the path on the trusted coarse Voronoi route up to margin relaxations, without copying the observed coordinates exactly.
Let \(z^{\mathrm{obs}}_{0:L}\) be an observed teacher-forced trajectory and let \(z^{\mathrm{cmp}}_{0:L}\) be any endpoint-matched comparison path. Assume \[z^{\mathrm{obs}}_0=z^{\mathrm{cmp}}_0, \qquad z^{\mathrm{obs}}_L=z^{\mathrm{cmp}}_L.\] Define \[g_\ell:=\langle \Delta z^{\mathrm{obs}}_\ell-\Delta z^{\mathrm{cmp}}_\ell,e_y\rangle.\] Then \[\sum_{\ell=0}^{L-1} g_\ell = 0, \qquad \sum_{\ell=0}^{j-1} g_\ell = \langle z^{\mathrm{obs}}_j-z^{\mathrm{cmp}}_j,e_y\rangle,\quad j=1,\dots,L.\]
Thus cumulative contribution-profile mismatch equals intermediate target-direction displacement between the observed path and any endpoint-matched comparison path.
Definition 3 (Margin-relaxed semantic tube). Let \(z^{\mathrm{obs}}_{0:L}\) be the observed trajectory and let \(\tau_\ell:=\kappa(h^{\mathrm{obs}}_\ell)\) be its Voronoi index. Choose a trusted layer set \(\mathcal{I}\subseteq\{1,\dots,L-1\}\), typically the layers whose margins dominate measurement error, and relaxation radii \(\rho_\ell\ge 0\). The semantic tube induced by the observed Voronoi trace is \[\mathcal{T}_\rho(z^{\mathrm{obs}}) :=\left\{z_{0:L}:\begin{array}{l} z_0=z^{\mathrm{obs}}_0,\quad z_L=z^{\mathrm{obs}}_L,\\ z_{\ell,\tau_\ell}\ge z_{\ell,j}-\rho_\ell\quad \forall \ell\in\mathcal{I},\;\forall j \end{array}\right\}.\] The constraints are linear dominance constraints in the semantic coordinates, so the tube is a closed convex polyhedron. For a realized trajectory and nonnegative radii, it contains \(z^{\mathrm{obs}}\). If \(\mathcal{I}=\varnothing\), the construction reduces to endpoint matching; inactive constraints leave the endpoint-only baseline unchanged.
Definition 4 (Path action). For a prediction event \((\theta,x,y)\), let \(e_y\in\mathbb{R}^K\) be the fixed unit direction used for target-direction diagnostics. We define an action \(S[z]\) on depth-discrete trajectories \(z_{0:L}\) by \[\label{eq:action} S[z] := \alpha \sum_{\ell=0}^{L-1}\lVert \Delta z_\ell\rVert^2 \;+\; \beta \sum_{\ell=1}^{L-1}\lVert \Delta^2 z_\ell\rVert^2 \;+\; \gamma \sum_{\ell=0}^{L-1}\bigl[-\langle \Delta z_\ell,e_y\rangle\bigr]_+,\qquad{(1)}\] with \(\alpha,\beta\ge 0\) not both zero, and \(\gamma\ge 0\). The terms penalize excess motion, oscillation, and target-direction backtracking.
If no target-direction diagnostic is used, set \(\gamma=0\) and omit the backtracking term. The action scores feasible comparison paths rather than training the model; Section 5 minimizes it inside the semantic tube.
We now pass from descriptive path comparison to a selected baseline. Minimizing the path action from Section 4 over the margin-relaxed semantic tube produces the canonical trace. Unless stated otherwise, Euler–Lagrange statements concern the quadratic regime \(\gamma=0\) on layers where tube inequalities are inactive.
Under teacher forcing, the model induces an observed semantic trajectory \(z^{\mathrm{TF}}_{0:L}\) with endpoints \(z^{\mathrm{TF}}_0=z_0\) and \(z^{\mathrm{TF}}_L=z_L\). With the semantic tube \(\mathcal{T}_\rho(z^{\mathrm{TF}})\) from Definition 3, define \[\label{eq:mintrace} z^\star \in \arg\min_{z_{0:L}\in\mathcal{T}_\rho(z^{\mathrm{TF}})} S[z].\tag{3}\] We call \(z^\star\) the canonical trace: the diagnostic minimum-action comparison path that follows the trusted coarse Voronoi route up to the relaxation radii.
Theorem 3 (Canonical-trace characterization). Consider 3 with the action ?? . The semantic tube from Definition 3 is a nonempty closed convex polyhedron for realized trajectories and nonnegative radii. If \((\alpha,\beta)\neq(0,0)\) with \(\alpha,\beta\ge 0\), the quadratic part is strictly convex on endpoint-constrained path variables. Hence the canonical trace is unique for \(\gamma=0\) and remains unique after adding the convex backtracking term for \(\gamma>0\). For \(\gamma=0\), writing \[\Delta^2 z_\ell := z_{\ell+1}-2z_\ell+z_{\ell-1}, \qquad \Delta^4 z_\ell := z_{\ell+2}-4z_{\ell+1}+6z_\ell-4z_{\ell-1}+z_{\ell-2},\] the coordinates satisfy, on every inactive interior layer, \[\beta\,\Delta^4 z^\star_\ell = \alpha\,\Delta^2 z^\star_\ell, \qquad \ell=2,\dots,L-2.\] At active tube constraints, KKT multipliers for the corresponding linear dominance inequalities are added.
Here \(\alpha\) prices total semantic travel and \(\beta\) penalizes oscillatory correction. If the tube has no active constraints, the quadratic canonical trace is straight endpoint interpolation. Otherwise \(z^\star\) is obtained from a sparse convex quadratic program with a banded Hessian and layer-local linear inequalities.
Excess action is a scalar deviation diagnostic: the action gap dominates quadratic deviation energy, with strong-convexity modulus stated in Proposition [prop:strongconv95appendix].
Let \(z^\star\) be a minimizer of 3 for the action ?? with \(\gamma\ge 0\), and let \(z\in\mathcal{T}_\rho(z^{\mathrm{TF}})\) be any feasible trajectory. Writing \(e_\ell:=z_\ell-z_\ell^\star\), one has \[S[z]-S[z^\star] \ge \alpha\sum_{\ell=0}^{L-1}\lVert \Delta e_\ell\rVert^2 + \beta\sum_{\ell=1}^{L-1}\lVert \Delta^2 e_\ell\rVert^2.\] When \(\gamma=0\) and the active normal-cone pairing vanishes in the direction \(z-z^\star\), the inequality is an equality.
Corollary 3 (Deviation controls from excess action). In the setting of Proposition [prop:exact95gap], define \(\Delta S:=S[z]-S[z^\star]\). Then \[\alpha\sum_{\ell=0}^{L-1}\lVert \Delta z_\ell-\Delta z_\ell^\star\rVert^2 \le \Delta S, \qquad \beta\sum_{\ell=1}^{L-1}\lVert \Delta^2 z_\ell-\Delta^2 z_\ell^\star\rVert^2 \le \Delta S.\] If the endpoint-constrained problem has strong-convexity modulus \(\mu>0\), then \[\sum_{\ell=1}^{L-1}\lVert z_\ell-z_\ell^\star\rVert^2\le \frac{2\Delta S}{\mu}.\] For \(z=z^{\mathrm{obs}}\), define \(g_\ell:=\langle \Delta z^{\mathrm{obs}}_\ell-\Delta z^\star_\ell,e_y\rangle\). If \(\alpha>0\), then \[\max_{1\le j\le L}\left|\sum_{\ell=0}^{j-1} g_\ell\right| \le \sum_{\ell=0}^{L-1}|g_\ell| \le \sqrt{\frac{L}{\alpha}}\,\Delta S^{1/2}.\] Thus excess action controls step mismatch, curvature mismatch, interior displacement, and cumulative target-direction contribution-profile mismatch whenever the corresponding weights are positive.
The preceding bounds quantify deviation from the canonical trace. We now ask when such a trace admits a compact description: low curvature means that few semantic knots suffice to approximate the full depthwise trajectory.
Definition 5 (Trajectory compressibility). A semantic trajectory \(z_{0:L}\in(\mathbb{R}^K)^{L+1}\) is \((m,\varepsilon)\)-compressible* if there exists a piecewise-linear trajectory \(\tilde{z}_{0:L}\) with at most \(m\) linear segments, whose breakpoints are the semantic knots, such that \[\frac{1}{L+1}\sum_{\ell=0}^L \lVert z_\ell-\tilde{z}_\ell\rVert^2 \le \varepsilon^2.\]*
Definition 6 (Frame-relative trace complexity). For a fixed admissible SemRF and tolerance \(\varepsilon\), define \(m_\varepsilon(z)\) as the smallest \(m\) for which \(z\) is \((m,\varepsilon)\)-compressible. Up to knot locations, the trace needs \(O(K\,m_\varepsilon(z))\) real values rather than \((L+1)K\). We use local knowledge density* only in this frame-relative sense: within the same frame and tolerance, smaller trace complexity means a more compact description of the same depthwise semantic transport.*
For a fixed architecture and dataset \(\mathcal{D}\), parameters \(\theta\) induce teacher-forced residual trajectories and hence SemRF trajectories \(z_{0:L}(x;\theta)\). Interface alignment, including the pseudo-inverse tying case in Section 3, stabilizes the semantic basis so that parameter settings can be compared through their induced semantic traces rather than through moving readout frames. Define the average trace complexity \[\bar m_\varepsilon(\theta;\mathcal{D}):=\frac{1}{|\mathcal{D}|}\sum_{x\in\mathcal{D}} m_\varepsilon\!\left(z^\star(x;\theta)\right),\] where \(z^\star(x;\theta)\) is the canonical trace inside the tube induced by the trajectory of \(\theta\) on \(x\). For a fixed frame dimension, a conditional semantic-density proxy is \[\mathrm{KD}_{\varepsilon}(\theta;\mathcal{D}):=\frac{L+1}{\bar m_\varepsilon(\theta;\mathcal{D})+1}.\] Equivalently, over a feasible family \(\Theta_{\mathcal{D}}\) of admissible parameters satisfying the same data constraints, one may use \[\theta^\dagger\in\arg\min_{\theta\in\Theta_{\mathcal{D}}} \frac{1}{|\mathcal{D}|}\sum_{x\in\mathcal{D}} S\!\left[z^\star(x;\theta)\right] +\lambda\,\bar m_\varepsilon(\theta;\mathcal{D})\] with \(\lambda\ge0\) as a semantic-efficiency selector. Lower action and lower trace complexity then indicate higher local knowledge density and a more parameter-efficient semantic realization, relative to the parameter-to-trajectory map \(\theta\mapsto z_{0:L}(x;\theta)\) and the fixed frame, data, and tolerance.
Theorem 4 (Curvature-to-compressibility bound). Let \(z_{0:L}\) be any trajectory with curvature energy \[E_2(z):=\sum_{\ell=1}^{L-1}\lVert \Delta^2 z_\ell\rVert^2.\] There exists a piecewise-linear \(\tilde{z}_{0:L}\) with at most \(m\) segments such that \[\frac{1}{L+1}\sum_{\ell=0}^L \lVert z_\ell-\tilde{z}_\ell\rVert^2 \le \frac{c\,L^3\,E_2(z)}{m^4},\] where \(c>0\) is a universal constant independent of \(K\) and \(L\). Hence \(z\) is \((m,\varepsilon)\)-compressible whenever \[m \ge \left(\frac{c\,L^3\,E_2(z)}{\varepsilon^2}\right)^{1/4}.\]
For a quadratic canonical trace with \(\gamma=0\) and \(\beta>0\), one has \(E_2(z^\star)\le S[z^\star]/\beta\). Theorem 4 therefore gives \((m,\varepsilon)\)-compressibility whenever \[m\ge \left(\frac{c\,L^3\,S[z^\star]}{\beta\,\varepsilon^2}\right)^{1/4}.\] This justifies the local knowledge-density claim: once measurement, tolerance, and tube constraints are fixed, low-curvature admissible traces need fewer values than the full depth-by-coordinate representation. The Voronoi trace gives the coarse route, while the canonical trace and excess action quantify the remaining continuous motion.
SemRF isolates a local regime in which residual-stream semantics is well posed. On a validated anchor span, coordinates are stable, the same anchors induce a semantic Voronoi diagram, and admissible frame changes preserve the geometry needed for path comparison. The observed Voronoi sequence defines a margin-relaxed semantic tube; minimizing path action inside it gives a diagnostic canonical trace. With at least one positive quadratic weight, the trace is unique, satisfies a discrete spline equation away from active constraints, and reduces to straight interpolation only when no coarse-state constraint is active. Excess action controls the corresponding step, curvature, interior, and target-direction deviations. Low curvature yields piecewise-linear compressibility, giving a local, frame- and tolerance-relative notion of knowledge density. Through the parameter-induced trajectory map, this connects to parameter efficiency by ranking lower-action and lower-complexity semantic transport within the same admissible frame and data constraints. These are local guarantees: they require controlled interface error, small projection residual, explicit tube constraints, and a fixed unit direction whenever target-direction diagnostics are used.
The additional proofs to our research are as follows, grouped by the corresponding technical role.
Proof. of Proposition [prop:pitz95canonical]. Because \(Z^\top Z=I_d\) and \(T\) is invertible, \[W_{\mathrm U}^{+}=(ZT^\top)^+=T^{-\top}Z^\top.\] Using \(W_{\mathrm{in}}=ZT^{-1}\) gives \(W_{\mathrm{in}}^\top=T^{-\top}Z^\top\), so \(W_{\mathrm U}^{+}=W_{\mathrm{in}}^\top\). ◻
Proof. of Proposition [prop:approx95pit]. Let \(E_{\mathcal{B}}^+=(E_{\mathcal{B}}^\top E_{\mathcal{B}})^{-1}E_{\mathcal{B}}^\top\). Since \(E_{\mathcal{B}}\) has full column rank, \(E_{\mathcal{B}}^+E_{\mathcal{B}}=I_K\). Hence \[U_{\mathcal{B}}^\top E_{\mathcal{B}}-I_K=(U_{\mathcal{B}}^\top-E_{\mathcal{B}}^+)E_{\mathcal{B}},\] so if \(\eta:=\|U_{\mathcal{B}}^\top-E_{\mathcal{B}}^+\|\), then \[\|U_{\mathcal{B}}^\top E_{\mathcal{B}}-I_K\|\le \eta\,\|E_{\mathcal{B}}\|.\] If instead \(U_{\mathcal{B}}^\top=\tilde{E}_{\mathcal{B}}^+\) for a nearby full-column-rank synthesis matrix \(\tilde{E}_{\mathcal{B}}\), then \(\tilde{E}_{\mathcal{B}}^+\tilde{E}_{\mathcal{B}}=I_K\) and therefore \[U_{\mathcal{B}}^\top E_{\mathcal{B}}-I_K =\tilde{E}_{\mathcal{B}}^+(E_{\mathcal{B}}-\tilde{E}_{\mathcal{B}}),\] which gives \[\|U_{\mathcal{B}}^\top E_{\mathcal{B}}-I_K\|\le \|\tilde{E}_{\mathcal{B}}^+\|\,\delta.\] ◻
Proof. of Proposition [prop:coord95distortion]. For \(h_c=E_{\mathcal{B}}c\), \[\phi_{\mathrm{rd}}(h_c)-c=(U_{\mathcal{B}}^\top E_{\mathcal{B}}-I)c,\] so the first claim follows from 1. For the general decomposition \(h=E_{\mathcal{B}}c+r\), \[\phi_{\mathrm{rd}}(h)-c=(U_{\mathcal{B}}^\top E_{\mathcal{B}}-I)c+U_{\mathcal{B}}^\top r,\] and the stated bound follows by the triangle inequality. ◻
Proof. of Corollary 1. The state bound is the second bound of Proposition [prop:coord95distortion], applied at layer \(\ell\) with \(h_\ell=E_{\mathcal{B}}c_\ell+r_\ell\). For the step bound, write \[\Delta z_\ell-\Delta c_\ell=(z_{\ell+1}-c_{\ell+1})-(z_\ell-c_\ell),\] and apply the state bound at layers \(\ell+1\) and \(\ell\). ◻
Proof. of Proposition [prop:first95order]. Apply 2 with \(h=h_\ell\) and \(\delta=\Delta h_\ell\): \[\phi(h_\ell+\Delta h_\ell)-\phi(h_\ell)=J_\phi(h_\ell)\Delta h_\ell+r(h_\ell,\Delta h_\ell),\] and rename the remainder \(\varepsilon_\ell\). ◻
Proof. of Proposition [prop:coord95equivalence]. Let \(A:=U_{\mathcal{B}}^\top E_{\mathcal{B}}\) and \(\tilde{A}:=\tilde{U}_{\mathcal{B}}^\top E_{\mathcal{B}}\). Since \(\|A-I\|\le \varepsilon<1\), the Neumann-series bound gives \(A\) invertible with \(\|A^{-1}\|\le 1/(1-\varepsilon)\). For \(h=E_{\mathcal{B}}c\), \[\phi(h)=Ac,\qquad \tilde{\phi}(h)=\tilde{A} c=\tilde{A} A^{-1}\phi(h).\] Hence \(\tilde{\phi}=(I+\Delta)\phi\) on the anchor span with \(\Delta:=\tilde{A} A^{-1}-I=(\tilde{A}-A)A^{-1}\). Using \[\|\tilde{A}-A\|\le \|\tilde{A}-I\|+\|A-I\|\le \tilde{\varepsilon}+\varepsilon,\] one obtains \[\|\Delta\|\le \frac{\varepsilon+\tilde{\varepsilon}}{1-\varepsilon}.\] ◻
Proof. of Corollary 2. By Proposition [prop:coord95equivalence], on the anchor span one has \(\tilde{\phi}=(I+\Delta)\phi\) with \[\|\Delta\|\le \eta:=\frac{\varepsilon+\tilde{\varepsilon}}{1-\varepsilon}.\] For any \(u,v\in \mathrm{span}(E_{\mathcal{B}})\), \[\langle \tilde{\phi}(u),\tilde{\phi}(v)\rangle-\langle \phi(u),\phi(v)\rangle =\langle \Delta\phi(u),\phi(v)\rangle+\langle \phi(u),\Delta\phi(v)\rangle+\langle \Delta\phi(u),\Delta\phi(v)\rangle.\] Applying the Cauchy–Schwarz inequality and the operator-norm bound on \(\Delta\) gives \[\bigl|\langle \tilde{\phi}(u),\tilde{\phi}(v)\rangle-\langle \phi(u),\phi(v)\rangle\bigr| \le (2\eta+\eta^2)\,\lVert \phi(u)\rVert\,\lVert \phi(v)\rVert.\] For distances, set \(w:=\phi(u)-\phi(v)\). Then \[\tilde{\phi}(u)-\tilde{\phi}(v)=(I+\Delta)w,\] so \[\lVert (I+\Delta)w\rVert\le (1+\eta)\lVert w\rVert.\] If \(\eta<1\), the reverse triangle inequality also gives \[\lVert (I+\Delta)w\rVert\ge \lVert w\rVert-\lVert \Delta w\rVert\ge (1-\eta)\lVert w\rVert.\] This proves the stated distance bounds. ◻
Proof. of Theorem 1. The statement is immediate from 3: \[\lVert \Delta z_\ell\rVert=\lVert \phi(h_\ell+\Delta h_\ell)-\phi(h_\ell)\rVert\le L_\phi\lVert \Delta h_\ell\rVert.\] The uniform bound follows by substituting \(\lVert \Delta h_\ell\rVert\le U\). ◻
Proof. of Theorem 2. For \(h=E_{\mathcal{B}}c\), \[\phi(h)-\tilde{\phi}(h)=(U_{\mathcal{B}}^\top-\tilde{U}_{\mathcal{B}}^\top)E_{\mathcal{B}}c=((U_{\mathcal{B}}^\top E_{\mathcal{B}}-I)- (\tilde{U}_{\mathcal{B}}^\top E_{\mathcal{B}}-I))c.\] Hence \[\lVert \phi(h)-\tilde{\phi}(h)\rVert\le (\varepsilon+\tilde{\varepsilon})\lVert c\rVert.\] Since \(\|h\|=\|E_{\mathcal{B}}c\|\ge \sigma_{\min}(E_{\mathcal{B}})\|c\|\), it follows that \[\lVert \phi(h)-\tilde{\phi}(h)\rVert\le \frac{\varepsilon+\tilde{\varepsilon}}{\sigma_{\min}(E_{\mathcal{B}})}\lVert h\rVert=C_{\mathrm{frame}}\lVert h\rVert.\] For the step bound, use linearity: \[(\phi(h+\Delta h)-\phi(h))-(\tilde{\phi}(h+\Delta h)-\tilde{\phi}(h))=(U_{\mathcal{B}}^\top-\tilde{U}_{\mathcal{B}}^\top)\Delta h.\] Writing \(\Delta h=E_{\mathcal{B}}\Delta c\) and repeating the previous argument yields \[\lVert (\phi(h+\Delta h)-\phi(h))-(\tilde{\phi}(h+\Delta h)-\tilde{\phi}(h))\rVert \le \frac{\varepsilon+\tilde{\varepsilon}}{\sigma_{\min}(E_{\mathcal{B}})}\lVert \Delta h\rVert=C_{\mathrm{frame}}\lVert \Delta h\rVert.\] ◻
Proof. of Proposition [prop:contrib95tel]. Let \(z^{\mathrm{cmp}}\) be any endpoint-matched comparison path, so \(z^{\mathrm{obs}}_0=z^{\mathrm{cmp}}_0\) and \(z^{\mathrm{obs}}_L=z^{\mathrm{cmp}}_L\). By definition, \[\sum_{\ell=0}^{L-1} g_\ell = \sum_{\ell=0}^{L-1}\langle \Delta z^{\mathrm{obs}}_\ell-\Delta z^{\mathrm{cmp}}_\ell,e_y\rangle = \langle z^{\mathrm{obs}}_L-z^{\mathrm{obs}}_0-(z^{\mathrm{cmp}}_L-z^{\mathrm{cmp}}_0),e_y\rangle =0.\] For every \(j\in\{1,\dots,L\}\), the same telescoping computation gives \[\sum_{\ell=0}^{j-1} g_\ell = \langle z^{\mathrm{obs}}_j-z^{\mathrm{obs}}_0-(z^{\mathrm{cmp}}_j-z^{\mathrm{cmp}}_0),e_y\rangle = \langle z^{\mathrm{obs}}_j-z^{\mathrm{cmp}}_j,e_y\rangle,\] because the comparison path shares the same initial state as the observed path. ◻
For fixed endpoints, collect the interior variables into \(\mathbf{z}:=(z_1,\dots,z_{L-1})\in(\mathbb{R}^K)^{L-1}\). The quadratic part is \[Q[z]=\alpha\,\mathbf{z}^\top (D_1^\top D_1\otimes I_K)\mathbf{z} + \beta\,\mathbf{z}^\top (D_2^\top D_2\otimes I_K)\mathbf{z} + \text{boundary terms},\] where \(D_1,D_2\) are endpoint-constrained difference matrices. Their nullspaces are trivial on zero-endpoint perturbations, so the quadratic Hessian is positive definite whenever \((\alpha,\beta)\neq(0,0)\); convex backtracking and tube constraints preserve uniqueness over the feasible set.
In the endpoint-constrained regime, the quadratic part of the action is \(\mu_0\)-strongly convex in the interior variables, where \[\mu_0 := 8\alpha\sin^2\!\Bigl(\frac{\pi}{2L}\Bigr) + 32\beta\sin^4\!\Bigl(\frac{\pi}{2L}\Bigr).\]
Proof. of Proposition [prop:strongconv95appendix]. The Hessian is \(H=2\alpha(D_1^\top D_1\otimes I_K)+2\beta(D_2^\top D_2\otimes I_K)\). The smallest eigenvalues of \(D_1^\top D_1\) and \(D_2^\top D_2\) are \(4\sin^2(\pi/(2L))\) and \(16\sin^4(\pi/(2L))\), respectively; summing the contributions gives the bound. ◻
Proof. of Theorem 3. With interior variables \(\mathbf{z}=(z_1,\dots,z_{L-1})\), the tube is convex and, for \(\gamma=0\), the objective has Hessian \(H=2\alpha(D_1^\top D_1\otimes I_K)+2\beta(D_2^\top D_2\otimes I_K)\). By the convexity facts above, \(H\) is positive definite whenever \((\alpha,\beta)\neq(0,0)\); adding the convex backtracking term preserves uniqueness over the convex tube. On inactive interior layers, varying \(z^\star+\tau v\) with \(v_0=v_L=0\) gives \[0=\alpha\sum_{\ell=0}^{L-1}\langle \Delta z_\ell^\star,\Delta v_\ell\rangle +\beta\sum_{\ell=1}^{L-1}\langle \Delta^2 z_\ell^\star,\Delta^2 v_\ell\rangle.\] Discrete summation by parts yields the coefficient \(-\alpha\Delta^2 z_\ell^\star+\beta\Delta^4 z_\ell^\star\), hence \(\beta\Delta^4 z_\ell^\star=\alpha\Delta^2 z_\ell^\star\). Active Voronoi-face constraints add KKT normal-cone terms. ◻
Proof. of Proposition [prop:exact95gap]. Write \(S=Q+\gamma B\), where \(Q\) is the quadratic part and \(B\) the convex backtracking term. Let \(C:=\mathcal{T}_\rho(z^{\mathrm{TF}})\) and \(e:=z-z^\star\). Optimality gives \(s\in\partial B(z^\star)\) and \(n\in N_C(z^\star)\) such that \[\nabla Q(z^\star)+\gamma s+n=0.\] Since \(\langle n,e\rangle\le0\) for feasible \(z\in C\), convexity of \(B\) gives \[\gamma\bigl(B[z]-B[z^\star]\bigr)\ge \gamma\langle s,e\rangle=-\langle \nabla Q(z^\star),e\rangle-\langle n,e\rangle.\] Therefore \[S[z]-S[z^\star] \ge Q[z]-Q[z^\star]-\langle \nabla Q(z^\star),e\rangle-\langle n,e\rangle \ge Q[z]-Q[z^\star]-\langle \nabla Q(z^\star),e\rangle.\] Expanding the quadratic part yields \[Q[z]-Q[z^\star]-\langle \nabla Q(z^\star),e\rangle = \alpha\sum_{\ell=0}^{L-1}\lVert \Delta e_\ell\rVert^2 +\beta\sum_{\ell=1}^{L-1}\lVert \Delta^2 e_\ell\rVert^2.\] This proves the lower bound. When \(\gamma=0\) and the active normal-cone pairing vanishes, the bound is tight. ◻
Proof. of Corollary 3. The step and curvature bounds are immediate from Proposition [prop:exact95gap]. Strong convexity on endpoint-constrained variables gives \[S[\mathbf{z}] \ge S[\mathbf{z}^\star] + \frac{\mu}{2}\lVert \mathbf{z}-\mathbf{z}^\star\rVert^2 = S[\mathbf{z}^\star]+\frac{\mu}{2}\sum_{\ell=1}^{L-1}\lVert z_\ell-z_\ell^\star\rVert^2.\] For the contribution-profile bound, let \(e_\ell:=z^{\mathrm{obs}}_\ell-z_\ell^\star\). Then \(e_0=e_L=0\) and \(g_\ell=\langle \Delta e_\ell,e_y\rangle\). Since \(\lVert e_y\rVert=1\), \[\sum_{\ell=0}^{L-1}|g_\ell| \le \sum_{\ell=0}^{L-1}\lVert \Delta e_\ell\rVert \le \sqrt{L}\Bigl(\sum_{\ell=0}^{L-1}\lVert \Delta e_\ell\rVert^2\Bigr)^{1/2}.\] Proposition [prop:exact95gap] bounds the final factor by \(\alpha^{-1/2}\Delta S^{1/2}\), and every partial sum is bounded by the total variation. ◻
Lemma 1 (Block interpolation bound). Let \(e_0,\dots,e_n\in\mathbb{R}^K\) satisfy \(e_0=e_n=0\). Then \[\sum_{k=0}^{n}\lVert e_k\rVert^2 \le 2n^4 \sum_{k=1}^{n-1}\lVert \Delta^2 e_k\rVert^2.\]
Proof. of Lemma 1. Write \(g_k:=e_k-e_{k-1}\) for \(k=1,\dots,n\). Since \(e_0=e_n=0\), one has \(\sum_{k=1}^n g_k=0\). Hence, for each \(k\), \[g_k=\frac{1}{n}\sum_{i=1}^n (g_k-g_i).\] If \(i<k\), then \(g_k-g_i=\sum_{j=i}^{k-1}\Delta^2 e_j\), while if \(i>k\), then \(g_k-g_i=-\sum_{j=k}^{i-1}\Delta^2 e_j\). Therefore \[\lVert g_k\rVert\le \sum_{j=1}^{n-1}\lVert \Delta^2 e_j\rVert\le \sqrt{n-1}\Bigl(\sum_{j=1}^{n-1}\lVert \Delta^2 e_j\rVert^2\Bigr)^{1/2}.\] Summing the first differences gives, for every \(k\), \[\lVert e_k\rVert=\Bigl\|\sum_{i=1}^{k} g_i\Bigr\|\le n^{3/2}\Bigl(\sum_{j=1}^{n-1}\lVert \Delta^2 e_j\rVert^2\Bigr)^{1/2}.\] Squaring and summing over \(k=0,\dots,n\) yields \[\sum_{k=0}^{n}\|e_k\|^2 \le (n+1)n^3 \sum_{j=1}^{n-1}\|\Delta^2 e_j\|^2 \le 2n^4 \sum_{j=1}^{n-1}\|\Delta^2 e_j\|^2,\] which proves the claim. ◻
Proof. of Theorem 4. Let the interval \(\{0,\dots,L\}\) be partitioned into \(m\) contiguous blocks \(I_j=[s_j,s_{j+1}]\) with lengths \(n_j:=s_{j+1}-s_j\le \lceil L/m\rceil\). Let \(\tilde{z}\) be the piecewise-linear interpolant that agrees with \(z\) at all block endpoints. On each block, define the interpolation error \(e_\ell:=z_\ell-\tilde{z}_\ell\). Then \(e_{s_j}=e_{s_{j+1}}=0\) and \(\Delta^2 e_\ell=\Delta^2 z_\ell\) on interior indices of the block. Applying Lemma 1 on each block gives \[\sum_{\ell\in I_j}\|e_\ell\|^2 \le 2n_j^4 \sum_{\ell=s_j+1}^{s_{j+1}-1}\|\Delta^2 z_\ell\|^2.\] Summing over blocks and using \(n_j\le \lceil L/m\rceil\) gives \[\sum_{\ell=0}^L\|z_\ell-\tilde{z}_\ell\|^2 \le 2\Bigl\lceil\frac{L}{m}\Bigr\rceil^4 E_2(z) \le c_1\frac{L^4}{m^4}E_2(z)\] for a universal constant \(c_1\). Dividing by \(L+1\) yields the stated bound after absorbing fixed numerical factors into the universal constant \(c\). ◻