March 20, 2026
We study the ill-posed problem of recovering a probability measure flow from finitely many moving localized sensors using a Bayes Hilbert framework. Relative to a fixed reference probability measure, a probability law is represented by its centered log-ratio coordinates, so that an evolving law becomes a path in a Hilbert space of functions. For sufficiently regular Bayes Hilbert paths, we construct a canonical minimum-energy transport realization of the path by solving a weighted Neumann problem at each time, which yields an intrinsic transport form on tangent directions.
We then formulate an inverse problem directly on Bayes Hilbert path space. Linearization of an observation operator yields an observability form, and recoverability is governed by its interaction with the transport geometry through a joint transport–observability form. In the ambient infinite-dimensional setting, we develop the corresponding regularized variational theory and identify a basic limitation of localized sensing: mobile sensors can make the joint form injective, but they do not in general yield a coercive stability estimate on the full state space.
This obstruction leads naturally to finite-dimensional Bayes Hilbert reductions. There the transport form becomes a kinetic tensor and the linearized observations become reduced sensing matrices, so recoverability can be expressed through explicit Gramian conditions. We show that localized bump sensors detect every fixed reduced direction, that finitely many suitably placed static sensors yield uniform reduced observability, and there exist path-dependent sensor trajectories such that even a single moving sensor can recover the reduced path. Finally, we show that these reduced recovery results lift to approximate ambient recovery for paths that are well approximated by the chosen finite-dimensional subspaces, yielding stable reconstruction up to projection error.
We study the recovery of a time-dependent probability law from indirect, time-dependent observations. Fix a bounded domain \(\Omega \subset \mathbb{R}^d\) and a reference probability measure \(\nu_0\). In Bayes Hilbert coordinates, a probability measure \(\rho\) is represented using the centered log-ratio transform [1] \[h=\mathop{\mathrm{clr}}(\rho)\in L^2_0(\nu_0), \qquad L^2_0(\nu_0) = \left\{ h \in L^2(\nu_0) : \int_\Omega h \; \mathrm{d}\nu_0 = 0\right\}.\] Using this transform, an evolving law \(\rho_t\) becomes a path \(h(t)\) in a Hilbert space of functions. The inverse problem is to reconstruct the unknown path \(h(\cdot)\), equivalently the law path \(t\mapsto \rho_{h(t)}\), from sensor data. We show that for any measure flow \(\rho_{h(t)}\) lying in a finite-dimensional Bayes Hilbert subspace \(V_m\), there exist a configuration of finitely many static localized sensors making the path fully observable at each time. Furthermore, for every finite-dimensional \(\rho_t\) there exists a (\(\rho_t\)-dependent) sensor trajectory in \(\Omega\) such that a single moving sensor can fully recover \(\rho_t\) even though the full state is not observed at any time. We then show that these finite-dimensional recovery results lift to approximate recovery of \(\rho_t\) in the full infinite dimensional setting, with error determined by the intrinsic error of projecting \(\rho_t\).
Unlike the finite-dimensional setting, localized sensing does not in general yield a coercive stability estimate on a full infinite-dimensional Bayes Hilbert state space. Mobile sensing can make the relevant bilinear form injective, but it does not by itself restore ambient coercivity except in the trivial fully observed case. This obstruction explains why finite-dimensional Bayes Hilbert reductions are not merely technical approximations, but the natural level at which localized sensing leads to genuine recovery theorems.
In this paper, we first construct a canonical transport field associated to a prescribed Bayes Hilbert path. For sufficiently regular \(h\), a canonical velocity is determined by a weighted Neumann problem at each time. This selects the unique transport realization of minimum kinetic energy compatible with the instantaneous evolution of the law and induces an intrinsic transport form \(\mathfrak g_h(\xi,\zeta)\) on tangent directions. Thus a Bayes Hilbert path carries both a law-valued evolution and a distinguished transport geometry.
The inverse problem is formulated directly on path space. If \(\mathcal{G}_t^y\) denotes the observation operator at time \(t\), with \(y\) the sensor trajectory, then linearization at a state \(h\) yields an observability form \[\mathfrak j_{t,h}^y(\xi,\zeta) := \bigl\langle D\mathcal{G}_t^y(h)\xi,\;D\mathcal{G}_t^y(h)\zeta\bigr\rangle .\] Recoverability is governed by the interaction between \(\mathfrak g_h\) and \(\mathfrak j_{t,h}^y\). The observation term determines which directions are visible to the data, while the transport term penalizes the time variation required for a perturbation to remain hidden as the sensing configuration changes. This leads naturally to a joint transport–observability form on perturbation paths, which is the central object driving our analysis.
This perspective also clarifies the role of sensor motion. In finite dimensions, moving sensors can rotate nullspaces of the observability form over time, so that directions invisible at any single time become recoverable over the observation horizon. In infinite dimensions, by contrast, localized sensing remains compact at each fixed time, and this compactness prevents full ambient coercivity. The finite-dimensional theory makes this mechanism explicit: on \(V_m\), the transport form becomes a reduced kinetic tensor and the linearized observations become reduced sensing matrices, so recoverability can be expressed through concrete conditions on a Gramian.
The Bayes Hilbert framework is well suited to this program because it cleanly separates three distinct objects: the measure path itself, its dynamical transport realization, and its recovery from data. The primary unknown is the coordinate path \(h(\cdot)\). The corresponding probability law is obtained by exponentiation and normalization, while the canonical dynamics are recovered from \(h(\cdot)\) through the weighted Neumann problem. This separation allows the forward and inverse theories to be developed using Hilbert space techniques that are often not available when working with probability measures; other geometries, such as Wasserstein and Fisher-Rao, construct infinite-dimensional Riemannian manifolds of probability measures rather than vector spaces.
The remainder of the paper follows a forward-to-inverse progression. Section 2 recalls basic facts of Bayes Hilbert spaces, the \(\mathop{\mathrm{clr}}\) map, and its inverse. Section 3 develops the canonical forward realization and the transport form \(\mathfrak g_h\). Section 4 introduces and analyzes our inverse problem in the ambient infinite-dimensional setting. We introduce the regularized variational problem, prove local stability, introduce the observability and joint transport–observability forms, and identify the limitations of localized sensing in infinite dimensions. Section 5 then passes to finite-dimensional Bayes Hilbert subspaces, proves the static- and moving-sensor recovery results, and lifts them to approximate ambient recovery.
Our Bayes Hilbert framework is related to the theory of metric gradient flows of probability measures [2]–[4], but the conceptual starting point of the present paper is different. We do not begin with a fixed energy functional and then derive its steepest-descent evolution under a prescribed metric. Instead, we begin with an arbitrary prescribed path \(t \mapsto h(t)\) in Bayes Hilbert coordinates and then solve a weighted Neumann problem to recover the unique gradient velocity field of minimum kinetic energy that realizes this path in the continuity equation. In this sense, our transport form \(\mathfrak g_h\) is an execution geometry induced by dynamical realization of a prescribed log-density path, rather than a Riemannian metric used to define a gradient flow of a fixed functional. The inverse problem is likewise formulated directly on path space, with observability encoded by the companion form \(\mathfrak j_h\); we are not aware of this path-space forward–inverse pairing having a direct analogue in the gradient-flow literature.
Some gradient flows, particularly Fisher–Rao gradient flows, arise as special cases of our framework. The geometric annealing path \[\rho_t \propto \rho_0^{\,1-t}\rho_1^{\,t}\] appearing in Fisher–Rao-based sampling and continuation methods is simply a straight-line in Bayes Hilbert coordinates: if \(h_i=\mathop{\mathrm{clr}}(\rho_i)\), then \[\mathop{\mathrm{clr}}(\rho_t) = (1-t)h_0 + t h_1.\] Thus the Fisher–Rao annealing path is contained in our framework as a distinguished special case. This observation is consistent with recent work showing that geometric annealing has a Fisher–Rao gradient-flow interpretation and can be dynamically realized by solving an elliptic or Poisson-type equation for a transport velocity [5]–[8]. In Section 3, we extend these techniques to paths that are not straight-lines in Bayes Hilbert space.
The relation to Wasserstein–Fisher–Rao (WFR), also called Hellinger–Kantorovich (HK), is different in a more fundamental way. The WFR/HK geometry is an unbalanced transport theory on nonnegative measures: it interpolates between quadratic Wasserstein transport and Fisher–Rao reaction, and its dynamic formulation allows source terms in the continuity equation [9]–[11]. By contrast, our present theory is formulated on normalized probability measures and, once a Bayes Hilbert path has been chosen, produces a conservative continuity equation \[\partial_t \rho_t + \nabla \cdot (\rho_t v_t) = 0.\] For this reason, our framework should not be viewed as a special case of WFR/HK gradient-flow theory. Rather, it provides a complementary log-density coordinate formalism for path design, minimum-energy dynamical realization, and inverse reconstruction on spaces of probability laws.
Bayes Hilbert spaces are closely related to the field of information geometry [12]–[14], but they encode a different geometric choice. In information geometry, a family of probability distributions is treated as a statistical manifold equipped with the Fisher metric and a dual pair of affine connections, with exponential and mixture coordinates playing a central role. In Bayes Hilbert space, by contrast, one fixes a reference measure and represents a law by its centered log-density, thereby obtaining a global Hilbert-space model in which addition is Bayes updating and affine subspaces correspond to exponential families. The common thread is the privileged role of log-densities and exponential families; the difference is that information geometry is primarily a local Riemannian differential geometry on statistical manifolds, whereas the Bayes Hilbert approach is a global linear functional-analytic geometry. For a more thorough comparison between the two geometries, we refer to [15].
Recent machine-learning-oriented uses of Bayes Hilbert spaces remain comparatively limited, but they point in several distinct directions. Barfoot and D’Eleuterio formulate variational inference in a Bayesian Hilbert space and show that, under suitable conditions, KL-based variational inference can be interpreted as iterative Euclidean projection of the posterior onto a chosen approximation family; they emphasize applications to high-dimensional robotic state estimation and sparsity-aware inference [16]. Wynne studies Bayes Hilbert spaces as a framework for posterior approximation, highlighting connections to Bayesian coreset constructions and kernel-based discrepancies for comparing posterior and pseudo-posterior measures [17]. More recently, Lach, Fottner, and Okhrin introduce pseudo \(f\)-divergences that extend classical \(f\)-divergences to include the metric induced by Bayes Hilbert spaces, and they develop a variational estimation framework with a generative-adversarial interpretation, reporting strong empirical performance relative to standard \(f\)-GAN variants and competitive results against Wasserstein GANs [18].
The literature already contains several nearby formulations of inverse problems for measures, though not, to our knowledge, the particular Bayes Hilbert path-recovery problem studied here. The closest precedent on the inverse-problems side is the work of Bredies and Fanzon on dynamic inverse problems in spaces of measures, where the unknown is a time-dependent curve of Radon measures and reconstruction is regularized by balanced or unbalanced dynamic optimal transport [19]. In a different but clearly related direction, Li, Oprea, Wang, and Yang study stochastic inverse problems in which the unknown is itself a probability law, and subsequently formulate inverse problems directly over probability measure space through pushforward constraints; these works are static rather than path-valued, but they place inverse problems for distributions on a rigorous infinite-dimensional footing [20], [21]. There is also a neighboring line of work on recovering dynamics from ensemble snapshot data, including system-identification formulations based on distributional evolution and, more recently, Schrödinger-bridge-based reconstruction from snapshot measurements [22], [23]. Relative to these works, our contribution is to formulate indirect recovery of a measure flow in Bayes Hilbert coordinates, to couple that recovery with a canonical minimum-energy dynamical realization obtained from weighted Neumann problems, and to identify the joint transport–observability form as the object governing recoverability in both the ambient infinite-dimensional setting and in concrete finite-dimensional reductions.
We provide a brief overview of the most important facts about Bayes Hilbert spaces for this work. Our goal is not to present the theory in full generality, but rather we tailor our presentation to streamline applicability to our later results. We refer to the original papers [1], [24] for additional details.
Bayes Hilbert spaces provide a linear coordinate model for strictly positive probability measures. Once a reference probability measure \(\nu_0\) is fixed, a probability measure can be encoded by its centered log-density using the \(\mathop{\mathrm{clr}}\)-transform (Section 2.1), and recovered by an exponentiation-normalization map (Section 2.2). This allows us to work with evolving measures through ordinary function-valued paths (Section 2.3). The key computation of this section is the differential of the exponential-normalization map (Proposition 4), which produces the centered forcing term driving the weighted Neumann problems in Section 3.
For the remainder of the paper, let \(\Omega \subset \mathbb{R}^d\) be open and connected, and let \(\nu_0 \ll \mathrm{d}x\) be a fixed probability measure with \(\mathop{\mathrm{supp}}(\nu_0)=\overline{\Omega}\). We define the centered \(L^2(\nu_0)\) functions by \[L^2_0(\nu_0) := \left\{ h\in L^2(\nu_0) : \int_\Omega h\,\mathrm{d}\nu_0 = 0 \right\}.\]
Let \(\rho\) and \(\eta\) be positive measures on \(\Omega\) that are each equivalent to \(\nu_0\) (that is, \(\rho \ll \nu_0\) and \(\nu_0 \ll \rho\), and similarly for \(\eta\)). We say that \(\rho\) and \(\eta\) are Bayes-equivalent, and write \(\rho\sim_B \eta\), if there exists \(c>0\) such that \(\rho=c\,\eta\). In particular, Bayes-equivalent measures are proportional, and the total mass of a measure is invisible to Bayes Hilbert methods.
We denote the Radon-Nikodym derivative of \(\rho\) with respect to \(\nu_0\) by \(\frac{\mathrm{d}\rho}{\mathrm{d}\nu_0}\).
Definition 1 (Bayes Hilbert space). The Bayes Hilbert space relative to \(\nu_0\) is \[B^2(\nu_0) := \left\{ \rho : \rho \ll \nu_0,\;\rho>0\;\nu_0\text{-a.e.},\; \log \frac{\mathrm{d}\rho}{\mathrm{d}\nu_0}\in L^2(\nu_0) \right\}\Big/\sim_B.\]
Definition 2 (Centered log-ratio transform). For \(\rho\in B^2(\nu_0)\), define \[\mathop{\mathrm{clr}}(\rho) := \log \frac{\mathrm{d}\rho}{\mathrm{d}\nu_0} - \int_\Omega \log \frac{\mathrm{d}\rho}{\mathrm{d}\nu_0}\,\mathrm{d}\nu_0.\]
The \(\mathop{\mathrm{clr}}\) transform is well defined on Bayes-equivalence classes and takes values in \(L^2_0(\nu_0)\). It is the basic coordinate map on \(B^2(\nu_0)\). With this in mind, we can use \(\mathop{\mathrm{clr}}\) to pull back the inner-product from \(L^2_0(\nu_0)\) to \(B^2(\nu_0)\), making the following proposition true.
Proposition 1 (Hilbert structure [1]). The map \[\mathop{\mathrm{clr}}: B^2(\nu_0)\to L^2_0(\nu_0)\] is an isometric isomorphism between Hilbert spaces.
Remark 2. The vector space operations on \(B^2(\nu_0)\) are not the usual pointwise addition and scalar multiplication, but rather perturbation and powering: \[\begin{align} & \rho \oplus \mu \text{ is represented by the density } \frac{\mathrm{d}\rho}{\mathrm{d}\nu_0} \frac{\mathrm{d}\mu}{\mathrm{d}\nu_0}, \qquad \rho, \mu \in B^2(\nu_0) \\ \qquad & \alpha \odot \rho \text{ is represented by the density } \left( \frac{\mathrm{d}\rho}{\mathrm{d}\nu_0} \right)^\alpha, \quad \alpha \in \mathbb{R}. \end{align}\] The perturbation \(\rho \oplus \mu\) corresponds to Bayesian updating—multiplying the density of \(\rho\) by that of \(\mu\)—while the powering \(\alpha \odot \rho\) raises the density to the power \(\alpha\). For this manuscript, it is our preference to work in the \(\mathop{\mathrm{clr}}\)-transformed space \(L^2_0(\nu_0)\) where the familiar addition and scalar multiplication apply, rather than directly in \(B^2(\nu_0)\) with \(\oplus\) and \(\odot\).
Remark 3. In the linear structure of \(B^2(\nu_0)\), the reference \(\nu_0\) becomes the neutral element (i.e. the zero vector). This fact is reflected in our choice of notation for \(\nu_0\).
Although \(B^2(\nu_0)\) is formally a space of Bayes-equivalence classes, and \(B^2(\nu_0)\) naturally includes both finite and infinite measures, in this paper we will work only with finite measures, and in particular with the unique probability representative of each equivalence class of finite measures.
To pass from centered log-density coordinates back to probability measures, we introduce the normalized exponential map. At the level of Bayes classes, the inverse of the \(\mathop{\mathrm{clr}}\) transform is the exponential map \(h \mapsto e^h \nu_0\); however, \(e^h \nu_0\) is not necessarily a finite measure. For simplicity, we therefore work in the bounded coordinate class \[\mathcal{H} := L^2_0(\nu_0) \cap L^\infty(\nu_0)\] equipped with the \(L^\infty(\nu_0)\)-norm (since \(\nu_0\) is a probability measure, this is stronger than the \(L^2(\nu_0)\) norm). For every \(h \in \mathcal{H}\), the measure \(e^h \nu_0\) is finite and can be normalized to a probability measure. This class is precisely the \(\mathop{\mathrm{clr}}\)-image of the bounded-density class [1] \[B^2_b(\nu_0) := \left\{ \rho \in B^2(\nu_0) : \exists\, b>0 \text{ such that } \frac{1}{b} < \frac{\mathrm{d}\rho}{\mathrm{d}\nu_0} < b \quad \nu_0\text{-a.e.} \right\}.\] Although \(B^2_b(\nu_0)\) does not contain all finite elements of \(B^2(\nu_0)\), it avoids technicalities when considering tangent directions for \(h\) in the Propositions below, and it contains all measures satisfying the additional regularity assumptions required in Section 3 onwards, so it is not unduly restrictive.
Definition 3 (Exponential-normalization map). For \(h \in \mathcal{H}\), define \[\label{eq:rho95h95def} \rho_h := \frac{e^h}{\int_\Omega e^h\,\mathrm{d}\nu_0}\,\nu_0.\tag{1}\]
The Radon-Nikodym derivative of \(\rho_h\) is \[\frac{\mathrm{d}\rho_h}{\mathrm{d}\nu_0} = \frac{e^h}{\int_\Omega e^h\,\mathrm{d}\nu_0}.\] By construction, \(\rho_h\) is a probability measure, \(\rho_h\ll \nu_0\), and \(\mathop{\mathrm{clr}}(\rho_h)=h\). The differential of the exponential-normalization map is one of the basic objects used later.
Proposition 4 (Fréchet differentiability of the exponential-normalization map). Define \(\mathcal{E}:\mathcal{H} \to L^1(\nu_0)\), \(\mathcal{E}(h):=\frac{\mathrm{d}\rho_h}{\mathrm{d}\nu_0}\). Then \(\mathcal{E}\) is Fréchet differentiable with derivative \[\label{eq:diff95exp95normalization} D\mathcal{E}(h)[\xi] = \frac{\mathrm{d}\rho_h}{\mathrm{d}\nu_0} \bigl(\xi-\mathbb{E}_{\rho_h}[\xi]\bigr).\qquad{(1)}\]
Proof. Write \[\mathcal{N}(h):=e^h \in L^1(\nu_0), \qquad Z(h):=\int_\Omega e^h\,\mathrm{d}\nu_0 \in (0,\infty).\] Then \[\mathcal{E}(h)=\frac{\mathcal{N}(h)}{Z(h)}.\]
We first show that \(\mathcal{N}:\mathcal{H}\to L^1(\nu_0)\) is Fréchet differentiable with \[D\mathcal{N}(h)[\xi]=e^h\xi.\] Indeed, \[\mathcal{N}(h+\xi)-\mathcal{N}(h)-e^h\xi = e^h\bigl(e^\xi-1-\xi\bigr).\] Using the elementary bound \[|e^r-1-r|\le \tfrac12 e^{|r|}r^2, \qquad r\in\mathbb{R},\] we obtain \[\|\mathcal{N}(h+\xi)-\mathcal{N}(h)-e^h\xi\|_{L^1(\nu_0)} \le \frac{1}{2} e^{\|h\|_{L^\infty}+\|\xi\|_{L^\infty}} \|\xi\|_{L^\infty}^2 = o(\|\xi\|_{L^\infty}).\] Here we used that \(\nu_0\) is a probability measure.
Similarly, \[Z(h+\xi)-Z(h)-\int_\Omega e^h\xi\,\mathrm{d}\nu_0 = \int_\Omega e^h\bigl(e^\xi-1-\xi\bigr)\,\mathrm{d}\nu_0,\] so \(Z:\mathcal{H}\to \mathbb{R}\) is Fréchet differentiable with \[DZ(h)[\xi]=\int_\Omega e^h\xi\,\mathrm{d}\nu_0.\]
Since \(Z(h)>0\) and inversion on \((0,\infty)\) is smooth, the quotient rule yields that \(\mathcal{E}=\mathcal{N}/Z\) is Fréchet differentiable, with \[D\mathcal{E}(h)[\xi] = \frac{e^h\xi}{Z(h)} - \frac{e^h}{Z(h)^2}\int_\Omega e^h\xi\,\mathrm{d}\nu_0 = \frac{e^h}{Z(h)} \left( \xi - \frac{1}{Z(h)}\int_\Omega e^h\xi\,\mathrm{d}\nu_0 \right) = \frac{\mathrm{d}\rho_h}{\mathrm{d}\nu_0} \bigl(\xi-\mathbb{E}_{\rho_h}[\xi]\bigr).\] ◻
Remark 5. The derivative formula ?? shows that a coordinate perturbation \(\xi\) induces the density variation \[\delta \rho = \rho_h\bigl(\xi-\mathbb{E}_{\rho_h}[\xi]\bigr).\] Thus Bayes Hilbert tangent directions are automatically centered with respect to the current law. This centered forcing term will be the source term in the weighted Neumann problems considered later.
We now pass from single states to continuous paths \(h(t)\) along a time interval \(t \in [0, T]\), and we derive the basic time-differentiation formula that drives the transport theory in Section 3.
Definition 4 (Regular coordinate path). A regular coordinate path is a map \(h\in C^1([0,T]; \mathcal{H})\). For such a path we define \(\rho_t := \rho_{h(t)}.\)
Proposition 6 (Log-density evolution along coordinate paths). Let \(h\in C^1([0,T];\mathcal{H})\) be a regular coordinate path. Then \[\label{eq:ambient95density95derivative} \partial_t \frac{\mathrm{d}\rho_t}{\mathrm{d}\nu_0} = \frac{\mathrm{d}\rho_t}{\mathrm{d}\nu_0} \Bigl(\dot{h}(t)-\mathbb{E}_{\rho_t}[\dot{h}(t)]\Bigr),\qquad{(2)}\] and hence \[\label{eq:ambient95log95density95derivative} \partial_t \log \frac{\mathrm{d}\rho_t}{\mathrm{d}\nu_0} = \dot{h}(t)-\mathbb{E}_{\rho_t}[\dot{h}(t)].\qquad{(3)}\]
Proof. Equation ?? follows from Proposition 4 and the Banach space chain rule. Equation ?? is obtained by dividing ?? by \(\mathrm{d}\rho_t/\mathrm{d}\nu_0\), which is positive \(\nu_0\)-a.e. ◻
In this section we develop the forward geometry of regular Bayes Hilbert paths in a form adapted to the bounded-domain setting used later in the paper. The guiding principle is classical in optimal transport: an infinitesimal density variation should be realized by the velocity field of least kinetic energy among all solutions of the continuity equation. The minimum-energy realization is in the spirit of the Benamou-Brenier/Otto dynamic viewpoint on transport [25], the difference here being the entire law path is prescribed, rather than just the endpoints. Our contribution here is to show that Bayes Hilbert tangent directions induce such minimum-energy realizations through a state-dependent weighted elliptic problem. Closely related results in the context of Reproducing Kernel Hilbert Spaces can also be found in the recent work [26].
For any finite measure \(\mu\ll \mathrm{d}x\), define \[H^1(\Omega;\mu) := \left\{ u\in H^1_{\mathrm{loc}}(\Omega): u\in L^2(\mu),\;\nabla u\in L^2(\mu;\mathbb{R}^d) \right\},\] with norm \[\|u\|_{H^1(\Omega;\mu)}^2 := \|u\|_{L^2(\mu)}^2+\|\nabla u\|_{L^2(\mu)}^2.\] We also define the weighted mean-zero space \[H^1_{\mu,\diamond}(\Omega) := \left\{ u\in H^1(\Omega;\mu): \int_\Omega u\,\mathrm{d}\mu =0 \right\}.\]
Assumption 1 (Reference spectral gap and strong state space). Assume:
there exists \(C_{\nu_0}>0\) such that \[\label{eq:reference-poincare} \int_\Omega \left|u-\mathbb{E}_{\nu_0}[u]\right|^2\,\mathrm{d}\nu_0 \le C_{\nu_0} \int_\Omega |\nabla u|^2\,\mathrm{d}\nu_0 \qquad \forall u\in H^1(\Omega;\nu_0);\tag{2}\]
\(\mathcal{Z}\) is a Banach space continuously embedded in \(L^\infty(\nu_0)\cap L^2(\nu_0)\).
\(C_c^\infty(\Omega)\) is dense in \(L^2(\nu_0)\).
Assumption 2 (Admissible state class). Let \(\mathcal{X}_{\mathrm{ad}}\subset \mathcal{Z}\cap L^2_0(\nu_0)\). Assume that there exist constants \(0<c<C<\infty\) such that, for every \(h\in\mathcal{X}_{\mathrm{ad}}\), the density \[w_h:=\frac{\mathrm{d}\rho_h}{\mathrm{d}\nu_0} = \frac{e^h}{\int_\Omega e^h\,\mathrm{d}\nu_0}\] satisfies \[\label{eq:ambient-uniform-density-bounds} c\le w_h(x)\le C \qquad\text{for }\nu_0\text{-a.e. }x\in\Omega.\tag{3}\]
Remark 7. The admissibility condition 3 implies a uniform \(L^\infty(\nu_0)\)-bound on \(h\). Indeed, \[\log w_h = h-\log\!\left(\int_\Omega e^h\,\mathrm{d}\nu_0\right),\] and since \(h\in L^2_0(\nu_0)\), \[h=\log w_h-\mathbb{E}_{\nu_0}[\log w_h].\] Hence \[\|h\|_{L^\infty(\nu_0)} \le 2\max\{|\log c|,\;|\log C|\}.\] Moreover, on the uniformly bounded \(L^\infty(\nu_0)\)-set \(\mathcal{X}_{\mathrm{ad}}\), the map \[h\longmapsto w_h=\frac{e^h}{\int_\Omega e^h\,\mathrm{d}\nu_0}\] is Lipschitz from \(L^\infty(\nu_0)\) to \(L^\infty(\nu_0)\). Since \(\mathcal{Z}\hookrightarrow L^\infty(\nu_0)\) continuously, it follows that \[\label{eq:ambient-weight-lipschitz} \|w_{h_1}-w_{h_2}\|_{L^\infty(\nu_0)} \le M_{\mathcal{X}_{\mathrm{ad}}}\|h_1-h_2\|_{\mathcal{Z}} \qquad \forall h_1,h_2\in\mathcal{X}_{\mathrm{ad}}\tag{4}\] for some constant \(M_{\mathcal{X}_{\mathrm{ad}}}>0\).
Lemma 1 (Equivalence of weighted Sobolev norms). Under Assumption 2, for every \(h\in\mathcal{X}_{\mathrm{ad}}\), \[H^1(\Omega;\rho_h)=H^1(\Omega;\nu_0)\] as sets, and the norms are uniformly equivalent: \[\label{eq:ambient-weighted-sobolev-equivalence} c\|u\|_{H^1(\Omega;\nu_0)}^2 \le \|u\|_{H^1(\Omega;\rho_h)}^2 \le C\|u\|_{H^1(\Omega;\nu_0)}^2 \qquad \forall u\in H^1(\Omega;\nu_0).\tag{5}\] In particular, \[H^1_{\rho_h,\diamond}(\Omega) = \left\{ u\in H^1(\Omega;\nu_0):\int_\Omega u\,\mathrm{d}\rho_h=0 \right\}.\]
Proof. Since \(\mathrm{d}\rho_h=w_h\,\mathrm{d}\nu_0\) and \(c\le w_h\le C\) \(\nu_0\)-a.e., we have \[c\int_\Omega |u|^2\,\mathrm{d}\nu_0 \le \int_\Omega |u|^2\,\mathrm{d}\rho_h \le C\int_\Omega |u|^2\,\mathrm{d}\nu_0,\] and similarly \[c\int_\Omega |\nabla u|^2\,\mathrm{d}\nu_0 \le \int_\Omega |\nabla u|^2\,\mathrm{d}\rho_h \le C\int_\Omega |\nabla u|^2\,\mathrm{d}\nu_0.\] Summing the two inequalities yields 5 . ◻
Proposition 8 (Weighted spectral gap inherited from the reference measure). Under Assumptions 1 and 2, for every \(h\in\mathcal{X}_{\mathrm{ad}}\) and every \(u\in H^1(\Omega;\rho_h)\), \[\label{eq:ambient-weighted-poincare} \int_\Omega \left|u-\mathbb{E}_{\rho_h}[u]\right|^2\,\mathrm{d}\rho_h \le \frac{C}{c}\,C_{\nu_0} \int_\Omega |\nabla u|^2\,\mathrm{d}\rho_h.\qquad{(4)}\] Consequently, the weighted Dirichlet form \[u\longmapsto \int_\Omega |\nabla u|^2\,\mathrm{d}\rho_h\] is coercive on \(H^1_{\rho_h,\diamond}(\Omega)\), with constants uniform in \(h\in\mathcal{X}_{\mathrm{ad}}\).
Proof. Fix \(h\in\mathcal{X}_{\mathrm{ad}}\) and \(u\in H^1(\Omega;\rho_h)\). Since \(u\mapsto \int |u-a|^2\,\mathrm{d}\rho_h\) is minimized at \(a=\mathbb{E}_{\rho_h}[u]\), for every constant \(a\in\mathbb{R}\), \[\int_\Omega \left|u-\mathbb{E}_{\rho_h}[u]\right|^2\,\mathrm{d}\rho_h \le \int_\Omega |u-a|^2\,\mathrm{d}\rho_h.\] Choosing \(a=\mathbb{E}_{\nu_0}[u]\) and using \(\mathrm{d}\rho_h=w_h\,\mathrm{d}\nu_0\), together with 3 , we obtain \[\begin{align} \int_\Omega \left|u-\mathbb{E}_{\rho_h}[u]\right|^2\,\mathrm{d}\rho_h & \le \int_\Omega \left|u-\mathbb{E}_{\nu_0}[u]\right|^2\,\mathrm{d}\rho_h \\ & \le C \int_\Omega \left|u-\mathbb{E}_{\nu_0}[u]\right|^2\,\mathrm{d}\nu_0 \\ & \le C C_{\nu_0}\int_\Omega |\nabla u|^2\,\mathrm{d}\nu_0 \\ & \le \frac{C}{c}\,C_{\nu_0}\int_\Omega |\nabla u|^2\,\mathrm{d}\rho_h, \end{align}\] which proves ?? . ◻
The Bayes Hilbert tangent direction \(\xi\in L^2_0(\nu_0)\) induces the centered density variation \[\rho_h\bigl(\xi-\mathbb{E}_{\rho_h}[\xi]\bigr),\] which is the forcing term in the weighted elliptic problem below.
Theorem 9 (Canonical Neumann potential). Assume Assumptions 1 and 2. Fix \(h\in \mathcal{X}_{\mathrm{ad}}\) and \(\xi\in L^2_0(\nu_0)\). Then there exists a unique \[\psi_{h,\xi}\in H^1_{\rho_h,\diamond}(\Omega)\] such that \[\label{eq:ambient-neumann-weak} \int_\Omega \nabla\psi_{h,\xi}\cdot\nabla\eta\,\mathrm{d}\rho_h = \int_\Omega \bigl(\xi-\mathbb{E}_{\rho_h}[\xi]\bigr)\eta\,\mathrm{d}\rho_h \qquad \forall \eta\in H^1_{\rho_h,\diamond}(\Omega).\tag{6}\] If, in addition, the density of \(\rho_h\) with respect to Lebesgue measure is regular enough, then 6 may be interpreted as the weak form of \[\label{eq:ambient-neumann-strong} -\nabla\cdot(\rho_h\nabla\psi_{h,\xi}) = \rho_h\bigl(\xi-\mathbb{E}_{\rho_h}[\xi]\bigr) \qquad\text{in }\Omega,\tag{7}\] together with the natural zero-flux boundary condition \[\label{eq:ambient-neumann-bc} \rho_h\,\partial_n\psi_{h,\xi}=0 \qquad\text{on }\partial\Omega\tag{8}\] when \(\partial\Omega\neq\varnothing\).
Proof. Define, on \(H^1_{\rho_h,\diamond}(\Omega)\), \[B_h(u,\eta):=\int_\Omega \nabla u\cdot \nabla\eta\,\mathrm{d}\rho_h, \qquad F_{h,\xi}(\eta) := \int_\Omega \bigl(\xi-\mathbb{E}_{\rho_h}[\xi]\bigr)\eta\,\mathrm{d}\rho_h.\] The bilinear form \(B_h\) is continuous. By Proposition 8, \[\|u\|_{L^2(\rho_h)}^2 \le \frac{C}{c}\,C_{\nu_0} \int_\Omega |\nabla u|^2\,\mathrm{d}\rho_h \qquad \forall u\in H^1_{\rho_h,\diamond}(\Omega),\] so \(B_h\) is coercive on \(H^1_{\rho_h,\diamond}(\Omega)\).
Next, \[\|\xi-\mathbb{E}_{\rho_h}[\xi]\|_{L^2(\rho_h)} \le 2\|\xi\|_{L^2(\rho_h)} \le 2\sqrt{C}\,\|\xi\|_{L^2(\nu_0)},\] and therefore, by Cauchy–Schwarz and Proposition 8, \[|F_{h,\xi}(\eta)| \le \|\xi-\mathbb{E}_{\rho_h}[\xi]\|_{L^2(\rho_h)} \|\eta\|_{L^2(\rho_h)} \le C'\|\xi\|_{L^2(\nu_0)}\|\eta\|_{H^1(\Omega;\rho_h)}.\] Thus \(F_{h,\xi}\) is continuous on \(H^1_{\rho_h,\diamond}(\Omega)\), and the conclusion follows from the Lax–Milgram theorem. The strong form is the usual divergence-form interpretation of 6 . ◻
Corollary 1 (Energy estimate for the weighted Neumann solve). Under the assumptions of Theorem 9, there exists \(M>0\), depending only on \(c\), \(C\), and \(C_{\nu_0}\), such that \[\label{eq:ambient-neumann-energy} \|\nabla\psi_{h,\xi}\|_{L^2(\rho_h)} \le M\|\xi\|_{L^2(\nu_0)} \qquad \forall h\in\mathcal{X}_{\mathrm{ad}},\;\forall \xi\in L^2_0(\nu_0).\tag{9}\]
Proof. Take \(\eta=\psi_{h,\xi}\) in 6 . Then \[\int_\Omega |\nabla\psi_{h,\xi}|^2\,\mathrm{d}\rho_h = \int_\Omega \bigl(\xi-\mathbb{E}_{\rho_h}[\xi]\bigr)\psi_{h,\xi}\,\mathrm{d}\rho_h.\] By Cauchy–Schwarz and Proposition 8, \[\int_\Omega |\nabla\psi_{h,\xi}|^2\,\mathrm{d}\rho_h \le \|\xi-\mathbb{E}_{\rho_h}[\xi]\|_{L^2(\rho_h)} \|\psi_{h,\xi}\|_{L^2(\rho_h)} \le M\|\xi\|_{L^2(\nu_0)} \|\nabla\psi_{h,\xi}\|_{L^2(\rho_h)}.\] Canceling the nonzero factor yields 9 . ◻
Remark 10 (Extension of the weak formulation to arbitrary test functions). Since both sides of 6 are unchanged when \(\eta\) is replaced by \(\eta-a\) for a constant \(a\in\mathbb{R}\), the weak formulation extends to all \(\eta\in H^1(\Omega;\rho_h)\): \[\label{eq:ambient-neumann-all-tests} \int_\Omega \nabla\psi_{h,\xi}\cdot\nabla\eta\,\mathrm{d}\rho_h = \int_\Omega \bigl(\xi-\mathbb{E}_{\rho_h}[\xi]\bigr)\eta\,\mathrm{d}\rho_h \qquad \forall \eta\in H^1(\Omega;\rho_h).\tag{10}\]
Proposition 11 (Minimum-energy characterization). Fix \(h\in\mathcal{X}_{\mathrm{ad}}\) and \(\xi\in L^2_0(\nu_0)\). Then \(\nabla\psi_{h,\xi}\) is the unique minimizer of \[\label{eq:ambient-min-energy} \inf_{v\in\mathcal{A}_{h,\xi}} \int_\Omega |v|^2\,\mathrm{d}\rho_h,\qquad{(5)}\] where \[\label{eq:ambient-admissible-velocities} \mathcal{A}_{h,\xi} := \left\{ v\in L^2(\rho_h;\mathbb{R}^d): \int_\Omega v\cdot\nabla\eta\,\mathrm{d}\rho_h = \int_\Omega \bigl(\xi-\mathbb{E}_{\rho_h}[\xi]\bigr)\eta\,\mathrm{d}\rho_h \;\forall \eta\in H^1(\Omega;\rho_h) \right\}.\qquad{(6)}\]
Proof. By Remark 10, \[\int_\Omega \nabla\psi_{h,\xi}\cdot\nabla\eta\,\mathrm{d}\rho_h = \int_\Omega \bigl(\xi-\mathbb{E}_{\rho_h}[\xi]\bigr)\eta\,\mathrm{d}\rho_h \qquad \forall \eta\in H^1(\Omega;\rho_h),\] so \(\nabla\psi_{h,\xi}\in\mathcal{A}_{h,\xi}\).
Now let \(v\in\mathcal{A}_{h,\xi}\). Taking \(\eta=\psi_{h,\xi}\) gives \[\int_\Omega \bigl(v-\nabla\psi_{h,\xi}\bigr)\cdot\nabla\psi_{h,\xi}\,\mathrm{d}\rho_h=0.\] Hence \[\begin{align} \int_\Omega |v|^2\,\mathrm{d}\rho_h & = \int_\Omega |\nabla\psi_{h,\xi}|^2\,\mathrm{d}\rho_h + \int_\Omega |v-\nabla\psi_{h,\xi}|^2\,\mathrm{d}\rho_h \\ & \ge \int_\Omega |\nabla\psi_{h,\xi}|^2\,\mathrm{d}\rho_h, \end{align}\] with equality if and only if \(v=\nabla\psi_{h,\xi}\) in \(L^2(\rho_h;\mathbb{R}^d)\). ◻
The weighted Neumann solve defines a canonical map from Bayes Hilbert tangent directions to minimum-energy velocity fields.
Definition 5 (Canonical transport map). For \(h\in\mathcal{X}_{\mathrm{ad}}\), define \[\mathcal{T}_h : L^2_0(\nu_0)\to L^2(\rho_h;\mathbb{R}^d), \qquad \mathcal{T}_h\xi := \nabla\psi_{h,\xi}.\]
Definition 6 (Intrinsic transport form). For \(h\in\mathcal{X}_{\mathrm{ad}}\) and \(\xi,\zeta\in L^2_0(\nu_0)\), define \[\label{eq:intrinsic-transport-form} \mathfrak g_h(\xi,\zeta) := \int_\Omega \mathcal{T}_h\xi\cdot \mathcal{T}_h\zeta\,\mathrm{d}\rho_h = \int_\Omega \nabla\psi_{h,\xi}\cdot\nabla\psi_{h,\zeta}\,\mathrm{d}\rho_h.\tag{11}\]
Proposition 12 (Basic properties of \(\mathfrak g_h\)). For each \(h\in\mathcal{X}_{\mathrm{ad}}\), the form \(\mathfrak g_h\) is a symmetric, nonnegative bilinear form on \(L^2_0(\nu_0)\). Moreover, \[\mathfrak g_h(\xi,\xi)=0 \quad\Longleftrightarrow\quad \xi=\mathbb{E}_{\rho_h}[\xi] \quad \rho_h\text{-a.e.}\] and hence, for \(\xi\in L^2_0(\nu_0)\), \[\mathfrak g_h(\xi,\xi)=0 \quad\Longleftrightarrow\quad \xi=0 \qquad \nu_0\text{-a.e.}\]
Proof. Bilinearity and symmetry follow from the linearity of \(\xi\mapsto \psi_{h,\xi}\) and the symmetry of the \(L^2(\rho_h)\) inner product. Nonnegativity is immediate from \[\mathfrak g_h(\xi,\xi)=\int_\Omega |\nabla\psi_{h,\xi}|^2\,\mathrm{d}\rho_h.\]
Assume \(\mathfrak g_h(\xi,\xi)=0\). Then \(\nabla\psi_{h,\xi}=0\) \(\rho_h\)-a.e., hence \(\psi_{h,\xi}=0\) in \(H^1_{\rho_h,\diamond}(\Omega)\). By 10 , \[\int_\Omega \bigl(\xi-\mathbb{E}_{\rho_h}[\xi]\bigr)\eta\,\mathrm{d}\rho_h =0 \qquad \forall \eta\in H^1(\Omega;\rho_h).\] In particular, the same holds for every \[\eta=\varphi-\mathbb{E}_{\rho_h}[\varphi], \qquad \varphi\in C_c^\infty(\Omega).\] Since \(C_c^\infty(\Omega)\) is dense in \(L^2(\nu_0)\) by Assumption 1, and since \(\mathrm{d}\rho_h=w_h\,\mathrm{d}\nu_0\) with \(w_h\) bounded above and below by Assumption 2, it is also dense in \(L^2(\rho_h)\). Hence the set \[\left\{ \varphi-\mathbb{E}_{\rho_h}[\varphi]:\varphi\in C_c^\infty(\Omega) \right\}\] is dense in \(L^2_0(\rho_h)\). Therefore \[\xi-\mathbb{E}_{\rho_h}[\xi]=0 \qquad \rho_h\text{-a.e.}\] The converse implication is immediate from 11 .
Finally, since \(\rho_h\) and \(\nu_0\) are equivalent and \(\xi\in L^2_0(\nu_0)\), the identity \(\xi=\mathbb{E}_{\rho_h}[\xi]\) \(\rho_h\)-a.e.implies that \(\xi\) is almost everywhere constant, hence zero by the \(\nu_0\)-mean-zero constraint. ◻
Proposition 13 (Stability of the weighted Neumann solve). Assume Assumptions 1 and 2. Then there exists \(M>0\), depending only on \(\nu_0\), \(c\), \(C\), and the embedding \(\mathcal{Z}\hookrightarrow L^\infty(\nu_0)\), such that for all \(h_1,h_2\in\mathcal{X}_{\mathrm{ad}}\) and all \(\xi_1,\xi_2\in L^2_0(\nu_0)\), \[\label{eq:ambient-neumann-stability-estimate} \inf_{a\in\mathbb{R}} \|\psi_{h_1,\xi_1}-\psi_{h_2,\xi_2}-a\|_{H^1(\Omega;\nu_0)} \le M\Big( \|\xi_1-\xi_2\|_{L^2(\nu_0)} + \|h_1-h_2\|_{\mathcal{Z}} \bigl(\|\xi_1\|_{L^2(\nu_0)}+\|\xi_2\|_{L^2(\nu_0)}\bigr) \Big).\qquad{(7)}\] In particular, \[\label{eq:ambient-neumann-gradient-stability} \|\nabla\psi_{h_1,\xi_1}-\nabla\psi_{h_2,\xi_2}\|_{L^2(\nu_0)} \le M\Big( \|\xi_1-\xi_2\|_{L^2(\nu_0)} + \|h_1-h_2\|_{\mathcal{Z}} \bigl(\|\xi_1\|_{L^2(\nu_0)}+\|\xi_2\|_{L^2(\nu_0)}\bigr) \Big).\qquad{(8)}\]
Proof. Fix \(h_1,h_2\in\mathcal{X}_{\mathrm{ad}}\) and \(\xi_1,\xi_2\in L^2_0(\nu_0)\). Write \[w_i:=\frac{\mathrm{d}\rho_{h_i}}{\mathrm{d}\nu_0}, \qquad q_i:=\xi_i-\mathbb{E}_{\rho_{h_i}}[\xi_i], \qquad \psi_i:=\psi_{h_i,\xi_i}, \qquad i=1,2.\] By 4 , \[\|w_1-w_2\|_{L^\infty(\nu_0)} \le M_{\mathcal{X}_{\mathrm{ad}}}\|h_1-h_2\|_{\mathcal{Z}}.\] Also, \[\|q_1-q_2\|_{L^2(\nu_0)} \le \|\xi_1-\xi_2\|_{L^2(\nu_0)} + \left| \mathbb{E}_{\rho_{h_1}}[\xi_1]-\mathbb{E}_{\rho_{h_2}}[\xi_2] \right|.\] The second term is bounded by \[\begin{align} \left| \int_\Omega \xi_1(w_1-w_2)\,\mathrm{d}\nu_0 \right| + \left| \int_\Omega (\xi_1-\xi_2)w_2\,\mathrm{d}\nu_0 \right| & \le \|\xi_1\|_{L^2(\nu_0)}\|w_1-w_2\|_{L^2(\nu_0)} + \sqrt{C}\,\|\xi_1-\xi_2\|_{L^2(\nu_0)} \\ & \le M\Big( \|\xi_1-\xi_2\|_{L^2(\nu_0)} + \|h_1-h_2\|_{\mathcal{Z}}\|\xi_1\|_{L^2(\nu_0)} \Big). \end{align}\] Hence \[\label{eq:ambient-centered-forcing-stability} \|q_1-q_2\|_{L^2(\nu_0)} \le M\Big( \|\xi_1-\xi_2\|_{L^2(\nu_0)} + \|h_1-h_2\|_{\mathcal{Z}} \bigl(\|\xi_1\|_{L^2(\nu_0)}+\|\xi_2\|_{L^2(\nu_0)}\bigr) \Big).\tag{12}\]
Let \(\delta\psi:=\psi_1-\psi_2\). Subtract the weak formulations: \[\int_\Omega w_1\nabla\delta\psi\cdot\nabla\eta\,\mathrm{d}\nu_0 = \int_\Omega (w_1q_1-w_2q_2)\eta\,\mathrm{d}\nu_0 - \int_\Omega (w_1-w_2)\nabla\psi_2\cdot\nabla\eta\,\mathrm{d}\nu_0\] for all \(\eta\in H^1(\Omega;\nu_0)\). Taking \[\eta:=\delta\psi-\mathbb{E}_{\nu_0}[\delta\psi]\] and using 2 , 9 , 12 , and the uniform bounds on the weights yields \[\|\nabla\delta\psi\|_{L^2(\nu_0)} \le M\Big( \|\xi_1-\xi_2\|_{L^2(\nu_0)} + \|h_1-h_2\|_{\mathcal{Z}} \bigl(\|\xi_1\|_{L^2(\nu_0)}+\|\xi_2\|_{L^2(\nu_0)}\bigr) \Big),\] which proves ?? . Subtracting the \(\nu_0\)-mean and using 2 gives ?? . ◻
Corollary 2 (Continuity of the transport map in the state variable). Assume Assumptions 1 and 2. Then there exists \(M>0\) such that for all \(h_1,h_2\in\mathcal{X}_{\mathrm{ad}}\), \[\|\mathcal{T}_{h_1}-\mathcal{T}_{h_2}\|_{\mathcal{L}(L^2_0(\nu_0),L^2(\nu_0;\mathbb{R}^d))} \le M\|h_1-h_2\|_{\mathcal{Z}}.\] In particular, if \(h_n\to h\) in \(\mathcal{Z}\), then \[\mathcal{T}_{h_n}\to \mathcal{T}_h \quad\text{in }\mathcal{L}(L^2_0(\nu_0),L^2(\nu_0;\mathbb{R}^d)).\]
Proof. Apply Proposition 13 with \(\xi_1=\xi_2=\xi\), then take the supremum over \(\|\xi\|_{L^2(\nu_0)}\le 1\). ◻
To differentiate the weighted Neumann solve with respect to the Bayes Hilbert state, we introduce one additional layer of regularity for admissible perturbation directions.
Assumption 3 (Differentiable admissible directions). There exists a Banach space \(\mathcal{X}\) continuously embedded in \(\mathcal{Z}\), with \[\mathcal{X} \hookrightarrow L^\infty(\nu_0)\cap L^2_0(\nu_0)\] continuously, such that for every \(h,\eta\in\mathcal{X}\) with \(h\in\mathcal{X}_{\mathrm{ad}}\), there exists \(\varepsilon_0>0\) for which \[h+\varepsilon\eta\in\mathcal{X}_{\mathrm{ad}} \qquad \text{for all }|\varepsilon|<\varepsilon_0.\]
Definition 7 (Weighted covariance). For \(h\in \mathcal{X}_{\mathrm{ad}}\) and \(f,g\in L^2(\rho_h)\), define \[\operatorname{Cov}_{\rho_h}(f,g) := \int_\Omega \bigl(f-\mathbb{E}_{\rho_h}[f]\bigr) \bigl(g-\mathbb{E}_{\rho_h}[g]\bigr)\,\mathrm{d}\rho_h.\]
Lemma 2 (Directional derivative of the normalized weight). Assume Assumption 3. Let \(h,\eta\in\mathcal{X}\), with \(h\in\mathcal{X}_{\mathrm{ad}}\), and assume that \(h+\varepsilon\eta\in\mathcal{X}_{\mathrm{ad}}\) for all \(|\varepsilon|<\varepsilon_0\). Define \[a_h(\eta):=\eta-\mathbb{E}_{\rho_h}[\eta].\] Then \[\label{eq:ambient-weight-derivative} \frac{w_{h+\varepsilon\eta}-w_h}{\varepsilon} \to w_h\,a_h(\eta) \qquad\text{in }L^\infty(\nu_0)\tag{13}\] as \(\varepsilon\to 0\). Consequently, for every \(\xi\in L^2(\nu_0)\), \[\label{eq:ambient-expectation-derivative} \frac{\mathbb{E}_{\rho_{h+\varepsilon\eta}}[\xi]-\mathbb{E}_{\rho_h}[\xi]}{\varepsilon} \to \operatorname{Cov}_{\rho_h}(\eta,\xi).\tag{14}\]
Proof. Write \[Z(h):=\int_\Omega e^h\,\mathrm{d}\nu_0, \qquad w_h=\frac{e^h}{Z(h)}.\] Since \(h,\eta\in L^\infty(\nu_0)\) and \(h+\varepsilon\eta\) remains in the uniformly bounded admissible set for \(|\varepsilon|<\varepsilon_0\), the map \[\varepsilon\longmapsto \frac{e^{h+\varepsilon\eta}}{Z(h+\varepsilon\eta)}\] is differentiable in \(L^\infty(\nu_0)\), with derivative \[w_h\bigl(\eta-\mathbb{E}_{\rho_h}[\eta]\bigr) = w_h\,a_h(\eta).\] This proves 13 . Then, for \(\xi\in L^2(\nu_0)\), \[\mathbb{E}_{\rho_{h+\varepsilon\eta}}[\xi] = \int_\Omega \xi\,w_{h+\varepsilon\eta}\,\mathrm{d}\nu_0,\] so 13 implies \[\begin{align} \frac{\mathbb{E}_{\rho_{h+\varepsilon\eta}}[\xi]-\mathbb{E}_{\rho_h}[\xi]}{\varepsilon} & \to \int_\Omega \xi\,w_h a_h(\eta)\,\mathrm{d}\nu_0 \\ & = \int_\Omega \xi\,a_h(\eta)\,\mathrm{d}\rho_h = \operatorname{Cov}_{\rho_h}(\eta,\xi), \end{align}\] which is 14 . ◻
Proposition 14 (Linearization of the weighted Neumann solve). Assume Assumptions 1, 2, and 3. Let \(h,\eta\in \mathcal{X}\) and \(\xi\in L^2_0(\nu_0)\), and assume that there exists \(\varepsilon_0>0\) such that \[h+\varepsilon \eta \in \mathcal{X}_{\mathrm{ad}} \qquad\text{for all } |\varepsilon|<\varepsilon_0.\] Then there exists a unique \[\chi_{h;\eta,\xi}\in H^1_{\rho_h,\diamond}(\Omega)\] such that \[\begin{align} \label{eq:linearized-neumann} \int_\Omega \nabla \chi_{h;\eta,\xi}\cdot \nabla \varphi\, \mathrm{d}\rho_h & = \int_\Omega \Big( a_h(\eta)\,q_h(\xi) - \operatorname{Cov}_{\rho_h}(\eta,\xi) \Big)\varphi\, \mathrm{d}\rho_h \notag \\ & \quad - \int_\Omega a_h(\eta) \nabla \psi_{h,\xi}\cdot \nabla \varphi\, \mathrm{d}\rho_h \qquad \forall \varphi\in H^1(\Omega;\rho_h), \end{align}\qquad{(9)}\] where \[a_h(\eta):=\eta-\mathbb{E}_{\rho_h}[\eta], \qquad q_h(\xi):=\xi-\mathbb{E}_{\rho_h}[\xi].\] Moreover, the gradient map is differentiable at \(\varepsilon=0\): \[\frac{\nabla\psi_{h+\varepsilon\eta,\xi}-\nabla\psi_{h,\xi}}{\varepsilon} \longrightarrow \nabla\chi_{h;\eta,\xi} \qquad\text{in } L^2(\nu_0;\mathbb{R}^d).\] Equivalently, \(\varepsilon\mapsto\psi_{h+\varepsilon\eta,\xi}\) is differentiable at \(\varepsilon=0\) in \(H^1(\Omega;\nu_0)/\mathbb{R}\).
Proof. The right-hand side of ?? is continuous on \(H^1(\Omega;\rho_h)\), so existence and uniqueness of \(\chi_{h;\eta,\xi}\) follow from Lax–Milgram exactly as in Theorem 9.
Let \[w_\varepsilon:=w_{h+\varepsilon\eta}, \qquad q_\varepsilon:=\xi-\mathbb{E}_{\rho_{h+\varepsilon\eta}}[\xi], \qquad \psi_\varepsilon:=\psi_{h+\varepsilon\eta,\xi}.\] Define the difference quotient \[\delta_\varepsilon := \frac{\psi_\varepsilon-\psi_{h,\xi}}{\varepsilon}.\] Subtract the weak formulations for \(\psi_\varepsilon\) and \(\psi_{h,\xi}\), divide by \(\varepsilon\), and compare with the weak equation for \(\chi_{h;\eta,\xi}\). Using Lemma 2, 14 , the boundedness of the weights, and the energy bound 9 , one obtains an error equation for \[r_\varepsilon:=\delta_\varepsilon-\chi_{h;\eta,\xi}\] of the form \[\int_\Omega \nabla r_\varepsilon\cdot\nabla\varphi\,\mathrm{d}\rho_{h+\varepsilon\eta} = R_\varepsilon(\varphi), \qquad \forall \varphi\in H^1(\Omega;\rho_{h+\varepsilon\eta}),\] where \(R_\varepsilon(\varphi)\to 0\) uniformly for \(\|\varphi\|_{H^1(\Omega;\rho_{h+\varepsilon\eta})}\le 1\). Taking \(\varphi=r_\varepsilon-\mathbb{E}_{\rho_{h+\varepsilon\eta}}[r_\varepsilon]\), applying Proposition 8, and using the equivalence of the \(\rho_{h+\varepsilon\eta}\)- and \(\nu_0\)-norms, we conclude that \[\|\nabla r_\varepsilon\|_{L^2(\nu_0)}\to 0.\] Subtracting an irrelevant constant and using 2 then yields \[\inf_{a\in\mathbb{R}} \|r_\varepsilon-a\|_{H^1(\Omega;\nu_0)} \to 0.\] In particular, \(\nabla r_\varepsilon\to 0\) in \(L^2(\nu_0;\mathbb{R}^d)\), which proves the gradient differentiability claim and the equivalent differentiability statement in \(H^1(\Omega;\nu_0)/\mathbb{R}\). ◻
Corollary 3 (Directional differentiability of the transport form). Under the assumptions of Proposition 14, the map \[\varepsilon\longmapsto \mathfrak g_{h+\varepsilon\eta}(\xi,\zeta)\] is differentiable at \(\varepsilon=0\) for every \(\xi,\zeta\in L^2_0(\nu_0)\), with derivative \[\begin{align} \label{eq:directional-derivative-gh} D_h \mathfrak g_h[\eta](\xi,\zeta) & = \int_\Omega a_h(\eta)\, \nabla\psi_{h,\xi}\cdot \nabla\psi_{h,\zeta}\, \mathrm{d}\rho_h \notag \\ & \quad + \int_\Omega \nabla\chi_{h;\eta,\xi}\cdot \nabla\psi_{h,\zeta}\, \mathrm{d}\rho_h + \int_\Omega \nabla\psi_{h,\xi}\cdot \nabla\chi_{h;\eta,\zeta}\, \mathrm{d}\rho_h. \end{align}\tag{15}\]
Proof. By definition, \[\mathfrak g_h(\xi,\zeta) = \int_\Omega \nabla\psi_{h,\xi}\cdot \nabla\psi_{h,\zeta}\, \mathrm{d}\rho_h.\] Differentiate this identity with respect to \(h\) in the direction \(\eta\). The derivative of \(\mathrm{d}\rho_h=w_h\,\mathrm{d}\nu_0\) is given by Lemma 2, while the derivatives of \(\psi_{h,\xi}\) and \(\psi_{h,\zeta}\) are given by Proposition 14. Passing to the limit in the resulting difference quotient yields 15 . ◻
Remark 15 (Pullback interpretation). The bilinear form \(\mathfrak g_h\) may be viewed as a pullback of continuity-equation transport geometry to Bayes Hilbert coordinates. A tangent direction \(\xi\) first produces the signed density variation \[\rho_h\bigl(\xi-\mathbb{E}_{\rho_h}[\xi]\bigr),\] and the weighted Neumann problem then selects the unique minimum-energy velocity field realizing that variation. The form \(\mathfrak g_h\) measures the kinetic energy of this realization.
We now pass from single tangent directions to time-dependent Bayes Hilbert paths.
Definition 8 (Regular admissible path). A path \[h:[0,T]\to \mathcal{X}_{\mathrm{ad}}\] is called regular admissible if \[h\in C^1([0,T];\mathcal{Z}).\] For such a path we define \[\rho_t := \rho_{h(t)}, \qquad v_t := \mathcal{T}_{h(t)}\dot{h}(t).\]
Theorem 16 (Canonical dynamical realization). Let \(h:[0,T]\to \mathcal{X}_{\mathrm{ad}}\) be a regular admissible path, and define \[\rho_t := \rho_{h(t)}, \qquad v_t := \mathcal{T}_{h(t)}\dot{h}(t).\] Then \((\rho_t,v_t)\) satisfies the continuity equation \[\label{eq:ambient-continuity-equation} \partial_t \rho_t + \nabla\cdot(\rho_t v_t)=0\tag{16}\] in the weak sense on \((0,T)\times\Omega\). Equivalently, for every \(\eta\in H^1(\Omega;\rho_t)\), \[\label{eq:ambient-continuity-weak} \frac{\mathrm{d}}{\mathrm{d}t}\int_\Omega \eta\, \mathrm{d}\rho_t = \int_\Omega \nabla \eta\cdot v_t\, \mathrm{d}\rho_t.\tag{17}\] If \(\Omega\) has boundary, the weak formulation encodes the natural zero-flux condition.
Proof. Since \(h\in C^1([0,T];\mathcal{Z})\) and \(\mathcal{Z}\hookrightarrow L^\infty(\nu_0)\cap L^2_0(\nu_0)\) continuously, \(h\) is a regular coordinate path in the sense of Section 2. By Proposition 6 from Section 2, \[\partial_t \frac{\mathrm{d}\rho_t}{\mathrm{d}\nu_0} = \frac{\mathrm{d}\rho_t}{\mathrm{d}\nu_0} \Bigl(\dot{h}(t)-\mathbb{E}_{\rho_t}[\dot{h}(t)]\Bigr).\] On the other hand, by definition of \(v_t\) and Remark 10, \[\int_\Omega \nabla\eta\cdot v_t\,\mathrm{d}\rho_t = \int_\Omega \Bigl(\dot{h}(t)-\mathbb{E}_{\rho_t}[\dot{h}(t)]\Bigr)\eta\,\mathrm{d}\rho_t\] for every \(\eta\in H^1(\Omega;\rho_t)\). Therefore \[\begin{align} \frac{\mathrm{d}}{\mathrm{d}t}\int_\Omega \eta\,\mathrm{d}\rho_t & = \int_\Omega \eta\, \partial_t\!\left(\frac{\mathrm{d}\rho_t}{\mathrm{d}\nu_0}\right)\,\mathrm{d}\nu_0 \\ & = \int_\Omega \eta \Bigl(\dot{h}(t)-\mathbb{E}_{\rho_t}[\dot{h}(t)]\Bigr)\,\mathrm{d}\rho_t \\ & = \int_\Omega \nabla\eta\cdot v_t\,\mathrm{d}\rho_t, \end{align}\] which is 17 . ◻
We conclude by recording the bounded-domain specialization used later in the paper.
Corollary 4 (Bounded Lipschitz domains). Let \(\Omega\subset\mathbb{R}^d\) be bounded, connected, and Lipschitz, and let \[\nu_0=|\Omega|^{-1}\,\mathrm{d}x.\] Fix exponents \[s>d, \qquad \frac{d}{2}<s'<\frac{s}{2},\] and define \[\mathcal{X}:=H^s(\Omega)\cap L^2_0(\nu_0), \qquad \mathcal{Z}:=H^{s'}(\Omega)\cap L^2_0(\nu_0).\] Then \(\mathcal{Z}\hookrightarrow L^\infty(\Omega)\cap L^2(\nu_0)\) continuously, and \(\nu_0\) satisfies the Poincaré inequality 2 . Hence any admissible class \(\mathcal{X}_{\mathrm{ad}}\subset\mathcal{Z}\) satisfying 3 falls under the abstract theory above. In this case, Theorem 9 is the weak weighted Neumann problem with natural zero-flux boundary condition.
If, in addition, \(\mathcal{X}_{\mathrm{ad}}\) is locally stable under \(\mathcal{X}\)-perturbations in the sense that for every \(h,\eta\in\mathcal{X}\) with \(h\in\mathcal{X}_{\mathrm{ad}}\) there exists \(\varepsilon_0>0\) such that \[h+\varepsilon\eta\in\mathcal{X}_{\mathrm{ad}} \qquad \text{for all }|\varepsilon|<\varepsilon_0,\] then Assumption 3 is satisfied with the above choice of \(\mathcal{X}\), and therefore the linearization results (Proposition 14 and Corollary 3) also apply in this bounded-domain setting. Moreover, Corollary 2 becomes an \(H^{s'}\)-continuity statement for the transport map.
Proof. The Sobolev embedding \(H^{s'}(\Omega)\hookrightarrow L^\infty(\Omega)\) holds because \(s'>d/2\), and the Poincaré inequality on bounded connected Lipschitz domains is classical. Since \(s>s'\), we also have the continuous embedding \(\mathcal{X}\hookrightarrow \mathcal{Z}\). Thus Assumption 1 holds with the chosen \(\mathcal{Z}\). Together with Assumption 2, this yields the abstract forward theory in the bounded-domain setting. The final statement is immediate from the local stability hypothesis. ◻
In Section 3, we associated to each regular Bayes Hilbert path \[h:[0,T]\to \mathcal{X}_{\mathrm{ad}}\] a canonical velocity field obtained from the weighted Neumann problem, together with the induced transport form \[\mathfrak g_h(\xi,\zeta).\] We now turn to the inverse problem. Rather than assuming that the path \(h(\cdot)\) is known, we ask how to reconstruct it from indirect time-dependent observations.
In many settings, instantaneous observations do not determine the state: for a prescribed mobile sensor path \(y\), the instantaneous linearized observation operator \[J_{t,h(t)}^{y}\] may have a nontrivial kernel at each time \(t\). Reconstruction then requires combining information from the entire observation record with dynamical structure linking different times. The central analytic object capturing this interplay is the joint transport–observability form, which couples the instantaneous observability form induced by the mobile sensors with the transport form induced by the forward geometry. Coercivity of this joint form is the natural replacement for pointwise observability, and it makes the pair \[(\mathfrak g_h,\mathfrak j_{t,h}^{y})\] into the engine of identifiability rather than merely an interpretive device.
We retain the setting of Section 3. Thus \(\Omega\subset\mathbb{R}^d\) is bounded, connected, and Lipschitz, \(\nu_0=|\Omega|^{-1}\mathrm{d}x\), the Sobolev exponents \[s>d, \qquad \frac{d}{2}<s'<\frac{s}{2}\] are fixed, and \[\mathcal{X} = H^s(\Omega)\cap L^2_0(\nu_0), \qquad \mathcal{Z} = H^{s'}(\Omega)\cap L^2_0(\nu_0), \qquad \mathcal{X}_{\mathrm{ad}}\subset \mathcal{X}\] is the admissible state class from Assumption 2.
For the inverse problem we impose one additional structural assumption on \(\mathcal{X}_{\mathrm{ad}}\).
Assumption 4 (Sobolev-regular admissible state class). Assume that \(\mathcal{X}_{\mathrm{ad}}\) is closed in \(\mathcal{Z}\).
We work throughout this section with a prescribed mobile-sensor observation model. Let \[\kappa\in L^\infty(\mathbb{R}^d)\cap C_c(\mathbb{R}^d)\] be a compactly supported sensor kernel, and let \[y=(y_1,\dots,y_r), \qquad y_j:[0,T]\to \Omega,\] be measurable sensor trajectories. For each \(t\in[0,T]\), define the instantaneous observation map \[\label{eq:mobile-sensor-observation} \mathcal{G}_t^{y}(h) := \left( \int_\Omega \kappa(x-y_1(t))\,\mathrm{d}\rho_h(x),\dots, \int_\Omega \kappa(x-y_r(t))\,\mathrm{d}\rho_h(x) \right)\in \mathbb{R}^r.\tag{18}\] Thus the observation operator itself depends on time through the sensor locations.
If \(h^\dagger(\cdot)\) denotes the unknown true path, then the ideal data are \[\mathsf{d}^\dagger(t)=\mathcal{G}_t^y(h^\dagger(t)),\] and the measured data are modeled as \[\mathsf{d}(t)=\mathsf{d}^\dagger(t)+\eta(t),\] where \(\eta\) is an observation error term.
Definition 9 (Observability differential). For \(t\in[0,T]\) and \(h\in\mathcal{X}_{\mathrm{ad}}\), assume \(\mathcal{G}_t^y\) is Fréchet differentiable at \(h\). The corresponding observability differential is \[J_{t,h}^{y}:=D\mathcal{G}_t^y(h):H^{s'}(\Omega)\to\mathbb{R}^r.\]
Definition 10 (Instantaneous observability form). For \(t\in[0,T]\) and \(h\in\mathcal{X}_{\mathrm{ad}}\), the associated instantaneous observability form is \[\mathfrak j_{t,h}^{y}(\xi,\zeta) := \langle J_{t,h}^{y}\xi,\;J_{t,h}^{y}\zeta\rangle_{\mathbb{R}^r}, \qquad \xi,\zeta\in H^{s'}(\Omega)\cap L^2_0(\nu_0).\]
Proposition 17 (Observability differential for mobile sensor averages). For each \(t\in[0,T]\), the map \[\mathcal{G}_t^{y}:\mathcal{X}_{\mathrm{ad}}\to\mathbb{R}^r\] defined by 18 is Fréchet differentiable with respect to the \(H^{s'}(\Omega)\)-topology, and for every \(h\in\mathcal{X}_{\mathrm{ad}}\) and \(\xi\in H^{s'}(\Omega)\cap L^2_0(\nu_0)\), \[\label{eq:mobile-sensor-differential-formula} J_{t,h}^{y}\xi = \Bigl( \operatorname{Cov}_{\rho_h}(\kappa(\cdot-y_1(t)),\xi),\dots, \operatorname{Cov}_{\rho_h}(\kappa(\cdot-y_r(t)),\xi) \Bigr).\qquad{(10)}\] Consequently, \[\label{eq:mobile-sensor-observability-form} \mathfrak j_{t,h}^{y}(\xi,\zeta) = \sum_{j=1}^r \operatorname{Cov}_{\rho_h}(\kappa(\cdot-y_j(t)),\xi)\, \operatorname{Cov}_{\rho_h}(\kappa(\cdot-y_j(t)),\zeta),\qquad{(11)}\] and in particular \[\label{eq:mobile-sensor-observability-energy} \mathfrak j_{t,h}^{y}(\xi,\xi) = \sum_{j=1}^r \operatorname{Cov}_{\rho_h}(\kappa(\cdot-y_j(t)),\xi)^2.\qquad{(12)}\]
Proof. For each \(j=1,\dots,r\), define \[\mathcal{G}_{t,j}^{y}(h):=\int_\Omega \kappa(x-y_j(t))\,\mathrm{d}\rho_h(x).\] Using the differential of the exponential-normalization map from Proposition 4, we have \[\begin{align} D\mathcal{G}_{t,j}^{y}(h)[\xi] & = \int_\Omega \kappa(x-y_j(t)) \bigl(\xi(x)-\mathbb{E}_{\rho_h}[\xi]\bigr)\,\mathrm{d}\rho_h(x) \\ & = \operatorname{Cov}_{\rho_h}(\kappa(\cdot-y_j(t)),\xi). \end{align}\] This proves ?? . The formulas ?? and ?? follow immediately from Definition 10. ◻
Remark 18 (Interpretation). For the mobile-sensor model 18 , a tangent direction \(\xi\) is seen through the data only via its covariance with the translated sensor kernels \(\kappa(\cdot-y_j(t))\) under the current law \(\rho_h\). Thus the visible directions vary in time both because the state \(h\) evolves and because the sensors move.
Proposition 19 (Instantaneous ill-posedness of mobile sensor observations). Fix \(t\in[0,T]\). For each \(h\in\mathcal{X}_{\mathrm{ad}}\), the differential \(J_{t,h}^{y}\) extends to a bounded linear operator \[J_{t,h}^{y}\in\mathcal{L}(L^2_0(\nu_0),\mathbb{R}^r),\] and \[\operatorname{rank}J_{t,h}^{y}\le r.\] In particular, \(J_{t,h}^{y}\) is finite-rank and therefore compact.
Consequently, the instantaneous invisible subspace \[\ker J_{t,h}^{y}\subset L^2_0(\nu_0)\] is closed and has codimension at most \(r\). If \(L^2_0(\nu_0)\) is infinite dimensional, then \(\ker J_{t,h}^{y}\) is infinite dimensional.
Proof. This is immediate from ?? . ◻
The form \(\mathfrak j_{t,h}^{y}\) is symmetric and nonnegative by construction. In the partially observable setting, \(\mathfrak j_{t,h}^{y}(\xi,\xi)\) may vanish for nonzero \(\xi\); equivalently, the instantaneous observation differential \(J_{t,h}^{y}\) may have a nontrivial kernel even when the sensor path is prescribed.
We reconstruct paths from the admissible class \[\mathcal{A}_{\mathrm{ad}} := \left\{ h\in L^2(0,T;\mathcal{X})\cap H^1(0,T;L^2_0(\nu_0)) : h(t)\in \mathcal{X}_{\mathrm{ad}} \text{ for a.e.\;} t\in[0,T] \right\}.\]
Remark 20 (Observation-side ill-posedness). Proposition 19 shows that the instantaneous linearized inverse problem for the mobile-sensor model is severely underdetermined in the ambient infinite-dimensional setting: at each fixed time \(t\), the data constrain at most \(r\) directions, while an infinite-dimensional family of tangent perturbations remains invisible. Thus identifiability cannot be expected from a single-time observation alone. The ambient inverse theory below addresses this by coupling instantaneous observability with the transport geometry through the joint form.
We now formulate the central observability condition for the partially observable inverse problem. The key idea is that directions invisible to the observation operator at a single time may become visible when integrated over the full time horizon, provided the transport geometry couples different times through the time derivative.
Definition 11 (Pathwise observability Gramian). For \(h\in \mathcal{A}_{\mathrm{ad}}\), the pathwise observability Gramian is the bilinear form \[\mathfrak J_y[h](\xi,\zeta) := \int_0^T \mathfrak j_{t,h(t)}^{y}\bigl(\xi(t),\zeta(t)\bigr)\,\mathrm{d}t = \int_0^T \langle J_{t,h(t)}^{y}\xi(t),\;J_{t,h(t)}^{y}\zeta(t)\rangle_{\mathbb{R}^r}\,\mathrm{d}t\] defined on path perturbations \[\xi,\zeta\in L^2(0,T;H^{s'}(\Omega)\cap L^2_0(\nu_0)).\]
In the fully observable case—when the instantaneous form \(\mathfrak j_{t,h}^{y}\) is coercive at each time—the pathwise Gramian \(\mathfrak J_y[h]\) is already coercive on its own. In the partially observable setting, however, \(\mathfrak J_y[h]\) may degenerate: there may exist nonzero path perturbations \(\xi(\cdot)\) with \[\xi(t)\in\ker J_{t,h(t)}^{y} \qquad\text{for a.e.\;}t,\] so that \(\mathfrak J_y[h](\xi,\xi)=0\). Reconstruction then requires additional structure linking different times.
This additional structure is provided by the transport form. The following definition combines the two ambient geometric objects into a single bilinear form on path perturbations.
Definition 12 (Joint transport–observability form). For \(h\in \mathcal{A}_{\mathrm{ad}}\), the joint transport–observability form is the bilinear form \[\label{eq:joint-form} Q_y[h](\xi,\zeta) := \int_0^T \mathfrak j_{t,h(t)}^{y}\bigl(\xi(t),\zeta(t)\bigr)\,\mathrm{d}t + \int_0^T \mathfrak g_{h(t)}\bigl(\dot{\xi}(t),\dot{\zeta}(t)\bigr)\,\mathrm{d}t\tag{19}\] defined on path perturbations \[\xi,\zeta \in L^2(0,T;H^{s'}(\Omega)\cap L^2_0(\nu_0)) \cap H^1(0,T;L^2_0(\nu_0)).\] Equivalently, \[Q_y[h](\xi,\zeta) = \mathfrak J_y[h](\xi,\zeta) + \int_0^T \mathfrak g_{h(t)}\bigl(\dot{\xi}(t),\dot{\zeta}(t)\bigr)\,\mathrm{d}t.\]
The form \(Q_y[h]\) encodes the interplay between the two ambient geometric objects \[(\mathfrak g_h,\mathfrak j_{t,h}^{y})\] at the level of path perturbations. The first term measures how strongly tangent directions are seen through the data; the second measures the dynamical cost of their time variation. Together, they quantify how much information about the perturbation \(\xi(\cdot)\) is available from the combination of observations and dynamical structure.
Remark 21 (Mechanism of partial observability). Suppose at each time \(t\), the kernel of \(J_{t,h(t)}^{y}\) is nontrivial but rotates as the state evolves and the sensors move. A nonzero perturbation \(\xi(\cdot)\) that tracks the kernel—staying in \(\ker J_{t,h(t)}^{y}\) for all \(t\)—must have a nontrivial time derivative \(\dot{\xi}\). The transport term \[\int_0^T\mathfrak g_{h(t)}(\dot{\xi},\dot{\xi})\,\mathrm{d}t\] then detects this variation. Conversely, a perturbation with \(\dot{\xi}=0\) is constant in time, so it can only hide from the data if it lies in \(\ker J_{t,h(t)}^{y}\) for all \(t\) simultaneously. Joint coercivity of \(Q_y[h]\) rules out both scenarios: perturbations must be visible either through the observations or through their dynamical cost.
Assumption 5 (Joint transport–observability coercivity). There exists \(\kappa>0\) such that for all \(h\in \mathcal{A}_{\mathrm{ad}}\) and all \[\xi\in L^2(0,T;H^{s'}(\Omega)\cap L^2_0(\nu_0)) \cap H^1(0,T;L^2_0(\nu_0)),\] we have \[\label{eq:joint-coercivity} Q_y[h](\xi,\xi) \ge \kappa \|\xi\|_{L^2(0,T;L^2(\nu_0))}^2.\tag{20}\]
Remark 22 (Comparison with pointwise observability). If the instantaneous observability form alone is coercive—that is, if there exists \(\kappa_0>0\) such that \[\mathfrak j_{t,h}^{y}(\xi,\xi)\ge \kappa_0 \|\xi\|_{L^2(\nu_0)}^2 \qquad \text{for all } t\in[0,T],\;h\in\mathcal{X}_{\mathrm{ad}},\; \xi\in H^{s'}(\Omega)\cap L^2_0(\nu_0),\] then Assumption 5 is automatically satisfied with \(\kappa=\kappa_0\), since the transport term is nonnegative. In this case, the data at each time already determine the state, and no dynamical coupling is needed. Conversely, when \(\ker J_{t,h(t)}^{y}\neq\{0\}\), Assumption 5 is strictly weaker than pointwise coercivity, and the transport term is essential: it compensates for directions that are invisible instantaneously by detecting their temporal variation. The \(\mu\)-regularization introduced in the variational formulation below provides analytic control over this transport term.
Proposition 23 (Upper bound on the transport term). There exists a constant \(C_{\mathrm{tr}}>0\), depending only on the admissible class \(\mathcal{X}_{\mathrm{ad}}\) and the domain \(\Omega\), such that for all \(h\in\mathcal{A}_{\mathrm{ad}}\) and all \[\xi\in H^1(0,T;L^2_0(\nu_0)),\] we have \[\int_0^T \mathfrak g_{h(t)}(\dot{\xi}(t),\dot{\xi}(t))\,\mathrm{d}t \le C_{\mathrm{tr}}\|\dot{\xi}\|_{L^2(0,T;L^2(\nu_0))}^2.\]
Proof. By definition of the transport form and the weighted Neumann solve, \[\mathfrak g_h(\zeta,\zeta) = \int_\Omega |\nabla\psi_{h,\zeta}|^2\,\mathrm{d}\rho_h.\] By the energy estimate from Theorem 9, together with the uniform density bounds on \(\mathcal{X}_{\mathrm{ad}}\), there exists \(C_{\mathrm{tr}}>0\) such that \[\mathfrak g_h(\zeta,\zeta) \le C_{\mathrm{tr}}\|\zeta\|_{L^2(\nu_0)}^2 \qquad \text{for all } h\in\mathcal{X}_{\mathrm{ad}},\;\zeta\in L^2_0(\nu_0).\] Applying this pointwise in time with \(\zeta=\dot{\xi}(t)\) and integrating over \([0,T]\) gives the claim. ◻
Remark 24 (Interpretation of the bounds). Assumption 5 gives a lower bound showing that the joint form \(Q_y[h]\) controls the \(L^2(0,T;L^2(\nu_0))\)-size of a path perturbation. On the other hand, Proposition 23 shows that the transport contribution to \(Q_y[h]\) is controlled by the natural \(H^1\)-in-time regularity of the perturbation.
For the prescribed mobile-sensor model, the observation differential is uniformly bounded on \(\mathcal{X}_{\mathrm{ad}}\): there exists \(C_{\mathrm{obs}}>0\), depending only on \(\kappa\), \(r\), \(\Omega\), and \(\mathcal{X}_{\mathrm{ad}}\), such that \[\|J_{t,h}^{y}\xi\|_{\mathbb{R}^r} \le C_{\mathrm{obs}}\|\xi\|_{L^2(\nu_0)} \qquad \text{for all } t\in[0,T],\;h\in\mathcal{X}_{\mathrm{ad}},\;\xi\in L^2_0(\nu_0).\] Hence \[Q_y[h](\xi,\xi) \lesssim \|\xi\|_{L^2(0,T;L^2(\nu_0))}^2 + \|\dot{\xi}\|_{L^2(0,T;L^2(\nu_0))}^2.\] Thus \(Q_y[h]\) should be viewed as a pathwise energy that is coercive in the \(L^2\)-state variable and compatible with the natural regularity of the admissible path space.
Proposition 25 (A sufficient criterion for joint transport–observability coercivity). Assume that for each \(h\in \mathcal{A}_{\mathrm{ad}}\) and a.e.\(t\in[0,T]\), the observability differential \(J_{t,h(t)}^{y}\) extends to a bounded operator \[J_{t,h(t)}^{y}\in \mathcal{L}\bigl(L^2_0(\nu_0),\mathbb{R}^r\bigr).\] Define the instantaneous invisible subspace \[K_{y,h}(t):=\ker J_{t,h(t)}^{y}\subset L^2_0(\nu_0),\] and let \[\Pi_{y,h}(t):L^2_0(\nu_0)\to K_{y,h}(t)\] denote the \(L^2(\nu_0)\)-orthogonal projection onto \(K_{y,h}(t)\). Assume moreover that for each \(h\in \mathcal{A}_{\mathrm{ad}}\), the map \[t\longmapsto \Pi_{y,h}(t)\] is strongly measurable as an \(\mathcal{L}(L^2_0(\nu_0),L^2_0(\nu_0))\)-valued map.
Suppose there exist constants \(c_{\mathrm{obs}},c_{\mathrm{dyn}}>0\) such that for every \(h\in \mathcal{A}_{\mathrm{ad}}\) the following hold:
(uniform transverse observability) for a.e.\(t\in[0,T]\) and every \(\zeta\in L^2_0(\nu_0)\), \[\label{eq:transverse-observability} \mathfrak j_{t,h(t)}^{y}(\zeta,\zeta) = \|J_{t,h(t)}^{y}\zeta\|_{\mathbb{R}^r}^2 \ge c_{\mathrm{obs}} \bigl\|(I-\Pi_{y,h}(t))\zeta\bigr\|_{L^2(\nu_0)}^2;\qquad{(13)}\]
(uniform dynamical detectability of invisible directions) for every \[\xi\in L^2(0,T;H^{s'}(\Omega)\cap L^2_0(\nu_0)) \cap H^1(0,T;L^2_0(\nu_0)),\] one has \[\label{eq:dynamical-detectability} \int_0^T \mathfrak g_{h(t)}(\dot{\xi}(t),\dot{\xi}(t))\,\mathrm{d}t \ge c_{\mathrm{dyn}} \int_0^T \bigl\|\Pi_{y,h}(t)\xi(t)\bigr\|_{L^2(\nu_0)}^2\,\mathrm{d}t.\qquad{(14)}\]
Then Assumption 5 holds with \[\kappa=\min\{c_{\mathrm{obs}},c_{\mathrm{dyn}}\}.\]
Proof. Fix \(h\in \mathcal{A}_{\mathrm{ad}}\), and let \[\xi\in L^2(0,T;H^{s'}(\Omega)\cap L^2_0(\nu_0)) \cap H^1(0,T;L^2_0(\nu_0)).\] Decompose \(\xi\) pointwise in time into its visible and invisible parts: \[\xi(t)=\xi_\perp(t)+\xi_0(t), \qquad \xi_\perp(t):=(I-\Pi_{y,h}(t))\xi(t), \qquad \xi_0(t):=\Pi_{y,h}(t)\xi(t).\] Since \(\Pi_{y,h}(t)\) is the \(L^2(\nu_0)\)-orthogonal projection onto \(K_{y,h}(t)\), we have \[\xi_\perp(t)\perp \xi_0(t) \qquad\text{in }L^2(\nu_0)\] for a.e.\(t\in[0,T]\), and therefore \[\label{eq:orthogonal-splitting} \|\xi\|_{L^2(0,T;L^2(\nu_0))}^2 = \int_0^T \|\xi_\perp(t)\|_{L^2(\nu_0)}^2\,\mathrm{d}t + \int_0^T \|\xi_0(t)\|_{L^2(\nu_0)}^2\,\mathrm{d}t.\tag{21}\]
Now apply the two hypotheses. By ?? , \[\int_0^T \mathfrak j_{t,h(t)}^{y}(\xi(t),\xi(t))\,\mathrm{d}t \ge c_{\mathrm{obs}} \int_0^T \|\xi_\perp(t)\|_{L^2(\nu_0)}^2\,\mathrm{d}t.\] By ?? , \[\int_0^T \mathfrak g_{h(t)}(\dot{\xi}(t),\dot{\xi}(t))\,\mathrm{d}t \ge c_{\mathrm{dyn}} \int_0^T \|\xi_0(t)\|_{L^2(\nu_0)}^2\,\mathrm{d}t.\] Adding the two inequalities yields \[\begin{align} Q_y[h](\xi,\xi) & = \int_0^T \mathfrak j_{t,h(t)}^{y}(\xi(t),\xi(t))\,\mathrm{d}t + \int_0^T \mathfrak g_{h(t)}(\dot{\xi}(t),\dot{\xi}(t))\,\mathrm{d}t \\ & \ge c_{\mathrm{obs}} \int_0^T \|\xi_\perp(t)\|_{L^2(\nu_0)}^2\,\mathrm{d}t + c_{\mathrm{dyn}} \int_0^T \|\xi_0(t)\|_{L^2(\nu_0)}^2\,\mathrm{d}t \\ & \ge \min\{c_{\mathrm{obs}},c_{\mathrm{dyn}}\} \left( \int_0^T \|\xi_\perp(t)\|_{L^2(\nu_0)}^2\,\mathrm{d}t + \int_0^T \|\xi_0(t)\|_{L^2(\nu_0)}^2\,\mathrm{d}t \right). \end{align}\] Using 21 , we conclude that \[Q_y[h](\xi,\xi) \ge \min\{c_{\mathrm{obs}},c_{\mathrm{dyn}}\} \|\xi\|_{L^2(0,T;L^2(\nu_0))}^2,\] which is exactly Assumption 5. ◻
Remark 26 (Interpretation of the criterion). Proposition 25 formalizes the heuristic in Remark 21. The component \((I-\Pi_{y,h}(t))\xi(t)\) is instantaneously visible and is controlled directly by the data through \(\mathfrak j_{t,h(t)}^{y}\). The component \(\Pi_{y,h}(t)\xi(t)\) lies in the instantaneous kernel of the observation operator and is therefore invisible at any fixed time; condition ?? requires that such directions be detected through the dynamical cost of their time variation measured by \(\mathfrak g_h\).
The natural data-misfit term for the prescribed mobile-sensor system is \[\frac{1}{2}\int_0^T \|\mathcal{G}_t^y(h(t))-\mathsf{d}(t)\|_{\mathbb{R}^r}^2\,\mathrm{d}t,\] and the natural dynamical penalty coming from the forward theory is the transport action \[\frac{1}{2}\int_0^T \mathfrak g_{h(t)}(\dot{h}(t),\dot{h}(t))\,\mathrm{d}t.\] In the ambient infinite-dimensional setting, however, the transport action alone does not provide sufficient compactness for the direct-method existence proof. For that reason, we add both a Bayes Hilbert \(H^1\)-in-time regularization term and a spatial Sobolev regularization term.
Definition 13 (Regularized inverse functional). Let \(\lambda,\mu,\gamma>0\). For \(h\in\mathcal{A}_{\mathrm{ad}}\), define \[\begin{align} \label{eq:ambient-inverse-functional} \mathcal{I}_{\lambda,\mu,\gamma}^{\,y}[h] & := \frac{1}{2}\int_0^T \|\mathcal{G}_t^y(h(t))-\mathsf{d}(t)\|_{\mathbb{R}^r}^2\,\mathrm{d}t + \frac{\lambda}{2}\int_0^T \mathfrak g_{h(t)}(\dot{h}(t),\dot{h}(t))\,\mathrm{d}t \notag \\ & \quad + \frac{\mu}{2}\int_0^T \Big( \|h(t)\|_{L^2(\nu_0)}^2+\|\dot{h}(t)\|_{L^2(\nu_0)}^2 \Big)\,\mathrm{d}t + \frac{\gamma}{2}\int_0^T \|h(t)\|_{H^s(\Omega)}^2\,\mathrm{d}t. \end{align}\tag{22}\]
Remark 27 (Dual role of the regularization terms). Under full pointwise observability, the \(\mu\)- and \(\gamma\)-terms serve a purely analytic purpose: the former provides \(H^1\)-in-time compactness in the Bayes Hilbert norm, while the latter provides the spatial compactness needed for strong convergence in \(C([0,T];H^{s'}(\Omega))\).
Under partial observability, the \(\mu\)-term acquires a second, structural role. The joint coercivity of \(Q_y[h]\) (Assumption 5) shows that unobserved directions are detected through their dynamical cost in the transport form \(\mathfrak g_h\). But the transport term in \(Q_y[h]\) involves \(\dot{\xi}\), and it is the \(\mu\)-weighted \(\|\dot{h}\|_{L^2(\nu_0)}^2\) penalty that provides a priori control over this quantity for minimizers. In this way, \(\mu\) acts not only as a compactness device but as the mechanism by which the dynamical structure propagates information across time.
The existence proof uses compactness of bounded-energy sequences and lower semicontinuity of each term in the functional.
Proposition 28 (Lower semicontinuity of the transport action under strong state convergence). Let \(h_n,h\in \mathcal{A}_{\mathrm{ad}}\) satisfy \[h_n\to h \quad\text{in } C([0,T];H^{s'}(\Omega)),\] and \[\dot{h}_n \rightharpoonup \dot{h} \quad\text{weakly in } L^2(0,T;L^2_0(\nu_0)).\] Then \[\int_0^T \mathfrak g_{h(t)}(\dot{h}(t),\dot{h}(t))\,\mathrm{d}t \le \liminf_{n\to\infty} \int_0^T \mathfrak g_{h_n(t)}(\dot{h}_n(t),\dot{h}_n(t))\,\mathrm{d}t.\]
Proof. For each \(h\in\mathcal{X}_{\mathrm{ad}}\), define \[\mathcal{S}_h : L^2_0(\nu_0)\to L^2(\nu_0;\mathbb{R}^d), \qquad \mathcal{S}_h \xi := w_h^{1/2}\,\mathcal{T}_h\xi,\] where \(w_h:=\mathrm{d}\rho_h/\mathrm{d}\nu_0\). Then \[\mathfrak g_h(\xi,\xi) = \|\mathcal{S}_h \xi\|_{L^2(\nu_0;\mathbb{R}^d)}^2.\]
We claim that \[\sup_{t\in[0,T]} \|\mathcal{S}_{h_n(t)}-\mathcal{S}_{h(t)}\|_{\mathcal{L}(L^2_0,L^2)} \to 0.\] Indeed, writing \[\mathcal{S}_{h_n}-\mathcal{S}_h = w_{h_n}^{1/2}(\mathcal{T}_{h_n}-\mathcal{T}_h) + (w_{h_n}^{1/2}-w_h^{1/2})\mathcal{T}_h,\] the first term vanishes by Corollary 2 and the second by the Sobolev embedding \(H^{s'}(\Omega)\hookrightarrow L^\infty(\Omega)\) and the Lipschitz dependence of the square-root weight on the state.
Now define \[u_n(t):=\mathcal{S}_{h_n(t)}\dot{h}_n(t), \qquad u(t):=\mathcal{S}_{h(t)}\dot{h}(t).\] The uniform operator convergence of \(\mathcal{S}_{h_n}\) to \(\mathcal{S}_h\), together with the weak convergence \(\dot{h}_n\rightharpoonup \dot{h}\) in \(L^2(0,T;L^2_0(\nu_0))\), implies \[u_n\rightharpoonup u \quad\text{weakly in } L^2(0,T;L^2(\nu_0;\mathbb{R}^d)).\] By weak lower semicontinuity of the norm, \[\|u\|_{L^2_tL^2_x}^2 \le \liminf_{n\to\infty} \|u_n\|_{L^2_tL^2_x}^2,\] which is the desired inequality. ◻
Proposition 29 (Compactness of bounded-energy sequences). Let \((h_n)\subset \mathcal{A}_{\mathrm{ad}}\) be a sequence satisfying \[\sup_n \left( \|h_n\|_{L^2(0,T;H^s(\Omega))} + \|h_n\|_{H^1(0,T;L^2_0(\nu_0))} \right) <\infty.\] Then there exist a subsequence, again denoted \((h_n)\), and a limit \[h\in L^2(0,T;H^s(\Omega))\cap H^1(0,T;L^2_0(\nu_0))\] such that \[h_n \rightharpoonup h \quad\text{weakly in } L^2(0,T;H^s(\Omega)),\] \[h_n \rightharpoonup h \quad\text{weakly in } H^1(0,T;L^2_0(\nu_0)),\] and \[h_n\to h \quad\text{in } C([0,T];H^{s'}(\Omega)).\] If, in addition, \(h_n(t)\in \mathcal{X}_{\mathrm{ad}}\) for a.e.\(t\), then \(h(t)\in \mathcal{X}_{\mathrm{ad}}\) for all \(t\in[0,T]\).
Proof. The assumed bound gives boundedness in \[L^2(0,T;H^s(\Omega))\cap H^1(0,T;L^2_0(\nu_0)).\] By the Lions–Magenes interpolation theorem, applied to the Hilbert triple \(H^s(\Omega)\hookrightarrow H^{s/2}(\Omega)\hookrightarrow L^2(\Omega)\), this space embeds continuously into \[C([0,T];H^{s/2}(\Omega)).\] Thus \((h_n)\) is bounded in \(C([0,T];H^{s/2}(\Omega))\). Since \(s'<s/2\), the embedding \(H^{s/2}(\Omega)\hookrightarrow H^{s'}(\Omega)\) is compact.
The \(H^1(0,T;L^2_0(\nu_0))\) bound gives equicontinuity in \(L^2(\nu_0)\): for all \(t_0,t_1\in[0,T]\), \[\|h_n(t_1)-h_n(t_0)\|_{L^2(\nu_0)} \le |t_1-t_0|^{1/2} \|\dot{h}_n\|_{L^2(0,T;L^2(\nu_0))}.\] Let \(\theta=2s'/s\in(0,1)\). Since \(\nu_0=|\Omega|^{-1}\mathrm{d}x\), the \(L^2(\nu_0)\) and Lebesgue \(L^2\) norms are equivalent, and interpolation gives \[\|f\|_{H^{s'}(\Omega)} \le C\|f\|_{L^2(\nu_0)}^{1-\theta} \|f\|_{H^{s/2}(\Omega)}^\theta.\] Applying this to \(f=h_n(t_1)-h_n(t_0)\), and using the uniform \(C([0,T];H^{s/2}(\Omega))\) bound, gives equicontinuity in \(H^{s'}(\Omega)\).
The Arzela–Ascoli theorem now gives relative compactness in \[C([0,T];H^{s'}(\Omega)),\] because pointwise relative compactness follows from \(H^{s/2}(\Omega)\hookrightarrow H^{s'}(\Omega)\) compactly, and equicontinuity was established above. The weak convergences follow from Banach–Alaoglu.
Finally, take the continuous representatives. Since \(h_n(t)\in\mathcal{X}_{\mathrm{ad}}\) for a.e.\(t\), continuity and the \(\mathcal{Z}\)-closedness of \(\mathcal{X}_{\mathrm{ad}}\) imply \(h_n(t)\in\mathcal{X}_{\mathrm{ad}}\) for all \(t\in[0,T]\). Passing to the uniform \(H^{s'}\)-limit and using closedness again gives \(h(t)\in\mathcal{X}_{\mathrm{ad}}\) for all \(t\in[0,T]\). ◻
Theorem 30 (Existence of ambient reconstructions). Let \(d\in L^2(0,T;\mathbb{R}^r)\). Then for every \(\lambda,\mu,\gamma>0\), the functional \(\mathcal{I}_{\lambda,\mu,\gamma}^{\,y}\) admits a minimizer over \(\mathcal{A}_{\mathrm{ad}}\).
Proof. Let \((h_n)\subset \mathcal{A}_{\mathrm{ad}}\) be a minimizing sequence. Since the first two terms in 22 are nonnegative, \[\mathcal{I}_{\lambda,\mu,\gamma}^{\,y}[h_n] \ge \frac{\mu}{2}\int_0^T \Big( \|h_n(t)\|_{L^2(\nu_0)}^2+\|\dot{h}_n(t)\|_{L^2(\nu_0)}^2 \Big)\,\mathrm{d}t + \frac{\gamma}{2}\int_0^T \|h_n(t)\|_{H^s(\Omega)}^2\,\mathrm{d}t.\] Thus \((h_n)\) is bounded in \[L^2(0,T;H^s(\Omega))\cap H^1(0,T;L^2_0(\nu_0)).\] By Proposition 29, after passing to a subsequence we obtain \[h_n\to h \quad\text{in } C([0,T];H^{s'}(\Omega)),\] \[h_n\rightharpoonup h \quad\text{weakly in } L^2(0,T;H^s(\Omega)),\] and \[\dot{h}_n\rightharpoonup \dot{h} \quad\text{weakly in } L^2(0,T;L^2_0(\nu_0))\] for some \(h\in \mathcal{A}_{\mathrm{ad}}\).
For the data term, fix \(j\in\{1,\dots,r\}\). Since \(\kappa\in L^\infty\), \[\left| \int_\Omega \kappa(x-y_j(t))\,\mathrm{d}\rho_{h_n(t)}(x) \right| \le \|\kappa\|_{L^\infty(\mathbb{R}^d)}\] for all \(n\) and a.e.\(t\). Moreover, \(h_n(t)\to h(t)\) in \(H^{s'}(\Omega)\) for each \(t\), hence \[\mathcal{G}_t^y(h_n(t))\to \mathcal{G}_t^y(h(t)) \qquad\text{for a.e.\;}t\in[0,T].\] By dominated convergence, \[\mathcal{G}_t^y(h_n(t))\to \mathcal{G}_t^y(h(t)) \quad\text{in } L^2(0,T;\mathbb{R}^r),\] so the data term is continuous. Proposition 28 gives lower semicontinuity of the transport action. The final two regularization terms are weakly lower semicontinuous by convexity. Therefore \[\mathcal{I}_{\lambda,\mu,\gamma}^{\,y}[h] \le \liminf_{n\to\infty} \mathcal{I}_{\lambda,\mu,\gamma}^{\,y}[h_n].\] Thus \(h\) is a minimizer. ◻
Remark 31 (On the strong state topology). The compactness result in Proposition 29 produces strong convergence in \(C([0,T];H^{s'}(\Omega))\), which is the topology naturally matched to the continuity theory of the weighted Neumann solve from Section 3. Since \(H^{s'}(\Omega)\hookrightarrow L^\infty(\Omega)\) continuously, this is in particular strong enough for all continuity statements involving the weights \(w_h\) and the transport map \(\mathcal{T}_h\).
We now prove the main stability result: under joint transport–observability coercivity, minimizers of \(\mathcal{I}_{\lambda,\mu,\gamma}^{\,y}\) are stable with respect to data perturbations, even when instantaneous observations do not determine the state. The proof combines three ingredients: a priori bounds from minimality, the joint coercivity of \(Q_y[h]\), and the compactness theory of Proposition 29.
For the local stability theory, it is convenient to isolate the subclass \[\mathcal{A}_{\mathrm{ad}}^{s'} := \mathcal{A}_{\mathrm{ad}} \cap C([0,T];H^{s'}(\Omega)).\] All local proximity conditions in this subsection are understood on \(\mathcal{A}_{\mathrm{ad}}^{s'}\), so that expressions of the form \[\sup_{t\in[0,T]}\|h(t)-k(t)\|_{H^{s'}(\Omega)}\] are well defined.
Throughout this subsection, we make the following regularity assumption on the mobile observation family.
Assumption 6 (Regularity of the mobile observation family). Assume that there exists an open neighborhood \(U\subset \mathcal{Z}\) of \(\mathcal{X}_{\mathrm{ad}}\), together with constants \(r_0>0\) and \(L>0\), such that for each \(t\in[0,T]\), \[\mathcal{G}_t^y:U\to\mathbb{R}^r\] is \(C^1\) with respect to the \(H^{s'}(\Omega)\)-topology, and its differential \[J_{t,h}^{y}=D\mathcal{G}_t^y(h)\] is locally Lipschitz in \(h\), uniformly in \(t\), as a map into \(\mathcal{L}(L^2(\nu_0),\mathbb{R}^r)\): for all \(h,k\in U\) with \(\|h-k\|_{H^{s'}(\Omega)}\le r_0\), \[\label{eq:J-lipschitz} \|J_{t,h}^{y}-J_{t,k}^{y}\|_{\mathcal{L}(L^2(\nu_0),\mathbb{R}^r)} \le L\|h-k\|_{H^{s'}(\Omega)} \qquad\text{for all } t\in[0,T].\tag{23}\]
Lemma 3 (A priori bounds from minimality). Let \(h^\dagger\in\mathcal{A}_{\mathrm{ad}}\), and let \[\mathsf{d}(t)=\mathcal{G}_t^y(h^\dagger(t))+\eta(t)\] with \[\|\eta\|_{L^2(0,T;\mathbb{R}^r)}\le \delta.\] Define the regularization energy of the true path by \[\label{eq:regularization-energy-true-path} R[h^\dagger] := \frac{\lambda}{2}\int_0^T \mathfrak g_{h^\dagger(t)}(\dot{h}^\dagger(t),\dot{h}^\dagger(t))\,\mathrm{d}t + \frac{\mu}{2}\int_0^T \Big( \|h^\dagger(t)\|_{L^2(\nu_0)}^2+\|\dot{h}^\dagger(t)\|_{L^2(\nu_0)}^2 \Big)\,\mathrm{d}t + \frac{\gamma}{2}\int_0^T \|h^\dagger(t)\|_{H^s(\Omega)}^2\,\mathrm{d}t.\tag{24}\] If \(h\in\mathcal{A}_{\mathrm{ad}}\) is a minimizer of \(\mathcal{I}_{\lambda,\mu,\gamma}^{\,y}\) for data \(\mathsf{d}\), then \[\label{eq:a-priori-functional-bound} \mathcal{I}_{\lambda,\mu,\gamma}^{\,y}[h] \le \tfrac12\delta^2+R[h^\dagger].\tag{25}\] In particular, \[\begin{align} \tag{26} \tfrac12\|\mathcal{G}_\cdot^y(h(\cdot))-d\|_{L^2(0,T;\mathbb{R}^r)}^2 & \le \tfrac12\delta^2+R[h^\dagger], \\ \tag{27} \tfrac{\mu}{2}\|\dot{h}\|_{L^2(0,T;L^2(\nu_0))}^2 & \le \tfrac12\delta^2+R[h^\dagger]. \end{align}\]
Proof. Since \(h\) minimizes \(\mathcal{I}_{\lambda,\mu,\gamma}^{\,y}\) and \(h^\dagger\in\mathcal{A}_{\mathrm{ad}}\), \[\mathcal{I}_{\lambda,\mu,\gamma}^{\,y}[h] \le \mathcal{I}_{\lambda,\mu,\gamma}^{\,y}[h^\dagger] = \tfrac12\|\eta\|_{L^2(0,T;\mathbb{R}^r)}^2+R[h^\dagger] \le \tfrac12\delta^2+R[h^\dagger].\] The bounds 26 and 27 follow because each term in \(\mathcal{I}_{\lambda,\mu,\gamma}^{\,y}[h]\) is nonnegative. ◻
Theorem 32 (Local pathwise stability under partial observability). Assume Assumptions 5 and 6. Let \(h^\dagger\in \mathcal{A}_{\mathrm{ad}}^{s'}\), and let \[\mathsf{d}(t)=\mathcal{G}_t^y(h^\dagger(t))+\eta(t)\] with \[\|\eta\|_{L^2(0,T;\mathbb{R}^r)}\le\delta.\] Let \(h\in\mathcal{A}_{\mathrm{ad}}^{s'}\) be a minimizer of \(\mathcal{I}_{\lambda,\mu,\gamma}^{\,y}\) for data \(\mathsf{d}\).
Then there exist constants \(r>0\) and \(C>0\), depending only on \(\mathcal{X}_{\mathrm{ad}}\), \(\Omega\), \(y\), the coercivity constant \(\kappa\), and the regularity constants in Assumption 6, such that if \[\label{eq:proximity-condition} \sup_{t\in[0,T]}\|h(t)-h^\dagger(t)\|_{H^{s'}(\Omega)}\le r,\tag{28}\] then \[\label{eq:pathwise-stability-estimate} \|h-h^\dagger\|_{L^2(0,T;L^2(\nu_0))}^2 \le \frac{C}{\kappa}(1+\mu^{-1}) \Big( \delta^2+R[h^\dagger] \Big).\tag{29}\]
Proof. Set \(\xi:=h-h^\dagger\). We estimate \(\|\xi\|_{L^2(0,T;L^2(\nu_0))}^2\) by applying the joint coercivity of \(Q_y[h^\dagger]\) and bounding the resulting terms using minimality and the proximity condition.
Step 1: joint coercivity. By Assumption 5 applied at the path \(h^\dagger\), \[\label{eq:stab-joint-coercivity} \kappa\|\xi\|_{L^2_tL^2_x}^2 \le \underbrace{ \int_0^T \|J_{t,h^\dagger(t)}^{y}\xi(t)\|_{\mathbb{R}^r}^2\,\mathrm{d}t }_{\mathrm{Term\;I}} + \underbrace{ \int_0^T \mathfrak g_{h^\dagger(t)}(\dot{\xi}(t),\dot{\xi}(t))\,\mathrm{d}t }_{\mathrm{Term\;II}}.\tag{30}\]
Step 2: control of Term I. Since \(h^\dagger\in \mathcal{A}_{\mathrm{ad}}^{s'}\), the image \(h^\dagger([0,T])\subset \mathcal{Z}\) is compact. Because \(U\subset\mathcal{Z}\) is an open neighborhood of \(\mathcal{X}_{\mathrm{ad}}\) and \(h^\dagger(t)\in\mathcal{X}_{\mathrm{ad}}\subset U\) for all \(t\), there exists \(r_U>0\) such that \[\bigcup_{t\in[0,T]} B_{\mathcal{Z}}(h^\dagger(t),r_U)\subset U.\] We choose \(r\) so that \[r\le \min\left\{r_0,\;r_U,\;\sqrt{\frac{\kappa}{L^2}}\right\}.\] Then for every \(t\in[0,T]\) and every \(\theta\in[0,1]\), \[h^\dagger(t)+\theta\xi(t)\in U,\] so the fundamental-theorem-of-calculus expansion and the Lipschitz estimate 23 may be applied along the segment between \(h^\dagger(t)\) and \(h(t)\).
By the fundamental theorem of calculus in Banach spaces, \[\mathcal{G}_t^y(h(t))-\mathcal{G}_t^y(h^\dagger(t)) = J_{t,h^\dagger(t)}^{y}\xi(t) + \int_0^1 \bigl(J_{t,h^\dagger(t)+\theta\xi(t)}^{y}-J_{t,h^\dagger(t)}^{y}\bigr)\xi(t)\,\mathrm{d}\theta.\] Denote the remainder by \[r(t) := \int_0^1 \bigl(J_{t,h^\dagger(t)+\theta\xi(t)}^{y}-J_{t,h^\dagger(t)}^{y}\bigr)\xi(t)\,\mathrm{d}\theta.\] By the Lipschitz condition 23 and the proximity condition 28 , \[\|r(t)\|_{\mathbb{R}^r} \le \int_0^1 L\theta\|\xi(t)\|_{H^{s'}(\Omega)}\|\xi(t)\|_{L^2(\nu_0)}\,\mathrm{d}\theta \le \frac{Lr}{2}\|\xi(t)\|_{L^2(\nu_0)}.\] Since \[J_{t,h^\dagger(t)}^{y}\xi(t) = \mathcal{G}_t^y(h(t))-\mathcal{G}_t^y(h^\dagger(t))-r(t),\] we obtain \[\|J_{t,h^\dagger(t)}^{y}\xi(t)\|_{\mathbb{R}^r}^2 \le 2\|\mathcal{G}_t^y(h(t))-\mathcal{G}_t^y(h^\dagger(t))\|_{\mathbb{R}^r}^2 + 2\|r(t)\|_{\mathbb{R}^r}^2.\] For the first piece, write \[\mathcal{G}_t^y(h(t))-\mathcal{G}_t^y(h^\dagger(t)) = \bigl(\mathcal{G}_t^y(h(t))-\mathsf{d}(t)\bigr)+\eta(t),\] so that \[\|\mathcal{G}_t^y(h(t))-\mathcal{G}_t^y(h^\dagger(t))\|_{\mathbb{R}^r}^2 \le 2\|\mathcal{G}_t^y(h(t))-\mathsf{d}(t)\|_{\mathbb{R}^r}^2+2\|\eta(t)\|_{\mathbb{R}^r}^2.\] Integrating in time and applying Lemma 3, \[\int_0^T \|\mathcal{G}_t^y(h(t))-\mathcal{G}_t^y(h^\dagger(t))\|_{\mathbb{R}^r}^2\,\mathrm{d}t \le 4\bigl(\tfrac12\delta^2+R[h^\dagger]\bigr)+2\delta^2 \le C_1\bigl(\delta^2+R[h^\dagger]\bigr).\] For the remainder, \[\int_0^T \|r(t)\|_{\mathbb{R}^r}^2\,\mathrm{d}t \le \frac{L^2 r^2}{4}\|\xi\|_{L^2_tL^2_x}^2.\] Therefore, \[\label{eq:stab-term-I-bound} \mathrm{Term\;I} \le C_1\bigl(\delta^2+R[h^\dagger]\bigr) + \frac{L^2 r^2}{2}\|\xi\|_{L^2_tL^2_x}^2.\tag{31}\]
Step 3: control of Term II. By Proposition 23, \[\mathrm{Term\;II} = \int_0^T \mathfrak g_{h^\dagger(t)}(\dot{\xi}(t),\dot{\xi}(t))\,\mathrm{d}t \le C_{\mathrm{tr}}\|\dot{\xi}\|_{L^2_tL^2_x}^2.\] Since \(\xi=h-h^\dagger\), we have \[\dot{\xi}=\dot{h}-\dot{h}^\dagger,\] and therefore \[\|\dot{\xi}\|_{L^2_tL^2_x}^2 \le 2\|\dot{h}\|_{L^2_tL^2_x}^2 + 2\|\dot{h}^\dagger\|_{L^2_tL^2_x}^2.\] By the a priori bound 27 , \[\frac{\mu}{2}\|\dot{h}\|_{L^2_tL^2_x}^2 \le \frac{1}{2}\delta^2+R[h^\dagger],\] so \[\|\dot{h}\|_{L^2_tL^2_x}^2 \le \frac{1}{\mu}\bigl(\delta^2+2R[h^\dagger]\bigr).\] Also, by the definition of \(R[h^\dagger]\), \[R[h^\dagger] \ge \frac{\mu}{2}\|\dot{h}^\dagger\|_{L^2_tL^2_x}^2,\] hence \[\|\dot{h}^\dagger\|_{L^2_tL^2_x}^2 \le \frac{2}{\mu}R[h^\dagger].\] Combining these estimates gives \[\label{eq:stab-term-II-bound} \mathrm{Term\;II} \le \frac{C_3}{\mu}\bigl(\delta^2+R[h^\dagger]\bigr)\tag{32}\] for a constant \(C_3>0\) depending only on \(C_{\mathrm{tr}}\).
Step 4: combine and absorb. Substituting 31 and 32 into 30 gives \[\kappa\|\xi\|_{L^2_tL^2_x}^2 \le \frac{L^2 r^2}{2}\|\xi\|_{L^2_tL^2_x}^2 + C_1\bigl(\delta^2+R[h^\dagger]\bigr) + \frac{C_3}{\mu}\bigl(\delta^2+R[h^\dagger]\bigr).\] Choose \[r \le \min\left\{r_0,\;r_U,\;\sqrt{\frac{\kappa}{L^2}}\right\},\] so that \(\frac{L^2 r^2}{2}\le \frac{\kappa}{2}\). Then \[\frac{\kappa}{2}\|\xi\|_{L^2_tL^2_x}^2 \le C(1+\mu^{-1})\bigl(\delta^2+R[h^\dagger]\bigr),\] and dividing by \(\kappa/2\) gives 29 . ◻
Remark 33 (The price of partial observability). Under full pointwise observability, Term II in the proof is unnecessary, and the stability estimate takes the form \[\|h-h^\dagger\|_{L^2_tL^2_x}^2 \le \frac{C}{\kappa}\bigl(\delta^2+R[h^\dagger]\bigr)\] with no factor of \(1/\mu\). Under partial observability, the transport term in \(Q_y[h]\) is needed to close the observability gap, and the \(\mu\)-regularization is what provides a priori control of this term. The resulting factor of \(1+\mu^{-1}\) is the quantitative cost of relying on dynamical coupling rather than instantaneous data; in the common regime \(0<\mu\le1\), this cost is equivalent to \(1/\mu\).
Remark 34 (On the proximity condition). The proximity condition 28 is a local hypothesis. Under fixed \(\lambda,\mu,\gamma>0\), vanishing-noise minimizers are precompact in \(C([0,T];H^{s'}(\Omega))\), and subsequential limits solve the noiseless regularized inverse problem. Thus the proximity condition is guaranteed in a small-noise regime provided the corresponding noiseless regularized minimizer lies within the ball 28 around \(h^\dagger\). In this sense, Theorem 32 is a local stability statement near the reference path.
Corollary 5 (Vanishing-noise convergence to a noiseless regularized minimizer). Assume the hypotheses of Theorem 30. Let \[\mathcal{R}_{\lambda,\mu,\gamma}[h] := \frac{\lambda}{2}\int_0^T \mathfrak g_{h(t)}(\dot{h}(t),\dot{h}(t))\,\mathrm{d}t + \frac{\mu}{2}\int_0^T \Big( \|h(t)\|_{L^2(\nu_0)}^2+\|\dot{h}(t)\|_{L^2(\nu_0)}^2 \Big)\,\mathrm{d}t + \frac{\gamma}{2}\int_0^T \|h(t)\|_{H^s(\Omega)}^2\,\mathrm{d}t.\] For a sequence \(\delta_n\to 0\), let \[\mathsf{d}_n(t)=\mathcal{G}_t^y(h^\dagger(t))+\eta_n(t), \qquad \|\eta_n\|_{L^2(0,T;\mathbb{R}^r)}\le \delta_n,\] and let \(h_n\in\mathcal{A}_{\mathrm{ad}}\) be a minimizer of \[\mathcal{I}_n[h] := \frac{1}{2}\|\mathcal{G}_\cdot^y(h(\cdot))-\mathsf{d}_n\|_{L^2(0,T;\mathbb{R}^r)}^2 + \mathcal{R}_{\lambda,\mu,\gamma}[h].\] Define the noiseless regularized functional \[\mathcal{I}^0[h] := \frac{1}{2}\|\mathcal{G}_\cdot^y(h(\cdot))-\mathcal{G}_\cdot^y(h^\dagger(\cdot))\|_{L^2(0,T;\mathbb{R}^r)}^2 + \mathcal{R}_{\lambda,\mu,\gamma}[h].\]
Then the sequence \((h_n)\) is bounded in \[L^2(0,T;H^s(\Omega))\cap H^1(0,T;L^2_0(\nu_0)).\] Consequently, after passing to a subsequence, \[h_n\to h_\ast \qquad\text{in } C([0,T];H^{s'}(\Omega))\] for some \(h_\ast\in\mathcal{A}_{\mathrm{ad}}\).
Moreover, \(h_\ast\) is a minimizer of \(\mathcal{I}^0\) over \(\mathcal{A}_{\mathrm{ad}}\).
If \(\mathcal{I}^0\) has a unique minimizer \(h^0_{\lambda,\mu,\gamma}\), then the whole sequence satisfies \[h_n\to h^0_{\lambda,\mu,\gamma} \qquad\text{in } C([0,T];H^{s'}(\Omega)).\]
Finally, assume the hypotheses of Theorem 32 hold, and let \(r>0\) be the radius from Theorem 32. If \[\sup_{t\in[0,T]} \|h^0_{\lambda,\mu,\gamma}(t)-h^\dagger(t)\|_{H^{s'}(\Omega)}<r,\] then for all sufficiently large \(n\), the minimizers \(h_n\) satisfy the proximity condition 28 , and therefore \[\|h_n-h^\dagger\|_{L^2(0,T;L^2(\nu_0))}^2 \le \frac{C}{\kappa}(1+\mu^{-1}) \bigl(\delta_n^2+R[h^\dagger]\bigr)\] for all sufficiently large \(n\).
Proof. Since \(h_n\) minimizes \(\mathcal{I}_n\) and \(h^\dagger\in\mathcal{A}_{\mathrm{ad}}\), we have \[\mathcal{I}_n[h_n] \le \mathcal{I}_n[h^\dagger] = \frac{1}{2}\|\eta_n\|_{L^2(0,T;\mathbb{R}^r)}^2 + \mathcal{R}_{\lambda,\mu,\gamma}[h^\dagger] \le \frac{1}{2}\delta_n^2+\mathcal{R}_{\lambda,\mu,\gamma}[h^\dagger].\] Since the data-misfit term and the transport term are nonnegative, this yields the uniform bound \[\sup_n \left( \|h_n\|_{L^2(0,T;H^s(\Omega))} + \|h_n\|_{H^1(0,T;L^2_0(\nu_0))} \right) <\infty.\] By Proposition 29, after passing to a subsequence, \[h_n\to h_\ast \quad\text{in } C([0,T];H^{s'}(\Omega)),\] \[h_n\rightharpoonup h_\ast \quad\text{weakly in } L^2(0,T;H^s(\Omega)),\] and \[\dot{h}_n\rightharpoonup \dot{h}_\ast \quad\text{weakly in } L^2(0,T;L^2_0(\nu_0))\] for some \(h_\ast\in\mathcal{A}_{\mathrm{ad}}\).
Because \(\mathsf{d}_n\to \mathcal{G}_\cdot^y(h^\dagger(\cdot))\) in \(L^2(0,T;\mathbb{R}^r)\), for each fixed \(h\in\mathcal{A}_{\mathrm{ad}}\) we have \[\mathcal{I}_n[h]\to \mathcal{I}^0[h].\] Also, by dominated convergence, \[\mathcal{G}_\cdot^y(h_n(\cdot))\to \mathcal{G}_\cdot^y(h_\ast(\cdot)) \quad\text{in } L^2(0,T;\mathbb{R}^r).\] Together with Proposition 28 and weak lower semicontinuity of the \(\mu\)- and \(\gamma\)-terms, this gives \[\mathcal{I}^0[h_\ast] \le \liminf_{n\to\infty}\mathcal{I}_n[h_n].\]
Now let \(h\in\mathcal{A}_{\mathrm{ad}}\) be arbitrary. Since \(h_n\) minimizes \(\mathcal{I}_n\), \[\mathcal{I}_n[h_n]\le \mathcal{I}_n[h].\] Passing to the limit superior on the right and combining with the previous liminf inequality yields \[\mathcal{I}^0[h_\ast]\le \mathcal{I}^0[h].\] Thus \(h_\ast\) is a minimizer of \(\mathcal{I}^0\).
If \(\mathcal{I}^0\) has a unique minimizer \(h^0_{\lambda,\mu,\gamma}\), then every convergent subsequence of \((h_n)\) has the same limit, so the whole sequence converges to \(h^0_{\lambda,\mu,\gamma}\) in \(C([0,T];H^{s'}(\Omega))\).
For the final claim, assume \[\sup_{t\in[0,T]} \|h^0_{\lambda,\mu,\gamma}(t)-h^\dagger(t)\|_{H^{s'}(\Omega)}<r.\] Since \(h_n\to h^0_{\lambda,\mu,\gamma}\) in \(C([0,T];H^{s'}(\Omega))\), it follows that for all sufficiently large \(n\), \[\sup_{t\in[0,T]} \|h_n(t)-h^\dagger(t)\|_{H^{s'}(\Omega)}<r.\] Hence the proximity condition 28 holds for \(h_n\), and Theorem 32 gives \[\|h_n-h^\dagger\|_{L^2(0,T;L^2(\nu_0))}^2 \le \frac{C}{\kappa}(1+\mu^{-1}) \bigl(\delta_n^2+R[h^\dagger]\bigr)\] for all sufficiently large \(n\). ◻
In particular, under fixed positive regularization parameters \(\lambda,\mu,\gamma\), vanishing-noise reconstructions converge to the noiseless regularized inverse problem. By contrast, Theorem 35 shows that a simultaneous vanishing-noise, vanishing-regularization regime yields convergence to an exact noiseless solution of minimum regularization energy.
Theorem 35 (Vanishing-regularization consistency toward a minimum-regularity exact solution). Assume the hypotheses of Theorem 30. Fix \[\bar\lambda,\bar\mu,\bar\gamma>0,\] and define the reference regularization functional \[\begin{align} \label{eq:reference-regularization-functional} \mathcal{R}[h] & := \frac{\bar\lambda}{2}\int_0^T \mathfrak g_{h(t)}(\dot{h}(t),\dot{h}(t))\,\mathrm{d}t + \frac{\bar\mu}{2}\int_0^T \Big( \|h(t)\|_{L^2(\nu_0)}^2+\|\dot{h}(t)\|_{L^2(\nu_0)}^2 \Big)\,\mathrm{d}t \notag \\ & \quad + \frac{\bar\gamma}{2}\int_0^T \|h(t)\|_{H^s(\Omega)}^2\,\mathrm{d}t . \end{align}\tag{33}\] Let \(\mathsf{d}^\dagger \in L^2(0,T;\mathbb{R}^r)\), and assume that the exact solution set \[\mathcal{S}(\mathsf{d}^\dagger) := \left\{ h\in \mathcal{A}_{\mathrm{ad}} : \mathcal{G}_\cdot^y(h(\cdot))=\mathsf{d}^\dagger \text{ in } L^2(0,T;\mathbb{R}^r) \right\}\] is nonempty.
Let \(\mathsf{d}_n\in L^2(0,T;\mathbb{R}^r)\) satisfy \[\|\mathsf{d}_n-\mathsf{d}^\dagger\|_{L^2(0,T;\mathbb{R}^r)}\le \delta_n, \qquad \delta_n\to 0,\] and let \(\alpha_n>0\) satisfy \[\alpha_n\to 0, \qquad \frac{\delta_n^2}{\alpha_n}\to 0.\] For each \(n\), define \[\label{eq:vanishing-regularization-functional} \mathcal{I}_n[h] := \frac{1}{2} \|\mathcal{G}_\cdot^y(h(\cdot))-\mathsf{d}_n\|_{L^2(0,T;\mathbb{R}^r)}^2 + \alpha_n \mathcal{R}[h], \qquad h\in \mathcal{A}_{\mathrm{ad}},\tag{34}\] and let \(h_n\in \mathcal{A}_{\mathrm{ad}}\) be a minimizer of \(\mathcal{I}_n\).
Then the sequence \((h_n)\) is bounded in \[L^2(0,T;H^s(\Omega)) \cap H^1(0,T;L^2_0(\nu_0)).\] Consequently, after passing to a subsequence, \[h_n \to h_\ast \qquad\text{in } C([0,T];H^{s'}(\Omega))\] for some \(h_\ast\in \mathcal{A}_{\mathrm{ad}}\).
Moreover, \(h_\ast\) is an exact solution: \[h_\ast \in \mathcal{S}(\mathsf{d}^\dagger),\] and it minimizes \(\mathcal{R}\) over the exact solution set: \[\mathcal{R}[h_\ast] = \min\left\{ \mathcal{R}[h] : h\in \mathcal{S}(\mathsf{d}^\dagger) \right\}.\]
If, in addition, the \(\mathcal{R}\)-minimizing exact solution is unique, then the whole sequence satisfies \[h_n \to h_\ast \qquad\text{in } C([0,T];H^{s'}(\Omega)).\]
Proof. Let \(\tilde{h}\in \mathcal{S}(\mathsf{d}^\dagger)\) be arbitrary. Since \(h_n\) minimizes \(\mathcal{I}_n\), \[\frac{1}{2} \|\mathcal{G}_\cdot^y(h_n(\cdot))-\mathsf{d}_n\|_{L^2(0,T;\mathbb{R}^r)}^2 + \alpha_n \mathcal{R}[h_n] \le \frac{1}{2} \|\mathcal{G}_\cdot^y(\tilde{h}(\cdot))-\mathsf{d}_n\|_{L^2(0,T;\mathbb{R}^r)}^2 + \alpha_n \mathcal{R}[\tilde{h}].\] Because \(\mathcal{G}_\cdot^y(\tilde{h}(\cdot))=\mathsf{d}^\dagger\), this gives \[\frac{1}{2} \|\mathcal{G}_\cdot^y(h_n(\cdot))-\mathsf{d}_n\|_{L^2(0,T;\mathbb{R}^r)}^2 + \alpha_n \mathcal{R}[h_n] \le \frac{1}{2} \|\mathsf{d}^\dagger-\mathsf{d}_n\|_{L^2(0,T;\mathbb{R}^r)}^2 + \alpha_n \mathcal{R}[\tilde{h}] \le \frac{1}{2} \delta_n^2 + \alpha_n \mathcal{R}[\tilde{h}].\] Hence \[\label{eq:vanishing-regularization-bound} \mathcal{R}[h_n] \le \mathcal{R}[\tilde{h}] + \frac{\delta_n^2}{2\alpha_n}.\tag{35}\] Since \(\delta_n^2/\alpha_n \to 0\), the sequence \((\mathcal{R}[h_n])\) is bounded.
Because \(\bar\mu,\bar\gamma>0\), boundedness of \(\mathcal{R}[h_n]\) implies boundedness of \[\|h_n\|_{L^2(0,T;H^s(\Omega))} \qquad\text{and}\qquad \|h_n\|_{H^1(0,T;L^2_0(\nu_0))}.\] Therefore Proposition 29 applies, and after passing to a subsequence, \[h_n \rightharpoonup h_\ast \quad\text{weakly in } L^2(0,T;H^s(\Omega)),\] \[h_n \rightharpoonup h_\ast \quad\text{weakly in } H^1(0,T;L^2_0(\nu_0)),\] and \[h_n \to h_\ast \quad\text{in } C([0,T];H^{s'}(\Omega))\] for some \(h_\ast\in \mathcal{A}_{\mathrm{ad}}\).
Next, from the minimizing inequality above we also have \[\frac{1}{2} \|\mathcal{G}_\cdot^y(h_n(\cdot))-\mathsf{d}_n\|_{L^2(0,T;\mathbb{R}^r)}^2 \le \frac{1}{2} \delta_n^2 + \alpha_n \mathcal{R}[\tilde{h}].\] Since \(\delta_n\to 0\) and \(\alpha_n\to 0\), it follows that \[\|\mathcal{G}_\cdot^y(h_n(\cdot))-\mathsf{d}_n\|_{L^2(0,T;\mathbb{R}^r)} \to 0.\] Therefore \[\|\mathcal{G}_\cdot^y(h_n(\cdot))-\mathsf{d}^\dagger\|_{L^2(0,T;\mathbb{R}^r)} \le \|\mathcal{G}_\cdot^y(h_n(\cdot))-\mathsf{d}_n\|_{L^2(0,T;\mathbb{R}^r)} + \|\mathsf{d}_n-\mathsf{d}^\dagger\|_{L^2(0,T;\mathbb{R}^r)} \to 0.\] Since \(h_n\to h_\ast\) in \(C([0,T];H^{s'}(\Omega))\), dominated convergence yields \[\mathcal{G}_\cdot^y(h_\ast(\cdot))=\mathsf{d}^\dagger.\] Thus \(h_\ast\in \mathcal{S}(\mathsf{d}^\dagger)\).
It remains to show that \(h_\ast\) minimizes \(\mathcal{R}\) over \(\mathcal{S}(\mathsf{d}^\dagger)\). Let \(z\in \mathcal{S}(\mathsf{d}^\dagger)\) be arbitrary. By minimality of \(h_n\), \[\frac{1}{2} \|\mathcal{G}_\cdot^y(h_n(\cdot))-\mathsf{d}_n\|_{L^2(0,T;\mathbb{R}^r)}^2 + \alpha_n \mathcal{R}[h_n] \le \frac{1}{2} \|\mathcal{G}_\cdot^y(z(\cdot))-\mathsf{d}_n\|_{L^2(0,T;\mathbb{R}^r)}^2 + \alpha_n \mathcal{R}[z].\] Using \(\mathcal{G}_\cdot^y(z(\cdot))=\mathsf{d}^\dagger\), we obtain \[\alpha_n \mathcal{R}[h_n] \le \frac{1}{2} \delta_n^2 + \alpha_n \mathcal{R}[z].\] Dividing by \(\alpha_n\) yields \[\mathcal{R}[h_n] \le \mathcal{R}[z] + \frac{\delta_n^2}{2\alpha_n}.\] Taking the limit superior and using \(\delta_n^2/\alpha_n\to 0\), we obtain \[\limsup_{n\to\infty}\mathcal{R}[h_n]\le \mathcal{R}[z].\]
On the other hand, Proposition 28 gives lower semicontinuity of the transport part of \(\mathcal{R}\), while the remaining two terms are weakly lower semicontinuous. Hence \[\mathcal{R}[h_\ast] \le \liminf_{n\to\infty}\mathcal{R}[h_n].\] Combining the last two inequalities gives \[\mathcal{R}[h_\ast]\le \mathcal{R}[z].\] Since \(z\in \mathcal{S}(\mathsf{d}^\dagger)\) was arbitrary, \(h_\ast\) minimizes \(\mathcal{R}\) over \(\mathcal{S}(\mathsf{d}^\dagger)\).
Finally, if the \(\mathcal{R}\)-minimizing exact solution is unique, then every convergent subsequence of \((h_n)\) has the same limit \(h_\ast\). Therefore the whole sequence converges to \(h_\ast\) in \(C([0,T];H^{s'}(\Omega))\). ◻
Remark 36 (Interpretation as Tikhonov consistency). Theorem 35 strengthens Corollary 5. For fixed positive regularization parameters, vanishing-noise reconstructions converge only to a minimizer of the noiseless regularized problem. By contrast, when the regularization strength \(\alpha_n\) tends to zero in such a way that \(\delta_n^2/\alpha_n\to 0\), the reconstructions converge to an exact solution of the noiseless inverse problem, and among exact solutions they select one of minimum regularization energy \(\mathcal{R}\). In this sense, the theorem gives a Tikhonov-type consistency statement for the Bayes Hilbert path-space inverse problem with prescribed mobile sensors.
The joint transport–observability coercivity condition (Assumption 5) is sufficient for the stability theory, but it is a strong hypothesis. In this subsection we remain entirely within the ambient mobile-sensor setting introduced above and analyze the joint form without assuming coercivity. We first characterize its null space, then prove that localized mobile sensors still yield a compact time-averaged observation operator in the infinite-dimensional ambient space, and finally record the resulting spectral decomposition. This identifies the directions that are effectively resolved and unresolved by the sensor motion over the full time horizon. The finite-dimensional recovery mechanism is deferred to Section 5.
We begin with the null space of the ambient joint form \(Q_y[h]\).
Proposition 37 (Null-space characterization). Let \(h\in \mathcal{A}_{\mathrm{ad}}\). Then \[\xi\in\ker Q_y[h] \quad\Longleftrightarrow\quad \left\{ \begin{align} & \xi(t)\in\ker J_{t,h(t)}^{y} \quad\text{for a.e.\;}t\in[0,T], \\ & \dot{\xi}(t)=0 \quad\text{for a.e.\;}t\in[0,T]. \end{align} \right.\] In particular, every element of \(\ker Q_y[h]\) is time-constant: \[\xi(t)\equiv \xi_0,\] with \[\label{eq:null-space-Q} \ker Q_y[h] = \left\{ \xi_0\in L^2_0(\nu_0): \xi_0\in\ker J_{t,h(t)}^{y} \text{ for a.e.\;}t\in[0,T] \right\}.\qquad{(15)}\]
Proof. If \(Q_y[h](\xi,\xi)=0\), then both terms in the definition of \(Q_y[h]\) vanish: \[\int_0^T \|J_{t,h(t)}^{y}\xi(t)\|_{\mathbb{R}^r}^2\,\mathrm{d}t=0, \qquad \int_0^T \mathfrak g_{h(t)}(\dot{\xi}(t),\dot{\xi}(t))\,\mathrm{d}t=0.\] The first identity gives \(\xi(t)\in\ker J_{t,h(t)}^{y}\) for a.e.\(t\). By Proposition 12, the second identity gives \(\dot{\xi}(t)=0\) for a.e.\(t\). Hence \(\xi\) is time-constant. The converse is immediate. ◻
Proposition 37 shows that the joint form has a much smaller null space than the instantaneous observation operator alone. At a fixed time \(t\), the kernel of \(J_{t,h(t)}^{y}\) may be large, but to lie in \(\ker Q_y[h]\) a perturbation must be simultaneously invisible at almost every time and constant in time. Thus the relevant obstruction is the common invisible subspace ?? .
The next result shows that, even with time-dependent sensor locations, localized mobile sensors do not produce ambient coercivity on the infinite-dimensional state space.
For \(h\in\mathcal{A}_{\mathrm{ad}}\) and sensor path \(y\), define the time-averaged observation form on time-constant perturbations by \[\label{eq:time-averaged-observation-form} \bar{\mathfrak j}_y[h](\xi_0,\eta_0) := \int_0^T \bigl\langle J_{t,h(t)}^{y}\xi_0,\;J_{t,h(t)}^{y}\eta_0\bigr\rangle_{\mathbb{R}^r}\,\mathrm{d}t, \qquad \xi_0,\eta_0\in L^2_0(\nu_0).\tag{36}\]
Proposition 38 (Compactness obstruction for localized mobile sensors). Let \(\kappa\in C_c(\mathbb{R}^d)\), let \(y_1,\dots,y_r:[0,T]\to\Omega\) be measurable, and let \(h\in\mathcal{A}_{\mathrm{ad}}\). Assume in addition that \(L^2_0(\nu_0)\) is infinite-dimensional. Then the form \(\bar{\mathfrak j}_y[h]\) is induced by a compact, self-adjoint, nonnegative operator \[\bar J_y[h]:L^2_0(\nu_0)\to L^2_0(\nu_0).\]
Consequently, there is no constant \(\kappa_{\mathrm{obs}}>0\) such that \[Q_y[h](\xi,\xi) \ge \kappa_{\mathrm{obs}} \|\xi\|_{L^2(0,T;L^2(\nu_0))}^2\] for every \[\xi\in L^2\bigl(0,T;H^{s'}(\Omega)\cap L^2_0(\nu_0)\bigr) \cap H^1\bigl(0,T;L^2_0(\nu_0)\bigr).\]
Proof. Define \[\mathcal{K}_y[h]:L^2_0(\nu_0)\to L^2(0,T;\mathbb{R}^r)\] by \[(\mathcal{K}_y[h]\xi_0)(t) := J_{t,h(t)}^{y}\xi_0.\] Then \[\bar{\mathfrak j}_y[h](\xi_0,\eta_0) = \langle \mathcal{K}_y[h]\xi_0,\mathcal{K}_y[h]\eta_0\rangle_{L^2(0,T;\mathbb{R}^r)},\] so \(\bar J_y[h]=\mathcal{K}_y[h]^\ast \mathcal{K}_y[h]\).
It remains to show that \(\mathcal{K}_y[h]\) is compact. For \(j=1,\dots,r\), write \[(\mathcal{K}_y[h]\xi_0)_j(t) = \int_\Omega b_j(t,x)\,\xi_0(x)\,\mathrm{d}\nu_0(x),\] where \[b_j(t,x) := \Bigl( \kappa(x-y_j(t)) - \mathbb{E}_{\rho_{h(t)}}[\kappa(\cdot-y_j(t))] \Bigr) \,w_{h(t)}(x), \qquad w_{h(t)}:=\frac{\mathrm{d}\rho_{h(t)}}{\mathrm{d}\nu_0}.\] Since \(\kappa\in L^\infty\) and the admissibility bounds give \[0<c\le w_{h(t)}(x)\le C<\infty \qquad\text{for a.e.\;}(t,x)\in[0,T]\times\Omega,\] we have \[|b_j(t,x)|\le 2C\|\kappa\|_{L^\infty(\mathbb{R}^d)},\] hence \(b_j\in L^2((0,T)\times\Omega)\). Therefore each component \[\xi_0\longmapsto (\mathcal{K}_y[h]\xi_0)_j\] is a Hilbert–Schmidt operator from \(L^2_0(\nu_0)\) to \(L^2(0,T)\), and so \(\mathcal{K}_y[h]\) itself is Hilbert–Schmidt, hence compact.
Thus \(\bar J_y[h]=\mathcal{K}_y[h]^\ast\mathcal{K}_y[h]\) is compact, self-adjoint, and nonnegative.
For the coercivity claim, suppose by contradiction that there exists \(\kappa_{\mathrm{obs}}>0\) such that \[Q_y[h](\xi,\xi) \ge \kappa_{\mathrm{obs}} \|\xi\|_{L^2(0,T;L^2(\nu_0))}^2\] for every admissible perturbation \(\xi\). Let \(\xi_0\in H^{s'}(\Omega)\cap L^2_0(\nu_0)\), and consider the time-constant path \(\xi(t)\equiv\xi_0\). Then \(\dot{\xi}=0\), so \[Q_y[h](\xi,\xi) = \bar{\mathfrak j}_y[h](\xi_0,\xi_0), \qquad \|\xi\|_{L^2(0,T;L^2(\nu_0))}^2 = T\|\xi_0\|_{L^2(\nu_0)}^2.\] Hence \[\bar{\mathfrak j}_y[h](\xi_0,\xi_0) \ge \kappa_{\mathrm{obs}}T\|\xi_0\|_{L^2(\nu_0)}^2 \qquad \forall \xi_0\in H^{s'}(\Omega)\cap L^2_0(\nu_0).\] Because \(H^{s'}(\Omega)\cap L^2_0(\nu_0)\) is dense in \(L^2_0(\nu_0)\) and \(\bar{\mathfrak j}_y[h]\) is continuous on \(L^2_0(\nu_0)\), this extends to all \(\xi_0\in L^2_0(\nu_0)\). Thus the compact operator \(\bar J_y[h]\) would be bounded below by a positive multiple of the identity on an infinite-dimensional Hilbert space, which is impossible. ◻
Remark 39 (Physical interpretation). The obstruction in Proposition 38 is a property of the observation modality, not of the sensor trajectory. Localized kernels act as spatial averaging operators. No matter how the sensors move, sufficiently oscillatory perturbations remain weakly visible, and the corresponding time-averaged observation eigenvalues must decay to zero.
Although ambient coercivity fails, Proposition 38 yields a precise spectral description of the directions that are strongly or weakly seen by the mobile sensors over the full time horizon.
Proposition 40 (Spectral decomposition of the averaged observation form). Let \(h\in\mathcal{A}_{\mathrm{ad}}\), and let \(y\) be a measurable sensor trajectory. Then there exist an orthonormal basis \(\{e_k\}_{k\ge 1}\) of \(L^2_0(\nu_0)\) and a nonincreasing sequence \(\sigma_k^2\downarrow 0\) such that \[\bar J_y[h]e_k=\sigma_k^2 e_k, \qquad k\ge 1,\] and \[\label{eq:spectral-expansion-mobile} \bar{\mathfrak j}_y[h](\xi_0,\xi_0) = \sum_{k=1}^\infty \sigma_k^2 \bigl|\langle \xi_0,e_k\rangle_{L^2(\nu_0)}\bigr|^2 \qquad \forall \xi_0\in L^2_0(\nu_0).\qquad{(16)}\]
Proof. This is the spectral theorem for compact self-adjoint operators applied to \(\bar J_y[h]\). ◻
Definition 14 (Resolved and unresolved directions). Let \(\mu>0\). A spectral direction \(e_k\) is called
resolved at scale \(\mu\) if \(\sigma_k^2\ge \mu\),
unresolved at scale \(\mu\) if \(\sigma_k^2<\mu\).
The corresponding effective resolved dimension is \[\label{eq:effective-dimension} m_{\mathrm{eff}}(\mu) := \#\{k:\sigma_k^2\ge \mu\}.\tag{37}\]
Remark 41 (Relation to the regularized inverse functional). The regularized inverse functional from Definition 13 contains the term \[\frac{\mu}{2}\int_0^T \|h(t)\|_{L^2(\nu_0)}^2\,\mathrm{d}t.\] For time-constant perturbations, the competition between the data term and this \(L^2\)-penalty is measured by the eigenvalues \(\sigma_k^2\) in ?? . Directions with \(\sigma_k^2\gg\mu\) are primarily constrained by the observations over the time horizon, whereas directions with \(\sigma_k^2\ll\mu\) are controlled predominantly by regularization. In this sense, \(m_{\mathrm{eff}}(\mu)\) counts the number of ambient directions that are observed at scale \(\mu\) by the mobile sensor system.
The ambient theory developed in Sections 3 and 4 is formulated on admissible Bayes Hilbert paths \[h\in \mathcal{A}_{\mathrm{ad}}.\] We now restrict that theory to finite-dimensional Bayes Hilbert subspaces. This serves two purposes.
First, it shows that the reduced transport and inverse objects are not ad hoc constructions. They are simply the coordinate representations of the ambient geometry after restricting the state variable \(h\) to a finite-dimensional subspace \(V_m\subset \mathcal{X}\). In particular, the reduced transport tensor and reduced sensing matrices introduced below are the finite-dimensional pullbacks of the ambient transport and observability forms.
Second, the finite-dimensional setting is the level at which localized sensing yields genuine recovery theorems. In the ambient infinite-dimensional problem, mobile sensing can remove common invisible directions and may even make the joint transport–observability form injective, but it does not in general produce a coercive stability estimate on the full state space. After restriction to \(V_m\), this compactness obstruction disappears, and one can formulate explicit reduced observability and recovery criteria in terms of finite-dimensional Gramian matrices.
There are, however, two distinct levels of result in this section, and it is useful to separate them from the start. The first level is truth-specific reduced observability: for a fixed reduced truth path, one can design localized sensors—including a single moving sensor—so that the resulting reduced Gramian is positive definite and the reduced joint form is coercive along that path. The second level is global reduced inversion on an admissible class: to solve a regularized inverse problem over an entire reduced class of coefficient paths, one needs an additional classwise reduced stability estimate for the chosen experiment. The last subsection makes this extra step explicit and then shows how reduced reconstructions lift to approximate ambient reconstructions.
The section has the following structure. We first pull back the ambient transport and observation geometry to a finite-dimensional coefficient space. We then show that a sufficiently localized bump sensor detects any fixed nonzero direction. Next, we prove that on a fixed \(m\)-dimensional reduced subspace, one can place finitely many stationary sensors so that the reduced state is instantaneously observable at every time. Finally, we show that a single moving sensor can recover the same reduced path by cycling through those locations, even though with one sensor the full reduced state need not be observable at any individual time.
Let \[V_m:=\operatorname{span}\{\phi_1,\dots,\phi_m\}\subset \mathcal{X}\] be an \(m\)-dimensional subspace of the ambient state space \(\mathcal{X}\), where \[\phi_1,\dots,\phi_m\in \mathcal{X}\] are linearly independent. For \(a=(a_1,\dots,a_m)\in\mathbb{R}^m\), define \[\label{eq:finite-dimensional-state} h(a):=\sum_{k=1}^m a_k\phi_k.\tag{38}\] Whenever \(h(a)\in \mathcal{X}_{\mathrm{ad}}\), we write \[\rho_a:=\rho_{h(a)}.\]
Thus a coefficient path \[a:[0,T]\to\mathbb{R}^m\] induces a Bayes Hilbert path \[h(t)=h(a(t))=\sum_{k=1}^m a_k(t)\phi_k\] and hence a law-valued path \[\rho_t=\rho_{a(t)}.\] The corresponding reduced admissible class is \[\mathcal{A}_{\mathrm{ad}}^{(m)} := \left\{ a\in H^1(0,T;\mathbb{R}^m) : h(a(\cdot))\in \mathcal{A}_{\mathrm{ad}} \right\}.\]
Remark 42. The finite-dimensional theory is obtained by restricting the ambient state variable \(h\) to the subspace \(V_m\). Every reduced transport, observation, and inverse object below is therefore the pullback of its ambient counterpart through the coordinate map \[\mathbb{R}^m\ni a\longmapsto h(a)\in V_m\subset \mathcal{X}.\]
Remark 43 (Finite-dimensional reductions as exponential families). Finite-dimensional affine subspaces of Bayes Hilbert spaces correspond to finite-dimensional exponential families relative to \(\nu_0\) [1]. In particular, the linear reduction \[h(a)=\sum_{k=1}^m a_k\phi_k\] represents the normalized exponential family \[\mathrm{d}\rho_a = \frac{ \exp\!\left(\sum_{k=1}^m a_k\phi_k\right) }{ \int_\Omega \exp\!\left(\sum_{k=1}^m a_k\phi_k\right) \,\mathrm{d}\nu_0 } \,\mathrm{d}\nu_0.\]
We begin by writing the ambient transport form and the ambient mobile-sensor observation map in coordinates on \(V_m\).
For \(a\in\mathbb{R}^m\) such that \(h(a)\in \mathcal{X}_{\mathrm{ad}}\), define \[\label{eq:reduced-transport-matrix} H(a)=\bigl(H_{k\ell}(a)\bigr)_{k,\ell=1}^m, \qquad H_{k\ell}(a):=\mathfrak g_{h(a)}(\phi_k,\phi_\ell).\tag{39}\]
Proposition 44 (Coordinate representation of the transport form). Let \(a\in\mathbb{R}^m\) with \(h(a)\in \mathcal{X}_{\mathrm{ad}}\), and let \[\alpha_h=\sum_{k=1}^m \alpha_k\phi_k, \qquad \beta_h=\sum_{k=1}^m \beta_k\phi_k\] be directions in \(V_m\), with coefficient vectors \[\alpha=(\alpha_1,\dots,\alpha_m)^\top, \qquad \beta=(\beta_1,\dots,\beta_m)^\top.\] Then \[\label{eq:coordinate-transport-form} \mathfrak g_{h(a)}(\alpha_h,\beta_h)=\alpha^\top H(a)\beta.\qquad{(17)}\] In particular, if \(a\in \mathcal{A}_{\mathrm{ad}}^{(m)}\), then \[\label{eq:coordinate-transport-action} \mathfrak g_{h(a(t))}\bigl(\dot{h}(t),\dot{h}(t)\bigr) = \dot{a}(t)^\top H(a(t))\dot{a}(t) \qquad\text{for a.e.\;}t\in[0,T].\qquad{(18)}\]
Proof. By bilinearity of \(\mathfrak g_{h(a)}\), \[\mathfrak g_{h(a)}(\alpha_h,\beta_h) = \sum_{k,\ell=1}^m \alpha_k\beta_\ell\,\mathfrak g_{h(a)}(\phi_k,\phi_\ell) = \sum_{k,\ell=1}^m \alpha_k\beta_\ell\,H_{k\ell}(a) = \alpha^\top H(a)\beta.\] Taking \[\dot{h}(t)=\sum_{k=1}^m \dot{a}_k(t)\phi_k\] gives ?? . ◻
We next specialize the mobile-sensor observation model of Section 4 to the reduced state space.
Definition 15 (Reduced mobile-sensor observation map). For \(t\in[0,T]\), a sensor trajectory \(y\), and \(a\in\mathbb{R}^m\) such that \(h(a)\in\mathcal{X}_{\mathrm{ad}}\), define \[\mathcal{G}_{m,t}^{y}(a):=\mathcal{G}_t^y(h(a))\in\mathbb{R}^r.\]
Proposition 45 (Reduced sensing matrix for mobile sensor averages). Let \[\xi=\sum_{k=1}^m \alpha_k\phi_k \in V_m, \qquad \alpha=(\alpha_1,\dots,\alpha_m)^\top\in \mathbb{R}^m.\] For \(t\in[0,T]\) and \(a\in\mathbb{R}^m\) with \(h(a)\in\mathcal{X}_{\mathrm{ad}}\), define \[\label{eq:reduced-mobile-sensing-matrix} M_y(t,a)\in\mathbb{R}^{r\times m}, \qquad [M_y(t,a)]_{jk} := \operatorname{Cov}_{\rho_{h(a)}}(\kappa(\cdot-y_j(t)),\phi_k).\qquad{(19)}\] Then \[\label{eq:reduced-mobile-differential} D\mathcal{G}_{m,t}^{y}(a)\,\alpha = M_y(t,a)\alpha.\qquad{(20)}\] Equivalently, \[J_{t,h(a)}^y\xi=M_y(t,a)\alpha.\] Hence \[\label{eq:reduced-mobile-observability-energy} \mathfrak j_{t,h(a)}^{y}(\xi,\xi) = \alpha^\top M_y(t,a)^\top M_y(t,a)\alpha.\qquad{(21)}\]
Proof. By Proposition 17, \[J_{t,h(a)}^y\xi = \Bigl( \operatorname{Cov}_{\rho_{h(a)}}(\kappa(\cdot-y_1(t)),\xi),\dots, \operatorname{Cov}_{\rho_{h(a)}}(\kappa(\cdot-y_r(t)),\xi) \Bigr).\] Substituting \(\xi=\sum_{k=1}^m \alpha_k\phi_k\) and using bilinearity of covariance gives \[\operatorname{Cov}_{\rho_{h(a)}}(\kappa(\cdot-y_j(t)),\xi) = \sum_{k=1}^m \alpha_k \operatorname{Cov}_{\rho_{h(a)}}(\kappa(\cdot-y_j(t)),\phi_k) = \sum_{k=1}^m [M_y(t,a)]_{jk}\alpha_k,\] which proves ?? . The identity ?? follows from Definition 10. ◻
The reduced joint transport–observability form is therefore the ambient form written in coordinates.
Proposition 46 (Pathwise reduced transport–observability). Let \[a\in \mathcal{A}_{\mathrm{ad}}^{(m)}, \qquad \alpha\in H^1(0,T;\mathbb{R}^m),\] and define the associated perturbation \[\xi_\alpha(t):=\sum_{k=1}^m \alpha_k(t)\phi_k \in V_m.\] Then \[\label{eq:pathwise-reduced-joint-form} Q_y[h(a)](\xi_\alpha,\xi_\alpha) = \int_0^T \alpha(t)^\top M_y(t,a(t))^\top M_y(t,a(t))\alpha(t)\,\mathrm{d}t + \int_0^T \dot{\alpha}(t)^\top H(a(t))\dot{\alpha}(t)\,\mathrm{d}t.\qquad{(22)}\]
Assume there exists \(\kappa_m>0\) such that for every \(a\in \mathcal{A}_{\mathrm{ad}}^{(m)}\) and every \(\alpha\in H^1(0,T;\mathbb{R}^m)\), \[\label{eq:pathwise-reduced-coercivity} \int_0^T \alpha(t)^\top M_y(t,a(t))^\top M_y(t,a(t))\alpha(t)\,\mathrm{d}t + \int_0^T \dot{\alpha}(t)^\top H(a(t))\dot{\alpha}(t)\,\mathrm{d}t \ge \kappa_m \int_0^T |\alpha(t)|^2\,\mathrm{d}t.\qquad{(23)}\] If, moreover, there exists \(C_m>0\) such that \[\label{eq:reduced-coefficient-norm-equivalence} \left\|\sum_{k=1}^m \beta_k \phi_k\right\|_{L^2(\nu_0)}^2 \le C_m |\beta|^2 \qquad\text{for all } \beta\in\mathbb{R}^m,\qquad{(24)}\] then Assumption 5 holds on the reduced class with \[\kappa=\frac{\kappa_m}{C_m}.\]
Proof. By Proposition 44, \[\mathfrak g_{h(a(t))}\bigl(\dot{\xi}_\alpha(t),\dot{\xi}_\alpha(t)\bigr) = \dot{\alpha}(t)^\top H(a(t))\dot{\alpha}(t).\] By Proposition 45, \[\mathfrak j_{t,h(a(t))}^{y}\bigl(\xi_\alpha(t),\xi_\alpha(t)\bigr) = \alpha(t)^\top M_y(t,a(t))^\top M_y(t,a(t))\alpha(t).\] Substituting these identities into Definition 12 gives ?? .
If ?? holds, then \[Q_y[h(a)](\xi_\alpha,\xi_\alpha) \ge \kappa_m \int_0^T |\alpha(t)|^2\,\mathrm{d}t.\] On the other hand, ?? gives \[\|\xi_\alpha(t)\|_{L^2(\nu_0)}^2 = \left\|\sum_{k=1}^m \alpha_k(t)\phi_k\right\|_{L^2(\nu_0)}^2 \le C_m |\alpha(t)|^2\] for a.e.\(t\in[0,T]\). Integrating in time yields \[\|\xi_\alpha\|_{L^2(0,T;L^2(\nu_0))}^2 \le C_m \int_0^T |\alpha(t)|^2\,\mathrm{d}t.\] Therefore \[Q_y[h(a)](\xi_\alpha,\xi_\alpha) \ge \frac{\kappa_m}{C_m} \|\xi_\alpha\|_{L^2(0,T;L^2(\nu_0))}^2,\] which is exactly Assumption 5 on the reduced class. ◻
Finally, the ambient inverse functional restricts to coefficient space.
Corollary 6 (Reduced inverse functional). Let \(a\in \mathcal{A}_{\mathrm{ad}}^{(m)}\), and set \[h(t):=h(a(t)).\] Then the ambient inverse functional \(\mathcal{I}_{\lambda,\mu,\gamma}\) takes the form \[\begin{align} \label{eq:reduced-inverse-functional} \mathcal{I}_{\lambda,\mu,\gamma}[a] & = \frac{1}{2}\int_0^T \|\mathcal{G}_{m,t}^{y}(a(t))-d(t)\|_{\mathbb{R}^r}^2\,\mathrm{d}t + \frac{\lambda}{2}\int_0^T \dot{a}(t)^\top H(a(t))\dot{a}(t)\,\mathrm{d}t \notag \\ & \quad + \frac{\mu}{2}\int_0^T \Bigl( \|h(a(t))\|_{L^2(\nu_0)}^2 + \|\dot{h}(t)\|_{L^2(\nu_0)}^2 \Bigr)\,\mathrm{d}t + \frac{\gamma}{2}\int_0^T \|h(a(t))\|_{H^s(\Omega)}^2\,\mathrm{d}t. \end{align}\tag{40}\] Because \(V_m\) is finite dimensional, all norms on \(V_m\) are equivalent. In particular, the last three terms are equivalent to Euclidean quadratic penalties in \(a(\cdot)\) and \(\dot{a}(\cdot)\).
Proof. Substitute \(h(t)=h(a(t))\) into Definition 13. The data term becomes \[\|\mathcal{G}_t^y(h(a(t)))-d(t)\|_{\mathbb{R}^r}^2 = \|\mathcal{G}_{m,t}^{y}(a(t))-d(t)\|_{\mathbb{R}^r}^2,\] and the transport term is identified by Proposition 44. Since \[h(a(t))=\sum_{k=1}^m a_k(t)\phi_k, \qquad \dot{h}(t)=\sum_{k=1}^m \dot{a}_k(t)\phi_k,\] the remaining terms are simply the pullbacks of ambient norms to the finite-dimensional space \(V_m\), hence are equivalent to Euclidean quadratic forms in \(a(t)\) and \(\dot{a}(t)\). ◻
Remark 47. The finite-dimensional reduction removes the spatial compactness issue that motivated the \(\gamma\)-term in the ambient existence theory. Thus \(\gamma\) may be set to zero on \(V_m\). The temporal \(L^2\)-regularization remains useful under partial observability, because it provides a priori control of \(\dot{a}\), which is exactly what bounds the transport contribution in the reduced joint form.
We now turn from abstract reduction to constructive observability. The first step is local: a sufficiently narrow bump sensor detects any fixed nonzero direction under mild regularity assumptions.
Proposition 48 (Lebesgue-point criterion for detection by localized bump sensors). Assume \(\Omega\subset\mathbb{R}^d\) is open and bounded, and let \(h\in\mathcal{X}_{\mathrm{ad}}\) be such that \(\rho_h\) admits a Lebesgue density \[r_h:=\frac{\mathrm{d}\rho_h}{\mathrm{d}x}\in L^\infty(\Omega).\] Let \(\kappa_0\in C_c^\infty(\mathbb{R}^d)\) be nonnegative, supported in \(B(0,1)\), and normalized by \[\int_{\mathbb{R}^d}\kappa_0(z)\,\mathrm{d}z=1.\] For \(\delta>0\), define \[\kappa_\delta(z):=\delta^{-d}\kappa_0(z/\delta).\] Fix \(\xi\in L^2_0(\nu_0)\), and set \[m_h(\xi):=\mathbb{E}_{\rho_h}[\xi], \qquad g_{h,\xi}(x):=r_h(x)\bigl(\xi(x)-m_h(\xi)\bigr)\in L^1(\Omega).\] Extend \(g_{h,\xi}\) by zero outside \(\Omega\).
Let \(y\in\Omega\) satisfy \[\mathop{\mathrm{dist}}(y,\partial\Omega)>0,\] and suppose that \(y\) is a Lebesgue point of \(g_{h,\xi}\). Then for every \(0<\delta<\mathop{\mathrm{dist}}(y,\partial\Omega)\), \[\label{eq:covariance-as-convolution} \operatorname{Cov}_{\rho_h}\!\bigl(\kappa_\delta(\cdot-y),\xi\bigr) = (\kappa_\delta * g_{h,\xi})(y),\qquad{(25)}\] and moreover \[\label{eq:lebesgue-point-detection-limit} \lim_{\delta\downarrow 0} \operatorname{Cov}_{\rho_h}\!\bigl(\kappa_\delta(\cdot-y),\xi\bigr) = g_{h,\xi}(y).\qquad{(26)}\] In particular, if \(g_{h,\xi}(y)\neq 0\), then there exists \(\delta_0>0\) such that \[\operatorname{Cov}_{\rho_h}\!\bigl(\kappa_\delta(\cdot-y),\xi\bigr)\neq 0 \qquad \forall\,0<\delta<\delta_0.\]
Proof. Since \(\int \kappa_\delta(\cdot-y)\,\mathrm{d}x=1\), one has \[\operatorname{Cov}_{\rho_h}\!\bigl(\kappa_\delta(\cdot-y),\xi\bigr) = \int_\Omega \kappa_\delta(x-y)\,\xi(x)\,\mathrm{d}\rho_h(x) - \Bigl(\int_\Omega \kappa_\delta(x-y)\,\mathrm{d}\rho_h(x)\Bigr)\,m_h(\xi).\] Using \(\mathrm{d}\rho_h(x)=r_h(x)\,\mathrm{d}x\), this becomes \[\operatorname{Cov}_{\rho_h}\!\bigl(\kappa_\delta(\cdot-y),\xi\bigr) = \int_\Omega \kappa_\delta(x-y)\,r_h(x)\bigl(\xi(x)-m_h(\xi)\bigr)\,\mathrm{d}x = \int_{\mathbb{R}^d}\kappa_\delta(y-x)\,g_{h,\xi}(x)\,\mathrm{d}x,\] which is exactly ?? .
Because \(y\) is a Lebesgue point of \(g_{h,\xi}\), the approximate-identity theorem implies \[(\kappa_\delta*g_{h,\xi})(y)\longrightarrow g_{h,\xi}(y) \qquad\text{as }\delta\downarrow 0,\] which proves ?? . If \(g_{h,\xi}(y)\neq 0\), then for all sufficiently small \(\delta\) the quantity \((\kappa_\delta*g_{h,\xi})(y)\) has the same sign as \(g_{h,\xi}(y)\), and in particular is nonzero. ◻
Corollary 7 (Single-direction detectability). Assume the hypotheses of Proposition 48, and suppose in addition that \(r_h>0\) almost everywhere in \(\Omega\). Let \(\xi\in L^2_0(\nu_0)\) be nonzero. Then there exist \(y\in\Omega\) and \(\delta_0>0\) such that \[\operatorname{Cov}_{\rho_h}\!\bigl(\kappa_\delta(\cdot-y),\xi\bigr)\neq 0 \qquad \forall\,0<\delta<\delta_0.\]
Proof. Since \(r_h\in L^\infty(\Omega)\) and \(\xi\in L^2(\nu_0)\), one has \(g_{h,\xi}\in L^1(\Omega)\). If \(g_{h,\xi}=0\) almost everywhere, then \[r_h(x)\bigl(\xi(x)-m_h(\xi)\bigr)=0 \qquad\text{for a.e. }x.\] Because \(r_h>0\) a.e., it follows that \(\xi=m_h(\xi)\) almost everywhere. Since \(\xi\in L^2_0(\nu_0)\), this forces \(\xi=0\) almost everywhere, a contradiction.
Hence \(g_{h,\xi}\) is not zero a.e. Its set of nonzero points has positive Lebesgue measure. By the Lebesgue differentiation theorem, almost every point of that set is a Lebesgue point of \(g_{h,\xi}\). Choose such a point \(y\in\Omega\) with \(g_{h,\xi}(y)\neq 0\) and \(\mathop{\mathrm{dist}}(y,\partial\Omega)>0\). Proposition 48 now yields the claim. ◻
Corollary 8 (Continuous-direction criterion). Assume the hypotheses of Proposition 48, and suppose in addition that \[r_h\in C(\overline{\Omega}), \qquad r_h(x)>0\;\text{ on }\overline{\Omega}, \qquad \xi\in C(\overline{\Omega})\cap L^2_0(\nu_0), \qquad \xi\not\equiv 0.\] Then there exist \(y\in\Omega\) and \(\delta_0>0\) such that \[\operatorname{Cov}_{\rho_h}\!\bigl(\kappa_\delta(\cdot-y),\xi\bigr)\neq 0 \qquad \forall\,0<\delta<\delta_0.\] Indeed, one may choose any \(y\in\Omega\) such that \[\xi(y)\neq \mathbb{E}_{\rho_h}[\xi].\]
Proof. Because \(\xi\) is continuous and nonzero with \(\xi\in L^2_0(\nu_0)\), it cannot be constant on \(\Omega\). Hence \(\xi-\mathbb{E}_{\rho_h}[\xi]\) is not identically zero, so there exists \(y\in\Omega\) such that \[\xi(y)\neq \mathbb{E}_{\rho_h}[\xi].\] For such \(y\), \[g_{h,\xi}(y)=r_h(y)\bigl(\xi(y)-\mathbb{E}_{\rho_h}[\xi]\bigr)\neq 0.\] Since both \(r_h\) and \(\xi\) are continuous, \(g_{h,\xi}\) is continuous, hence every point is a Lebesgue point of \(g_{h,\xi}\). The conclusion follows from Proposition 48. ◻
The previous subsection treats one direction at a time. We now show that on a fixed \(m\)-dimensional reduced subspace one can place finitely many stationary sensors so that the whole reduced state is visible at each time.
Lemma 4 (Evaluation points for linearly independent continuous functions). Let \(K\) be a compact Hausdorff space, and let \[f_0,\dots,f_m\in C(K)\] be linearly independent over \(\mathbb{R}\). Then there exist points \[y_0,\dots,y_m\in K\] such that the square matrix \[A_Y := \bigl(f_\ell(y_j)\bigr)_{j,\ell=0}^m \in\mathbb{R}^{(m+1)\times(m+1)}\] is invertible.
Proof. For \(y\in K\), define the evaluation vector \[v(y):=\bigl(f_0(y),\dots,f_m(y)\bigr)\in\mathbb{R}^{m+1}.\] If the span of \(\{v(y):y\in K\}\) had dimension strictly smaller than \(m+1\), then there would exist \(c=(c_0,\dots,c_m)\neq 0\) such that \[c\cdot v(y)=\sum_{\ell=0}^m c_\ell f_\ell(y)=0 \qquad\forall y\in K,\] contradicting the linear independence of \(f_0,\dots,f_m\). Hence \(\{v(y):y\in K\}\) spans \(\mathbb{R}^{m+1}\), so one can choose \(y_0,\dots,y_m\in K\) such that \(v(y_0),\dots,v(y_m)\) form a basis of \(\mathbb{R}^{m+1}\). Equivalently, the matrix \(A_Y\) is invertible. ◻
Theorem 49 (Uniform instantaneous reduced observability by fixed static sensors). Let \[V_m=\operatorname{span}\{\phi_1,\dots,\phi_m\}\subset C(K_{\mathrm{sens}})\cap L^2_0(\nu_0),\] where \(K_{\mathrm{sens}}\subset\Omega\) is a nonempty compact sensor-feasible region. Assume that the restrictions of \[1,\phi_1,\dots,\phi_m\] to \(K_{\mathrm{sens}}\) are linearly independent. Let \(U\subset\mathbb{R}^m\) be compact, and define \[h(a):=\sum_{k=1}^m a_k\phi_k, \qquad a\in U.\] Assume:
\(h(a)\in\mathcal{X}_{\mathrm{ad}}\) for every \(a\in U\);
for every \(a\in U\), the probability measure \(\rho_{h(a)}\) has a Lebesgue density \[r_a:=\frac{\mathrm{d}\rho_{h(a)}}{\mathrm{d}x}\in C(\overline{\Omega}),\] and \((a,x)\mapsto r_a(x)\) is continuous on \(U\times\overline{\Omega}\);
there exists \(c_\rho>0\) such that \[r_a(x)\ge c_\rho \qquad \forall (a,x)\in U\times K_{\mathrm{sens}};\]
\(\kappa_\delta(z)=\delta^{-d}\kappa_0(z/\delta)\), where \(\kappa_0\in C_c^\infty(\mathbb{R}^d)\) is nonnegative, supported in \(B(0,1)\), and \(\int_{\mathbb{R}^d}\kappa_0=1\).
Then there exist points \[y_0,\dots,y_m\in K_{\mathrm{sens}}\] and \(\delta_0>0\) such that for every \(0<\delta<\delta_0\) and every \(a\in U\), the static reduced sensing matrix \[M_Y^\delta(a)\in\mathbb{R}^{(m+1)\times m}, \qquad [M_Y^\delta(a)]_{jk} := \operatorname{Cov}_{\rho_{h(a)}}(\kappa_\delta(\cdot-y_j),\phi_k),\] has full column rank \(m\). Moreover, there exists \(\sigma_*>0\) such that \[\label{eq:uniform-static-gramian-lower-bound} M_Y^\delta(a)^\top M_Y^\delta(a)\succeq \sigma_* I_m \qquad \forall a\in U,\;\forall 0<\delta<\delta_0.\tag{41}\]
Proof. Apply Lemma 4 to the functions \[f_0:=1,\qquad f_k:=\phi_k,\quad k=1,\dots,m,\] on the compact space \(K_{\mathrm{sens}}\). This yields points \(y_0,\dots,y_m\in K_{\mathrm{sens}}\) such that the augmented evaluation matrix \[A_Y := \begin{pmatrix} 1 & \phi_1(y_0) & \cdots & \phi_m(y_0) \\ \vdots & \vdots & & \vdots \\ 1 & \phi_1(y_m) & \cdots & \phi_m(y_m) \end{pmatrix} \in\mathbb{R}^{(m+1)\times(m+1)}\] is invertible.
For \(a\in U\), define \[c(a):= \bigl(c_1(a),\dots,c_m(a)\bigr)^\top, \qquad c_k(a):=\mathbb{E}_{\rho_{h(a)}}[\phi_k].\] Also define \[X_Y:=\bigl(\phi_k(y_j)\bigr)_{j=0,\dots,m}^{k=1,\dots,m}\in\mathbb{R}^{(m+1)\times m}, \qquad \mathbf{1}:=(1,\dots,1)^\top\in\mathbb{R}^{m+1},\] and \[D_Y(a):=\operatorname{diag}(r_a(y_0),\dots,r_a(y_m)).\] Consider the ideal point-sensor matrix \[\widetilde{M}_Y(a):=D_Y(a)\bigl(X_Y-\mathbf{1}\,c(a)^\top\bigr).\]
We first show that \(\widetilde{M}_Y(a)\) has full column rank \(m\) for every \(a\in U\). Suppose \[\widetilde{M}_Y(a)\alpha=0 \qquad\text{for some }\alpha\in\mathbb{R}^m.\] Since \(D_Y(a)\) is invertible by the uniform positivity of \(r_a\) on \(K_{\mathrm{sens}}\), this implies \[\bigl(X_Y-\mathbf{1}\,c(a)^\top\bigr)\alpha=0,\] hence \[X_Y\alpha=(c(a)^\top\alpha)\mathbf{1}.\] Therefore \[A_Y \binom{-c(a)^\top\alpha}{\alpha} = 0.\] Because \(A_Y\) is invertible, we conclude \(\alpha=0\). Thus \(\operatorname{rank}\widetilde{M}_Y(a)=m\).
Next, the map \(a\mapsto \widetilde{M}_Y(a)\) is continuous on \(U\): the values \(r_a(y_j)\) are continuous in \(a\), and the averages \(c_k(a)=\int \phi_k\,\mathrm{d}\rho_{h(a)}\) are continuous by dominated convergence. Hence the smallest singular value of \(\widetilde{M}_Y(a)\) is a positive continuous function on the compact set \(U\), so there exists \(\sigma_*>0\) such that \[\widetilde{M}_Y(a)^\top \widetilde{M}_Y(a)\succeq 2\sigma_* I_m \qquad \forall a\in U.\]
It remains to compare the actual bump-sensor matrix \(M_Y^\delta(a)\) with \(\widetilde{M}_Y(a)\). Because \(K_{\mathrm{sens}}\subset\Omega\) is compact, there exists \[d_*:=\mathop{\mathrm{dist}}(K_{\mathrm{sens}},\partial\Omega)>0.\] For \(0<\delta<d_*\), the support of \(\kappa_\delta(\cdot-y_j)\) stays inside \(\Omega\) for every \(j\). For each \(j,k\), \[[M_Y^\delta(a)]_{jk} = \int_\Omega \kappa_\delta(x-y_j)\,r_a(x)\bigl(\phi_k(x)-c_k(a)\bigr)\,\mathrm{d}x.\] Define \[g_{a,k}(x):=r_a(x)\bigl(\phi_k(x)-c_k(a)\bigr).\] Since \((a,x)\mapsto r_a(x)\) is continuous on \(U\times\overline{\Omega}\), each \(\phi_k\) is continuous on \(K_{\mathrm{sens}}\), and \(a\mapsto c_k(a)\) is continuous, the family \(\{g_{a,k}:a\in U\}\) is uniformly continuous in a neighborhood of \(K_{\mathrm{sens}}\). Hence the approximate-identity property yields \[\sup_{a\in U}\bigl|[M_Y^\delta(a)]_{jk}-g_{a,k}(y_j)\bigr|\longrightarrow 0 \qquad\text{as }\delta\downarrow 0.\] But \(g_{a,k}(y_j)=[\widetilde{M}_Y(a)]_{jk}\), so \[\sup_{a\in U}\|M_Y^\delta(a)-\widetilde{M}_Y(a)\|\longrightarrow 0 \qquad\text{as }\delta\downarrow 0.\] Choosing \(\delta_0>0\) sufficiently small, we may arrange \[M_Y^\delta(a)^\top M_Y^\delta(a)\succeq \sigma_* I_m \qquad \forall a\in U,\;\forall 0<\delta<\delta_0.\] This proves both full column rank and 41 . ◻
Remark 50. Theorem 49 places the reduced problem in a fully observable regime: with \(m+1\) suitably chosen stationary sensors, the reduced sensing matrix is uniformly coercive on the whole reduced family \(U\). The next subsection shows that one may trade this instantaneous full observability for time-aggregated observability by moving a single sensor through those same locations.
We now return to the pathwise viewpoint. In finite dimensions, the correct design target for mobile sensing is the averaged reduced Gramian along the reduced path.
Proposition 51 (Time-averaged coercivity on constant reduced modes). Fix \(a\in \mathcal{A}_{\mathrm{ad}}^{(m)}\) and a measurable sensor trajectory \(y\). Define the averaged reduced Gramian \[\label{eq:averaged-reduced-gramian} \overline{G}_y[a] := \int_0^T M_y(t,a(t))^\top M_y(t,a(t))\,\mathrm{d}t.\qquad{(27)}\] Then the following are equivalent:
\(\overline{G}_y[a]\) is positive definite;
there exists \(\kappa_{\mathrm{obs}}^{(m)}>0\) such that \[\int_0^T \|J_{t,h(a(t))}^{y}\xi_0\|_{\mathbb{R}^r}^2\,\mathrm{d}t \ge \kappa_{\mathrm{obs}}^{(m)}\|\xi_0\|_{L^2(\nu_0)}^2 \qquad \forall \xi_0\in V_m;\]
\(\ker Q_y[h(a)]\cap V_m=\{0\}\), where \(V_m\) is identified with time-constant perturbations.
Proof. Let \[B_m:=\bigl(\langle \phi_k,\phi_\ell\rangle_{L^2(\nu_0)}\bigr)_{k,\ell=1}^m \in \mathbb{R}^{m\times m}.\] Since \(\{\phi_1,\dots,\phi_m\}\) is a basis of \(V_m\), \(B_m\) is symmetric positive definite.
Write \[\xi_0=\sum_{k=1}^m c_k\phi_k, \qquad c=(c_1,\dots,c_m)^\top\in\mathbb{R}^m.\] Then \[\label{eq:xi0-norm-gram} \|\xi_0\|_{L^2(\nu_0)}^2 = c^\top B_m c.\tag{42}\] By Proposition 45, \[J_{t,h(a(t))}^{y}\xi_0 = M_y(t,a(t))\,c,\] and therefore \[\begin{align} \int_0^T \|J_{t,h(a(t))}^{y}\xi_0\|_{\mathbb{R}^r}^2\,\mathrm{d}t & = \int_0^T c^\top M_y(t,a(t))^\top M_y(t,a(t))\,c\,\mathrm{d}t \notag \\ & = c^\top \overline{G}_y[a]\,c. \label{eq:obs-energy-gram} \end{align}\tag{43}\]
Assume \((1)\). Since both \(B_m\) and \(\overline{G}_y[a]\) are symmetric positive definite, the matrix \[A_m:=B_m^{-1/2}\,\overline{G}_y[a]\,B_m^{-1/2}\] is symmetric positive definite. Hence \[\lambda_*:=\lambda_{\min}(A_m)>0.\] Using 42 and 43 , \[c^\top \overline{G}_y[a]\,c = (B_m^{1/2}c)^\top A_m (B_m^{1/2}c) \ge \lambda_*\,|B_m^{1/2}c|^2 = \lambda_*\,\|\xi_0\|_{L^2(\nu_0)}^2.\] Thus \((2)\) holds.
Conversely, if \((2)\) holds and \(\overline{G}_y[a]\) were not positive definite, there would exist \(c\neq 0\) such that \[c^\top \overline{G}_y[a]\,c=0.\] Then the associated \(\xi_0=\sum_{k=1}^m c_k\phi_k\) is nonzero, while 43 contradicts \((2)\). Hence \((1)\) follows.
Finally, a time-constant perturbation \(\xi(t)\equiv \xi_0\in V_m\) has \(\dot{\xi}=0\), so its transport contribution vanishes. Therefore \[Q_y[h(a)](\xi_0,\xi_0) = \int_0^T \|J_{t,h(a(t))}^{y}\xi_0\|_{\mathbb{R}^r}^2\,\mathrm{d}t.\] Thus \[\ker Q_y[h(a)]\cap V_m = \left\{ \xi_0\in V_m: \int_0^T \|J_{t,h(a(t))}^{y}\xi_0\|_{\mathbb{R}^r}^2\,\mathrm{d}t=0 \right\},\] which shows \((2)\Longleftrightarrow(3)\). ◻
The previous proposition identifies the reduced design target: one should choose the sensor trajectory so that \(\overline{G}_y[a]\) is positive definite. In the single-sensor setting, this means that the time-average of the rank-one reduced observation matrices must span the whole reduced coefficient space.
For the remainder of this subsection we consider a single moving sensor, so \(r=1\). We write \[\mathcal{Y}_{\mathrm{ad}}^{(1)} := \left\{ y\in H^1(0,T;\mathbb{R}^d): y(t)\in K_{\mathrm{sens}} \text{ for a.e.\;}t\in[0,T] \right\}.\] We also work on the global reduced admissible class \[\mathcal{A}_m(U) := \left\{ a\in H^1(0,T;\mathbb{R}^m): a(t)\in U \text{ for a.e.\;}t\in[0,T] \right\},\] where \(U\subset \mathbb{R}^m\) is compact.
Remark 52. The class \(\mathcal{A}_m(U)\) is the natural global reduced admissible class for the variational theory, since Remark 58 shows that the reduced regularization is equivalent to a Euclidean \(H^1\)-in-time penalty on the coefficient path.
To pass from the static observability theorem to a single moving sensor, we need a mild geometric assumption on the feasible sensor region: any two feasible positions can be joined by an \(H^1\)-curve that remains feasible.
Theorem 53 (Path-dependent single-sensor trajectory design). Let \(U\subset\mathbb{R}^m\) be compact, and assume the hypotheses of Theorem 49. Assume in addition that the feasible sensor region \(K_{\mathrm{sens}}\subset \Omega\) is \(H^1\)-path connected in the sense that for every \(p,q\in K_{\mathrm{sens}}\) there exists \[\gamma_{p,q}\in H^1([0,1];\mathbb{R}^d)\] such that \[\gamma_{p,q}(0)=p,\qquad \gamma_{p,q}(1)=q,\qquad \gamma_{p,q}(s)\in K_{\mathrm{sens}} \quad\forall s\in[0,1].\]
Fix any \[a\in \mathcal{A}_m(U).\] Then there exists a single-sensor trajectory \[y[a]\in \mathcal{Y}_{\mathrm{ad}}^{(1)}\] such that the averaged reduced Gramian \[\overline{G}_{y[a]}[a] = \int_0^T M_{y[a]}(t,a(t))^\top M_{y[a]}(t,a(t))\,\mathrm{d}t\] is positive definite.
Proof. Choose once and for all a kernel width \(0<\delta<\delta_0\), where \(\delta_0\) is the radius given by Theorem 49. We suppress the dependence on \(\delta\) in the notation.
By Theorem 49, there exist points \[z_0,\dots,z_m\in K_{\mathrm{sens}}\] and a constant \(\sigma_*>0\) such that for every \(b\in U\), \[\label{eq:static-uniform-coercivity-proof} \sum_{j=0}^m M_{z_j}(b)^\top M_{z_j}(b) \succeq \sigma_* I_m,\tag{44}\] where \(M_{z_j}(b)\in \mathbb{R}^{1\times m}\) denotes the reduced sensing row for a single static sensor at the location \(z_j\).
For each \(j=0,\dots,m-1\), choose an \(H^1\)-curve in \(K_{\mathrm{sens}}\) connecting \(z_j\) to \(z_{j+1}\), and choose an \(H^1\)-curve connecting \(z_m\) back to \(z_0\). Because the set of curves is finite, we may regard these as fixed transition templates.
For the fixed path \(a\in \mathcal{A}_m(U)\), define \[G_j(t):=M_{z_j}(a(t))^\top M_{z_j}(a(t)) \in \mathbb{R}^{m\times m}, \qquad G_\Sigma(t):=\sum_{j=0}^m G_j(t).\] Since \(a\in H^1(0,T;\mathbb{R}^m)\subset C([0,T];\mathbb{R}^m)\) and \(b\mapsto M_{z_j}(b)\) is continuous on \(U\), each \(G_j\) is continuous on \([0,T]\), hence uniformly continuous. Also, by 44 , \[\label{eq:Gsigma-lower-bound} G_\Sigma(t)\succeq \sigma_* I_m \qquad\forall t\in[0,T].\tag{45}\]
Let \(q:=m+1\), and choose \(\eta>0\) so small that \[q\eta \le \frac{\sigma_*}{2}.\] By uniform continuity, there exists a partition \[0=t_0<t_1<\cdots<t_N=T\] such that on each interval \(I_n:=[t_{n-1},t_n]\), \[\label{eq:Gj-small-variation} \|G_j(t)-G_j(s)\|\le \eta \qquad \forall s,t\in I_n,\;\forall j=0,\dots,m.\tag{46}\]
We now construct \(y[a]\) interval by interval. On each \(I_n\), first let the sensor dwell at \(z_0,\dots,z_m\) in that order, each for a time interval of length \[\frac{|I_n|}{2q}.\] Use the remaining half of \(I_n\) to traverse the chosen transition curves between successive locations, scaled in time so as to fit. This gives a piecewise \(H^1\) path on each \(I_n\), and by construction the endpoint of the last transition on \(I_n\) matches the initial position of the first dwell on \(I_{n+1}\). Hence the resulting global path \(y[a]\) belongs to \(\mathcal{Y}_{\mathrm{ad}}^{(1)}\).
Let \(D_{n,j}\subset I_n\) denote the dwell interval on which \(y[a](t)=z_j\). Since the transition contributions are positive semidefinite, we have \[\overline{G}_{y[a]}[a] \succeq \sum_{n=1}^N \sum_{j=0}^m \int_{D_{n,j}} G_j(t)\,\mathrm{d}t.\] Fix \(n\in\{1,\dots,N\}\) and choose any \(s_n\in I_n\). By 46 , \[G_j(t)\succeq G_j(s_n)-\eta I_m \qquad \forall t\in I_n,\;\forall j=0,\dots,m.\] Therefore \[\begin{align} \sum_{j=0}^m \int_{D_{n,j}} G_j(t)\,\mathrm{d}t & \succeq \sum_{j=0}^m \frac{|I_n|}{2q}\bigl(G_j(s_n)-\eta I_m\bigr) \\ & = \frac{|I_n|}{2q} \left( G_\Sigma(s_n)-q\eta I_m \right). \end{align}\] Using 45 and the choice of \(\eta\), \[G_\Sigma(s_n)-q\eta I_m \succeq \frac{\sigma_*}{2} I_m.\] Hence \[\sum_{j=0}^m \int_{D_{n,j}} G_j(t)\,\mathrm{d}t \succeq \frac{|I_n|\,\sigma_*}{4q} I_m.\] Summing over \(n\) gives \[\overline{G}_{y[a]}[a] \succeq \sum_{n=1}^N \frac{|I_n|\,\sigma_*}{4q} I_m = \frac{T\,\sigma_*}{4(m+1)} I_m.\] Thus \(\overline{G}_{y[a]}[a]\) is positive definite. ◻
Remark 54 (Path-dependent design). The trajectory \(y[a]\) in Theorem 53 is allowed to depend on the reduced path \(a\). Thus, when the theorem is applied to a ground-truth path \(a^\dagger\), the resulting sensor trajectory may depend on that ground truth. At this point we are not claiming that a single trajectory can be designed a priori, independently of the unknown reduced path, so as to give the same averaged-Gramian positivity uniformly over a whole admissible class.
Once a path \(a^\dagger\) and an associated trajectory \(y^\dagger\) have been fixed, positivity of the averaged reduced Gramian upgrades immediately to full pathwise coercivity of the reduced joint form. This is the finite-dimensional mechanism by which the transport term controls oscillatory modes and the averaged observation term controls the surviving constant mode.
Theorem 55 (Fixed-path reduced coercivity from a positive averaged Gramian). Let \(U\subset\mathbb{R}^m\) be compact, let \[a^\dagger\in \mathcal{A}_m(U), \qquad y^\dagger\in \mathcal{Y}_{\mathrm{ad}}^{(1)},\] and assume:
the reduced transport tensor \(H(a)\) is continuous on \(U\) and uniformly positive definite: \[\label{eq:uniform-transport-lower-bound} H(a)\succeq c_{\mathrm{tr}} I_m \qquad\text{for all }a\in U\tag{47}\] for some \(c_{\mathrm{tr}}>0\);
the reduced sensing row along the fixed pair \((a^\dagger,y^\dagger)\) satisfies \[t\longmapsto M_{y^\dagger}(t,a^\dagger(t)) \in L^\infty(0,T;\mathbb{R}^{1\times m});\]
the averaged reduced Gramian along the fixed path is positive definite: \[\label{eq:positive-averaged-gramian-reference} \overline{G}_{y^\dagger}[a^\dagger] = \int_0^T M_{y^\dagger}(t,a^\dagger(t))^\top M_{y^\dagger}(t,a^\dagger(t))\,\mathrm{d}t \succ 0.\tag{48}\]
Then there exists \(\kappa_m^\dagger>0\) such that \[\label{eq:fixed-path-reduced-coercivity} \int_0^T \alpha(t)^\top M_{y^\dagger}(t,a^\dagger(t))^\top M_{y^\dagger}(t,a^\dagger(t)) \alpha(t)\,\mathrm{d}t + \int_0^T \dot{\alpha}(t)^\top H(a^\dagger(t))\,\dot{\alpha}(t)\,\mathrm{d}t \ge \kappa_m^\dagger \int_0^T |\alpha(t)|^2\,\mathrm{d}t\tag{49}\] for every \(\alpha\in H^1(0,T;\mathbb{R}^m)\).
Proof. Suppose the conclusion fails. Then there exists a sequence \[\alpha_n\in H^1(0,T;\mathbb{R}^m), \qquad \|\alpha_n\|_{L^2(0,T;\mathbb{R}^m)}=1,\] such that \[\label{eq:fixed-path-coercivity-contradiction} \int_0^T \alpha_n^\top M_{y^\dagger}(t,a^\dagger(t))^\top M_{y^\dagger}(t,a^\dagger(t)) \alpha_n\,\mathrm{d}t + \int_0^T \dot{\alpha}_n^\top H(a^\dagger(t))\,\dot{\alpha}_n\,\mathrm{d}t \longrightarrow 0.\tag{50}\]
By 47 , \[c_{\mathrm{tr}}\int_0^T |\dot{\alpha}_n(t)|^2\,\mathrm{d}t \le \int_0^T \dot{\alpha}_n(t)^\top H(a^\dagger(t))\,\dot{\alpha}_n(t)\,\mathrm{d}t \longrightarrow 0,\] so \[\|\dot{\alpha}_n\|_{L^2(0,T;\mathbb{R}^m)}\to 0.\] Let \[\bar\alpha_n:=\frac{1}{T}\int_0^T \alpha_n(t)\,\mathrm{d}t.\] By the Poincaré–Wirtinger inequality, \[\|\alpha_n-\bar\alpha_n\|_{L^2(0,T;\mathbb{R}^m)} \le C_P \|\dot{\alpha}_n\|_{L^2(0,T;\mathbb{R}^m)} \to 0.\] Since \((\bar\alpha_n)\) is bounded in \(\mathbb{R}^m\), after passing to a subsequence we may assume \[\bar\alpha_n\to c\in\mathbb{R}^m.\] Therefore \[\alpha_n\to c \qquad\text{strongly in }L^2(0,T;\mathbb{R}^m),\] where \(c\) is identified with the corresponding time-constant function on \([0,T]\). Since \(\|\alpha_n\|_{L^2}=1\), we have \[\|c\|_{L^2(0,T;\mathbb{R}^m)}=1,\] so \(c\neq 0\).
Set \[M^\dagger(t):=M_{y^\dagger}(t,a^\dagger(t)).\] Because \(M^\dagger\in L^\infty(0,T;\mathbb{R}^{1\times m})\), \[M^\dagger \alpha_n \to M^\dagger c \qquad\text{strongly in }L^2(0,T;\mathbb{R}).\] Passing to the limit in the first term of 50 gives \[0 = \int_0^T |M^\dagger(t)c|^2\,\mathrm{d}t = c^\top \overline{G}_{y^\dagger}[a^\dagger]\,c.\] This contradicts 48 , since \(\overline{G}_{y^\dagger}[a^\dagger]\succ 0\) and \(c\neq 0\). Hence 49 holds. ◻
Remark 56. Theorem 53 is global in the reduced path class \(\mathcal{A}_m(U)\), but the resulting trajectory \(y[a]\) depends on the chosen path \(a\). Theorem 55 then gives a global coercivity statement along that fixed path, rather than a neighborhood theorem uniform over nearby reduced paths. This is exactly the form needed when the sensor design is allowed to depend on the true reduced path.
We illustrate the preceding theory on a simple two-dimensional reduction with a single moving localized sensor.
Let \[\Omega=(0,1)^2, \qquad \nu_0 = \mathrm{d}x,\] and consider the two-dimensional reduced space \[V_2:=\operatorname{span}\{\phi_1,\phi_2\}, \qquad \phi_1(x)=x_1-\tfrac12, \qquad \phi_2(x)=x_2-\tfrac12.\] For \(a=(a_1,a_2)\in\mathbb{R}^2\), define \[h(a):=a_1\phi_1+a_2\phi_2.\] The associated probability density is \[\rho_{h(a)}(x) = \frac{\exp(h(a)(x))}{\int_\Omega \exp(h(a)(z))\,\mathrm{d}z}.\]
Choose a nonzero reference reduced path \[a^\dagger(t) = \bigl(\bar a_1+\varepsilon \cos(2\pi t/T),\; \bar a_2+\varepsilon \sin(2\pi t/T)\bigr), \qquad t\in[0,T],\] where \(\bar a=(\bar a_1,\bar a_2)\neq 0\) and \(\varepsilon>0\) is small. Thus \[h^\dagger(t,x) = \bigl(\bar a_1+\varepsilon \cos(2\pi t/T)\bigr)\Bigl(x_1-\tfrac12\Bigr) + \bigl(\bar a_2+\varepsilon \sin(2\pi t/T)\bigr)\Bigl(x_2-\tfrac12\Bigr).\]
Let \(\kappa_0\in C_c^\infty(\mathbb{R}^2)\) be nonnegative, radial, supported in \(B(0,1)\), and normalized by \(\int_{\mathbb{R}^2}\kappa_0(z)\,\mathrm{d}z=1\). For \(\delta>0\), define \[\kappa_\delta(x):=\delta^{-2}\kappa_0(x/\delta).\] Given a sensor position \(y\in \Omega\), define \[M(y,a) := \bigl( \operatorname{Cov}_{\rho_{h(a)}}(\kappa_\delta(\,\cdot-y),\phi_1),\; \operatorname{Cov}_{\rho_{h(a)}}(\kappa_\delta(\,\cdot-y),\phi_2) \bigr) \in \mathbb{R}^{1\times 2}.\] For a sensor trajectory \(y:[0,T]\to\Omega\), write \[M_y(t,a):=M(y(t),a).\]
We choose two dwell locations, \[p^{(1)}=\Bigl(\tfrac78,\tfrac12\Bigr), \qquad p^{(2)}=\Bigl(\tfrac12,\tfrac78\Bigr),\] and let \(y\in H^1(0,T;\mathbb{R}^2)\) be any trajectory such that \[y(t)=p^{(1)} \quad\text{for }t\in\Bigl[0,\frac{T}{2}-\tau\Bigr], \qquad y(t)=p^{(2)} \quad\text{for }t\in\Bigl[\frac{T}{2}+\tau,T\Bigr],\] with a smooth transition on \([T/2-\tau,T/2+\tau]\), where \(0<\tau<T/4\).
The key point is that the two dwell locations favor the two reduced directions differently. At the unperturbed state \(a=0\) and in the point-sensor limit \(\delta\to 0\), \[M(p^{(1)},0)\approx \bigl(\phi_1(p^{(1)}),\phi_2(p^{(1)})\bigr) = \Bigl(\tfrac38,0\Bigr),\] and \[M(p^{(2)},0)\approx \bigl(\phi_1(p^{(2)}),\phi_2(p^{(2)})\bigr) = \Bigl(0,\tfrac38\Bigr).\] Hence the ideal averaged reduced Gramian is \[\overline{G}_{\mathrm{ideal}} = \int_0^{T/2} \begin{pmatrix}\tfrac38\\[1mm]0\end{pmatrix} \begin{pmatrix}\tfrac38 & 0\end{pmatrix}\,\mathrm{d}t + \int_{T/2}^{T} \begin{pmatrix}0\\[1mm]\tfrac38\end{pmatrix} \begin{pmatrix}0 & \tfrac38\end{pmatrix}\,\mathrm{d}t = \frac{9T}{128}\,I_2,\] which is positive definite.
Since \(a\mapsto \rho_{h(a)}\) and \((y,a)\mapsto M(y,a)\) are continuous on compact sets, the actual averaged reduced Gramian stays positive definite for small perturbations of this ideal configuration.
Proposition 57. There exist parameters \[\bar a\in\mathbb{R}^2\setminus\{0\}, \qquad \varepsilon>0, \qquad \delta>0, \qquad \tau\in(0,T/4),\] such that the above nonzero reference path \(a^\dagger\) and moving sensor trajectory \(y\) satisfy \[\overline{G}_y[a^\dagger]\succ 0.\] Consequently, if \(U\subset\mathbb{R}^2\) is a compact neighborhood of the image of \(a^\dagger\) on which the reduced transport tensor \(H(a)\) is continuous and uniformly positive definite, then the conclusion of Theorem 55 applies along \(a^\dagger\).
Proof. The idealized Gramian computed above is \(\frac{9T}{128}I_2\), hence positive definite. By continuity of the reduced sensing map with respect to the state, sensor location, and kernel width, the actual averaged Gramian is a small perturbation of this ideal matrix whenever \(|\bar a|\), \(\varepsilon\), \(\delta\), and \(\tau\) are sufficiently small. Positive definiteness is stable under small perturbations, so \(\overline{G}_y[a^\dagger]\succ 0\). The final claim then follows from Theorem 55. ◻
The finite-dimensional theory is a direct restriction of the ambient theory, but its recovery consequences are stronger.
Indeed, the reduced transport matrix \(H(a)\) is the coordinate form of the ambient transport geometry \(\mathfrak g_h\), and \(M_y(t,a)^\top M_y(t,a)\) is the coordinate form of the instantaneous ambient observability form \(\mathfrak j_{t,h(a)}^y\). Thus the reduced joint form is exactly the pullback of the ambient form \(Q_y[h]\) to the coefficient space. In this sense, the finite-dimensional model is not a separate inverse problem: it is the ambient inverse problem restricted to a finite-dimensional Bayes Hilbert subspace.
At the same time, restriction changes the observability mechanism in an essential way. In the ambient infinite-dimensional setting, localized mobile sensors can shrink the common invisible subspace but cannot produce coercivity on the full state space, because the time-averaged observation operator remains compact. In the reduced setting, that compactness obstruction disappears. One can first detect individual directions by localized bump sensors, then choose finitely many static locations that make the reduced state instantaneously observable, and finally trade those multiple static sensors for a single moving sensor whose time-averaged reduced Gramian is positive definite. When the reduced transport tensor is uniformly positive definite, the transport term upgrades that averaged observability to coercivity of the reduced joint form along the chosen reduced path.
It is important, however, to distinguish what has and has not been proved at this stage. The sensor-design theorems in this section are pathwise: they produce observability and coercivity statements along a fixed reduced truth path, and in the moving-sensor case the trajectory may depend on that truth path. They do not by themselves yield a single global inverse theorem over an entire reduced admissible class. That stronger conclusion requires an additional classwise reduced stability property for the experiment, which we formulate explicitly in the next subsection.
This is the sense in which the finite-dimensional theory is both a specialization of the ambient theory and a concrete recovery theory in its own right.
The previous subsections establish a truth-specific reduced observability theory. For each reduced truth path \(a^\dagger\in \mathcal{A}_m(U)\), one may design localized sensors—in particular, a truth-dependent single-sensor trajectory \(y[a^\dagger]\)—for which the averaged reduced Gramian is positive definite. Along that fixed truth path, this positivity combines with the reduced transport tensor to yield coercivity of the reduced joint transport–observability form on all reduced perturbations.
To convert that pathwise observability statement into a genuine inverse problem posed over a full reduced admissible class, one needs a stronger classwise reduced stability property for the chosen experiment. We therefore do not claim that the earlier moving-sensor design theorems alone yield global reduced reconstruction on \(\mathcal{A}_m(U)\). Instead, we formulate the required reduced stability estimate directly at the reduced level. Once that stronger property is available, the reduced Tikhonov problem admits global minimizers on the whole admissible coefficient class, and those reduced reconstructions lift to approximate ambient reconstructions with explicit projection-error, induced data-mismatch, and regularization-bias terms.
Throughout this subsection we fix the path-space norm \[\mathfrak H := L^2(0,T;L^2(\nu_0)) \cap H^1(0,T;L^2_0(\nu_0)), \qquad \|h\|_{\mathfrak H}^2 := \|h\|_{L^2(0,T;L^2(\nu_0))}^2 + \|\dot{h}\|_{L^2(0,T;L^2_0(\nu_0))}^2.\]
Let \[V_m=\operatorname{span}\{\phi_1,\dots,\phi_m\}\subset \mathcal{X}\] be the nested reduced spaces from Section 5, and let \[P_m:L^2_0(\nu_0)\to V_m\] be the \(L^2(\nu_0)\)-orthogonal projection, extended pointwise in time to path space by \[(P_m h)(t):=P_m(h(t)).\] For \(a=(a_1,\dots,a_m)\in\mathbb{R}^m\), we continue to write \[h(a):=\sum_{k=1}^m a_k\phi_k,\] and for a coefficient path \(a\in H^1(0,T;\mathbb{R}^m)\) we write \[\partial_t h(a)(t):=\sum_{k=1}^m \dot{a}_k(t)\phi_k.\]
Given a sensor trajectory \(y\), we retain the reduced pathwise forward map \[\mathcal{G}_m^y[a](t):=\mathcal{G}_t^y(h(a(t)))\in \mathbb{R}^r, \qquad \mathcal{Y}_T:=L^2(0,T;\mathbb{R}^r).\]
Definition 16 (Reduced regularization energy and reduced Tikhonov functional). For \(a\in \mathcal{A}_m(U)\) such that \(h(a(\cdot))\in \mathcal{A}_{\mathrm{ad}}\), define \[\begin{align} \label{eq:reduced-regularization-energy} \mathcal{R}_{\lambda,\mu}^{(m)}[a] & := \frac{\lambda}{2}\int_0^T \dot{a}(t)^\top H(a(t))\dot{a}(t)\,\mathrm{d}t \notag \\ & \quad + \frac{\mu}{2}\int_0^T \Bigl( \|h(a(t))\|_{L^2(\nu_0)}^2 + \|\partial_t h(a)(t)\|_{L^2(\nu_0)}^2 \Bigr)\,\mathrm{d}t . \end{align}\tag{51}\] Given data \(d\in \mathcal{Y}_T\), the corresponding reduced Tikhonov functional is \[\label{eq:reduced-tikhonov} \mathcal{J}_{\lambda,\mu}^{(m),y}[a;d] := \frac{1}{2}\|\mathcal{G}_m^y[a]-d\|_{\mathcal{Y}_T}^2 + \mathcal{R}_{\lambda,\mu}^{(m)}[a].\tag{52}\]
Remark 58. Because \(V_m\) is finite dimensional, all norms on \(V_m\) are equivalent. In particular, there exist constants \(c_m,C_m>0\) such that for every \(b\in\mathbb{R}^m\), \[c_m |b|^2 \le \left\|\sum_{k=1}^m b_k\phi_k\right\|_{L^2(\nu_0)}^2 \le C_m |b|^2.\] Hence \(\mathcal{R}_{\lambda,\mu}^{(m)}\) is equivalent to a Euclidean \(H^1\)-in-time penalty on the coefficient path \(a(\cdot)\). We keep the form 51 because it is exactly the pullback of the ambient regularization to \(V_m\).
Proposition 59 (Existence of global reduced minimizers). Let \(U\subset\mathbb{R}^m\) be compact, assume \(h(a)\in\mathcal{X}_{\mathrm{ad}}\) for every \(a\in U\), and assume the map \[\mathcal{A}_m(U)\ni a \longmapsto \mathcal{G}_m^y[a]\in \mathcal{Y}_T\] is continuous with respect to strong convergence in \(C([0,T];\mathbb{R}^m)\). Then for every \(d\in \mathcal{Y}_T\) and every \(\lambda,\mu>0\), the reduced functional \[a\longmapsto \mathcal{J}_{\lambda,\mu}^{(m),y}[a;d]\] attains a minimum on \(\mathcal{A}_m(U)\).
Proof. Let \((a_n)\subset \mathcal{A}_m(U)\) be a minimizing sequence. Since \(\mu>0\), the definition of \(\mathcal{J}_{\lambda,\mu}^{(m),y}\) and Remark 58 imply that \((a_n)\) is bounded in \(H^1(0,T;\mathbb{R}^m)\). Passing to a subsequence if necessary, \[a_n \rightharpoonup a \quad\text{weakly in }H^1(0,T;\mathbb{R}^m), \qquad a_n\to a \quad\text{strongly in }C([0,T];\mathbb{R}^m).\] Since \(U\) is compact, \(a(t)\in U\) for all \(t\in[0,T]\), so \(a\in \mathcal{A}_m(U)\).
By continuity of the reduced forward map under strong \(C([0,T];\mathbb{R}^m)\)-convergence, \[\mathcal{G}_m^y[a_n]\to \mathcal{G}_m^y[a] \qquad\text{in }\mathcal{Y}_T.\] The transport term in \(\mathcal{R}_{\lambda,\mu}^{(m)}\) is weakly lower semicontinuous because \(H(\cdot)\) is continuous on the compact set \(U\), hence bounded and uniformly positive semidefinite there. The remaining two regularization terms are convex quadratic forms in \(a\) and \(\dot{a}\), and therefore also weakly lower semicontinuous. Consequently, \[\mathcal{J}_{\lambda,\mu}^{(m),y}[a;d] \le \liminf_{n\to\infty}\mathcal{J}_{\lambda,\mu}^{(m),y}[a_n;d],\] so \(a\) is a minimizer. ◻
The next theorem is the reduced global reconstruction statement. It is formulated under a classwise reduced stability hypothesis for the chosen experiment. This is the natural reduced analogue of the ambient stability estimate from Section 4, but now imposed directly on the full reduced admissible class.
Theorem 60 (Global reduced regularized reconstruction). Let \(U\subset\mathbb{R}^m\) be compact, let \[a^\dagger\in \mathcal{A}_m(U), \qquad h_m^\dagger:=h(a^\dagger)\in \mathcal{A}_{\mathrm{ad}}^{s'},\] and fix a sensor trajectory \[y^\dagger\in \mathcal{Y}_{\mathrm{ad}}^{(1)}.\] Assume:
\(h(a)\in \mathcal{A}_{\mathrm{ad}}^{s'}\) for every \(a\in \mathcal{A}_m(U)\);
the reduced Tikhonov functional \[\mathcal{J}_{\lambda,\mu}^{(m),y^\dagger}[\cdot;d]\] admits minimizers on \(\mathcal{A}_m(U)\) for every \(d\in \mathcal{Y}_T\);
there exists \(C_m^{\mathrm{stab}}>0\) such that for every \(a\in \mathcal{A}_m(U)\), \[\label{eq:global-reduced-stability} \|h(a)-h(a^\dagger)\|_{L^2(0,T;L^2(\nu_0))}^2 \le C_m^{\mathrm{stab}} \left( \|\mathcal{G}_m^{y^\dagger}[a]-\mathcal{G}_m^{y^\dagger}[a^\dagger]\|_{\mathcal{Y}_T}^2 + \mathcal{R}_{\lambda,\mu}^{(m)}[a] \right).\tag{53}\]
Then for every noisy data \(d^\delta\in \mathcal{Y}_T\) and every minimizer \[\hat{a}_m^\delta\in \mathcal{A}_m(U)\] of \(\mathcal{J}_{\lambda,\mu}^{(m),y^\dagger}[\cdot;d^\delta]\), one has \[\label{eq:global-reduced-regularized-estimate} \|h(\hat{a}_m^\delta)-h(a^\dagger)\|_{L^2(0,T;L^2(\nu_0))}^2 \le 5\,C_m^{\mathrm{stab}} \left( \|d^\delta-\mathcal{G}_m^{y^\dagger}[a^\dagger]\|_{\mathcal{Y}_T}^2 + \mathcal{R}_{\lambda,\mu}^{(m)}[a^\dagger] \right).\tag{54}\]
Proof. By the minimizing property of \(\hat{a}_m^\delta\), \[\frac{1}{2}\|\mathcal{G}_m^{y^\dagger}[\hat{a}_m^\delta]-d^\delta\|_{\mathcal{Y}_T}^2 + \mathcal{R}_{\lambda,\mu}^{(m)}[\hat{a}_m^\delta] \le \frac{1}{2}\|\mathcal{G}_m^{y^\dagger}[a^\dagger]-d^\delta\|_{\mathcal{Y}_T}^2 + \mathcal{R}_{\lambda,\mu}^{(m)}[a^\dagger].\] Hence \[\label{eq:reduced-minimizer-basic-bound} \mathcal{R}_{\lambda,\mu}^{(m)}[\hat{a}_m^\delta] \le \frac{1}{2}\|d^\delta-\mathcal{G}_m^{y^\dagger}[a^\dagger]\|_{\mathcal{Y}_T}^2 + \mathcal{R}_{\lambda,\mu}^{(m)}[a^\dagger]\tag{55}\] and \[\label{eq:reduced-data-residual-basic-bound} \|\mathcal{G}_m^{y^\dagger}[\hat{a}_m^\delta]-d^\delta\|_{\mathcal{Y}_T}^2 \le \|d^\delta-\mathcal{G}_m^{y^\dagger}[a^\dagger]\|_{\mathcal{Y}_T}^2 + 2\,\mathcal{R}_{\lambda,\mu}^{(m)}[a^\dagger].\tag{56}\]
Now \[\begin{align} \|\mathcal{G}_m^{y^\dagger}[\hat{a}_m^\delta]-\mathcal{G}_m^{y^\dagger}[a^\dagger]\|_{\mathcal{Y}_T}^2 & \le 2\|\mathcal{G}_m^{y^\dagger}[\hat{a}_m^\delta]-d^\delta\|_{\mathcal{Y}_T}^2 + 2\|d^\delta-\mathcal{G}_m^{y^\dagger}[a^\dagger]\|_{\mathcal{Y}_T}^2 \\ & \le 4\|d^\delta-\mathcal{G}_m^{y^\dagger}[a^\dagger]\|_{\mathcal{Y}_T}^2 + 4\,\mathcal{R}_{\lambda,\mu}^{(m)}[a^\dagger], \end{align}\] where in the last step we used 56 .
Applying the global reduced stability estimate 53 with \(a=\hat{a}_m^\delta\), and then using the last inequality together with 55 , we obtain \[\begin{align} \|h(\hat{a}_m^\delta)-h(a^\dagger)\|_{L^2(0,T;L^2(\nu_0))}^2 & \le C_m^{\mathrm{stab}} \left( \|\mathcal{G}_m^{y^\dagger}[\hat{a}_m^\delta]-\mathcal{G}_m^{y^\dagger}[a^\dagger]\|_{\mathcal{Y}_T}^2 + \mathcal{R}_{\lambda,\mu}^{(m)}[\hat{a}_m^\delta] \right) \\ & \le C_m^{\mathrm{stab}} \left( 4\|d^\delta-\mathcal{G}_m^{y^\dagger}[a^\dagger]\|_{\mathcal{Y}_T}^2 + 4\,\mathcal{R}_{\lambda,\mu}^{(m)}[a^\dagger] + \frac{1}{2}\|d^\delta-\mathcal{G}_m^{y^\dagger}[a^\dagger]\|_{\mathcal{Y}_T}^2 + \mathcal{R}_{\lambda,\mu}^{(m)}[a^\dagger] \right) \\ & \le 5\,C_m^{\mathrm{stab}} \left( \|d^\delta-\mathcal{G}_m^{y^\dagger}[a^\dagger]\|_{\mathcal{Y}_T}^2 + \mathcal{R}_{\lambda,\mu}^{(m)}[a^\dagger] \right), \end{align}\] which is 54 . ◻
We now lift the preceding reduced estimate to the ambient path space. Since the sensor trajectory may depend on the reduced truth, the lifted statement is formulated for a truth-dependent sequence of reduced experiments.
Theorem 61 (Approximate ambient reconstruction from reduced regularized minimizers). Let \[\mathcal{K}\subset \mathcal{A}_{\mathrm{ad}}^{s'}\cap \mathfrak H\] be a \(\mathfrak H\)-compact admissible path class. Assume:
for every \(m\in\mathbb{N}\), \[P_m\mathcal{K}\subset \mathcal{A}_{\mathrm{ad}}^{s'}\cap \mathfrak H,\] and \[\sup_{h\in\mathcal{K}}\|(I-P_m)h\|_{\mathfrak H}\longrightarrow 0 \qquad\text{as }m\to\infty;\]
for each \(m\) and each admissible sensor trajectory \(y\) used below, the pathwise observation map is Lipschitz on \(\mathcal{K}\cup P_m\mathcal{K}\): there exists \(L_{\mathcal{K},m,y}>0\) such that \[\label{eq:ambient-observation-lipschitz-last-subsection} \|\mathcal{G}^{y}[h_1]-\mathcal{G}^{y}[h_2]\|_{\mathcal{Y}_T} \le L_{\mathcal{K},m,y}\|h_1-h_2\|_{\mathfrak H}\tag{57}\] for all \(h_1,h_2\in \mathcal{K}\cup P_m\mathcal{K}\).
Fix \[h^\dagger\in\mathcal{K}.\] For each \(m\), write \[z_m^\dagger:=P_m h^\dagger\in V_m, \qquad z_m^\dagger=h(a_m^\dagger)\] for the unique coefficient path \(a_m^\dagger\in \mathcal{A}_m(U_m)\), where \(U_m\subset\mathbb{R}^m\) is a compact set containing the image of \(a_m^\dagger\). Let \[y_m^\dagger\in \mathcal{Y}_{\mathrm{ad}}^{(1)}\] be the reduced experiment chosen for the reduced truth \(a_m^\dagger\), and assume the hypotheses of Theorem 60 hold on \(\mathcal{A}_m(U_m)\) with this trajectory \(y_m^\dagger\). Let \[d_m^\delta\in \mathcal{Y}_T, \qquad \|d_m^\delta-\mathcal{G}^{y_m^\dagger}[h^\dagger]\|_{\mathcal{Y}_T}\le \delta,\] and let \[\hat{a}_m^\delta\in \mathcal{A}_m(U_m)\] be a minimizer of the reduced Tikhonov functional \(\mathcal{J}_{\lambda,\mu}^{(m),y_m^\dagger}[\cdot;d_m^\delta]\). Define the lifted reduced reconstruction \[\hat{h}_m^\delta:=h(\hat{a}_m^\delta)\in V_m.\]
Then \[\begin{align} \label{eq:ambient-lifted-regularized-estimate} \|\hat{h}_m^\delta-h^\dagger\|_{L^2(0,T;L^2(\nu_0))} & \le \|(I-P_m)h^\dagger\|_{L^2(0,T;L^2(\nu_0))} \notag \\ & \quad + \Biggl[ 5\,C_m^{\mathrm{stab}} \Bigl( \bigl(\delta + L_{\mathcal{K},m,y_m^\dagger}\|(I-P_m)h^\dagger\|_{\mathfrak H}\bigr)^2 + \mathcal{R}_{\lambda,\mu}^{(m)}[a_m^\dagger] \Bigr) \Biggr]^{1/2}. \end{align}\tag{58}\]
Proof. Set \[d_m^\dagger:=\mathcal{G}_m^{y_m^\dagger}[a_m^\dagger] =\mathcal{G}^{y_m^\dagger}[z_m^\dagger].\] By the observation Lipschitz estimate 57 , \[\|\mathcal{G}^{y_m^\dagger}[h^\dagger]-d_m^\dagger\|_{\mathcal{Y}_T} = \|\mathcal{G}^{y_m^\dagger}[h^\dagger]-\mathcal{G}^{y_m^\dagger}[z_m^\dagger]\|_{\mathcal{Y}_T} \le L_{\mathcal{K},m,y_m^\dagger}\|h^\dagger-z_m^\dagger\|_{\mathfrak H} = L_{\mathcal{K},m,y_m^\dagger}\|(I-P_m)h^\dagger\|_{\mathfrak H}.\] Therefore \[\label{eq:effective-reduced-data-error-global} \|d_m^\delta-d_m^\dagger\|_{\mathcal{Y}_T} \le \|d_m^\delta-\mathcal{G}^{y_m^\dagger}[h^\dagger]\|_{\mathcal{Y}_T} + \|\mathcal{G}^{y_m^\dagger}[h^\dagger]-d_m^\dagger\|_{\mathcal{Y}_T} \le \delta + L_{\mathcal{K},m,y_m^\dagger}\|(I-P_m)h^\dagger\|_{\mathfrak H}.\tag{59}\]
Now apply Theorem 60 with reference path \(a_m^\dagger\), chosen experiment \(y_m^\dagger\), and reduced data \(d_m^\delta\). Using 59 , we obtain \[\|\hat{h}_m^\delta-z_m^\dagger\|_{L^2(0,T;L^2(\nu_0))}^2 \le 5\,C_m^{\mathrm{stab}} \Bigl( \bigl(\delta + L_{\mathcal{K},m,y_m^\dagger}\|(I-P_m)h^\dagger\|_{\mathfrak H}\bigr)^2 + \mathcal{R}_{\lambda,\mu}^{(m)}[a_m^\dagger] \Bigr).\] Finally, use the triangle inequality: \[\|\hat{h}_m^\delta-h^\dagger\|_{L^2(0,T;L^2(\nu_0))} \le \|\hat{h}_m^\delta-z_m^\dagger\|_{L^2(0,T;L^2(\nu_0))} + \|z_m^\dagger-h^\dagger\|_{L^2(0,T;L^2(\nu_0))},\] and note that \(z_m^\dagger=P_m h^\dagger\). This gives 58 . ◻
Remark 62 (Interpretation of the lifted estimate). The estimate 58 separates three effects.
First, the term \[\|(I-P_m)h^\dagger\|_{L^2(0,T;L^2(\nu_0))}\] is the unavoidable Galerkin projection error.
Second, the quantity \[L_{\mathcal{K},m,y_m^\dagger}\|(I-P_m)h^\dagger\|_{\mathfrak H}\] is the induced data/model mismatch: even in the absence of measurement noise, the \(m\)-th reduced inverse problem uses data generated by the full ambient truth rather than by its projection.
Third, the term \[\mathcal{R}_{\lambda,\mu}^{(m)}[a_m^\dagger]\] is the regularization bias of the reduced reference path. At fixed \(\lambda,\mu>0\), the theorem therefore gives an approximate ambient reconstruction estimate for reduced regularized minimizers, not an exact consistency statement. Removing that bias would require an additional reduced vanishing-regularization analysis, which we do not pursue here.
Remark 63. The role of Section 5 can now be stated precisely.
First, the finite-dimensional reduction identifies the reduced tensors and sensing matrices as coordinate pullbacks of the ambient transport and observability geometry.
Second, Theorem 53 and Theorem 55 provide a truth-specific reduced observability theory: for each fixed reduced truth path, one may design a localized sensing experiment—including a single moving sensor—whose reduced averaged Gramian is positive definite and whose reduced joint form is coercive along that path.
Third, the present subsection isolates the additional ingredient needed to pass from those pathwise observability statements to a global reduced inverse problem on \(\mathcal{A}_m(U)\): one needs a classwise reduced stability estimate for the chosen experiment. Under that stronger assumption, reduced regularized minimizers exist globally and lift to approximate ambient reconstructions with explicit projection-error, induced data-mismatch, and regularization-bias terms.
In particular, the finite-dimensional theory should be read in two layers: constructive reduced observability along fixed reduced truths, and global reduced inversion once an additional classwise stability property is imposed.