We consider a two-dimensional periodic reaction-diffusion system under natural conditions on the reaction function and with initial condition \(\theta\). We show that on the global attractor \(\mathcal{A}\) of the resulting dynamical system \((u_\theta(t):t>0)\), a reverse Poincaré inequality holds true, and that as a consequence the map \(\theta \mapsto
u_\theta(t)\) satisfies a \(L^2\)-Lipschitz stability estimate on \(\mathcal{A}\) for any \(t>0\) fixed. We then show that statistical recovery of
an initial condition \(\theta\) in the attractor \(\mathcal{A}\), as well as prediction of the states \(u_\theta\), is possible from discrete measurements of
the system at ‘fast’ near parametric convergence rates.
Consider a non-linear dynamical system \((u(t, \cdot): t >0)\) evolving in the Hilbert space \(L^2(\Omega)\) of square integrable functions over some bounded domain \(\Omega\) in Euclidean space. The focus here will be the situation when the infinitesimal dynamics are described by a non-linear parabolic partial differential equation (PDE) of the form \[\label{pde0}
\begin{align} \frac{\partial u}{\partial t} - \Delta u - F(u) &=0, \\ u(0,\cdot)&=\theta,
\end{align}\tag{1}\] where \(\Delta\) is the Laplacian (acting on \(x \in \Omega\)), where the functional \(F\) is given and models the
non-linearity, and where \(\theta \in L^2(\Omega)\) is an unknown initial condition. We consider the scenario where the system is in its equilibrium dynamics, specifically, when it evolves in its global attractor,
which consists of all the states of the system that can be reached as \(t \to \infty\). If we denote by \(\{S(t):t\ge 0\}\) the semigroup \(S(t)\theta=u_\theta(t)\) associated with the above PDE, then under certain assumptions on \(F\) the long-time dynamics ‘collapse’ into a compact subset of \(L^2(\Omega)\). Specifically, if \(B\) is a compact set in \(L^2(\Omega)\) such that \(S(t)X\subset B\) for all bounded \(X\subset L^2(\Omega)\) and all \(t\ge t_0(X)\) large enough, then the set \[\label{lazydimitri}
\mathcal{A}
\equiv
\bigcap_{t> 0} S(t) B\tag{2}\] is nonempty, compact and invariant in the sense that \(S(t)\mathcal{A}=\mathcal{A}\) for all \(t\ge 0\). Thus once the dynamics reach \(\mathcal{A}\), they never leave this region again, and \(\mathcal{A}\) describes a family of equilibrium states known as the ‘global attractor’ of the dynamical system – they can be seen to be
independent of the choice of \(B\). It can be shown to exist in a variety of non-linear ‘dissipative’ PDEs including reaction-diffusion and Navier-Stokes equations, and generally has a finite but positive Hausdorff
dimension, see [1]–[3].
In this article we are concerned with the problem of how to estimate or predict the states \((u_\theta(\tau):\tau \ge 0)\) of the system from discrete statistical observations, thus contributing to a recently emerging
literature on nonlinear statistical inference problems with parabolic PDEs, see [4]–[14] and references therein. We consider the case where the system has reached its ‘equilibrium dynamics’, that is, when \(u_\theta\) evolves in \(\mathcal{A}\). The measurement setting we consider is the following: For \(t>0\) fixed, data \(Z^{(N)} = (Y_i, X_i)_{i=1}^N\) is obtained through the
regression model \[\label{model}
Y_i
=
u_{\theta}(t, X_i)+ \varepsilon_i
,\tag{3}\] where the initial condition \(\theta\) is known to lie in \(\mathcal{A}\), where the \(X_i\)’s are independent and identically
distributed (i.i.d.) drawn from the uniform distribution over \(\Omega\), and the \(\varepsilon_i\)’s are i.i.d. standard normal \(\mathscr{N}(0,1)\) random
variables accounting for measurement errors, independent of the \(X_i\)’s. Let \(P_\theta\) denote the law in \(\mathbb{R}\times \Omega\) of \((Y_1, X_1)\) with corresponding infinite product measure \(P_\theta^{\mathbb{N}}\) whose \(N\)-dimensional marginal distributions describe the law of \(Z^{(N)}\).
It is a common practice in data assimilation models such as (3 ) to assign a Gaussian random field to the initial condition \(\theta\) and to then use the pushforward under the dynamics as a
‘prior’ probability measure in the space of trajectories of \((u_\theta(t, \cdot): t>0)\). A typical situation [15]–[18] is to maintain an infinite-dimensional model for the initial condition \(\theta\) which can be an arbitrary element of a Sobolev space \(\theta \in H^\alpha\) for suitable \(\alpha>0\). But if it is believed that the system has already equilibrated in its global attractor \(\mathcal{A}\), then
the prior distribution \(\Pi\) for the law of \(\theta\) should be chosen to reflect this, ideally by verifying \(\Pi(\mathcal{A})=1\). It is generally hard
to analytically describe \(\mathcal{A}\) and hence to suggest computationally implementable priors \(\Pi\) – we shall describe below two constructions of probability measures that do
concentrate all their mass either on \(\mathcal{A}\) itself or on the slightly larger ‘inertial manifold’ \(\mathcal{M}\) engulfing \(\mathcal{A}\).
In proving theorems about the statistical performance of posterior measures arising from such priors, the inverse problem of identifying the initial condition \(\theta\) from the measurements of \(u_\theta\) plays a central role. Recovery of the initial condition is generally a severely ill-posed problem in the measurement model (3 ) with \(t>0\) fixed, and the
best available minimax statistical convergence rates for \(\theta\) belonging to a fixed Sobolev ball are ‘slow’, that is, of order \(1/\log N\) where \(N\)
is the sample size. This was proved rigorously in Theorems 2 and 4 in [5] for the periodic 2D Navier-Stokes equations but holds also more generally (at
least in perturbative neighbourhoods of the standard heat equation, i.e., when \(F\) in (1 ) is small in an appropriate sense). When measurements for times near zero are available, faster (algebraic
in \(1/N\)) convergence rates are available as was proved in the recent articles [6], [19]. In the present work we identify another situation where slow logarithmic rates can be improved upon, namely when the initial condition \(\theta\) is known to belong to the global
attractor \(\mathcal{A}\), so that the system is in equilibrium. We will show that in this case statistical estimators (i.e., measurable functions of the data \(Z^{(N)}\)) exist that
themselves belong to \(\mathcal{A}\) and that infer the initial condition and then also \(u_\theta(\tau), \tau>0,\) at near parametric convergence rates \(\sqrt{\log N}/\sqrt N\), even when no samples near \(t=0\) are available.
The mathematical proofs of these facts will be executed in a specific periodic two-dimensional reaction-diffusion model, where we show that a key reverse Poincaré inequality, which gives a Lipschitz stability estimate for the inverse problem similar in
spirit to Theorem 1B in [5], can be established. To introduce the PDE, for \(\theta\in L^2(\Omega)\), consider the
unique solution \(u=u_\theta:[0,\infty)\times \Omega\to \mathbb{R}\) to the reaction-diffusion equation \[\label{eq:ReacDiff}
\begin{align}
\frac{\partial u}{\partial t}(t,x) -\Delta u(t,x) &= f(u(t,x))
\quad \text{on } (0,\infty)\times \Omega, \\
u(0,x) &= \theta(x)
\quad\quad~~~~ \text{on } \Omega
\end{align}\tag{4}\] where \(\Omega = (0,1]^2\) is the two-dimensional torus with unit area and opposite points identified. The reaction term \(f:\mathbb{R}\to\mathbb{R}\) is
of class \(C^3(\mathbb{R})\) – throughout, \(C^k(\mathbb{R})\) denotes the class of \(k\)-times continuously differentiable maps defined on \(\mathbb{R}\), without any boundedness assumption – and further satisfies, for some \(p>2\), and constants \(\alpha_1 > \alpha_2 > 0\), \(k\ge 0\) and \(L\in \mathbb{R}\), the inequalities \[\label{Condf}
-k - \alpha_1 |s|^p
\le
f(s)s
\le
k - \alpha_2 |s|^p,
\quad
\textrm{and}
\quad
f'(s)\le L
,\quad\quad
\forall\;s\in \mathbb{R}.\tag{5}\] Under these assumptions, the system of equations in (4 ) possesses a unique (weak) solution satisfying \(u_{\theta} \in C([0,\infty),
L^2(\Omega))\); we review this in Proposition 10 below. For each \(t\ge 0\) define the solution map \(S(t)
: L^2(\Omega)\to L^2(\Omega),\;\theta \mapsto u_{\theta}(t,\cdot)\). Then \(\{S(t) : t\ge 0\}\) is a semigroup on \(L^2(\Omega)\), and the semidynamical system \((L^2(\Omega), \{S_t\}_{t\ge 0}\}\) admits a global attractor \(\mathcal{A}\) as we explain in Proposition 12
below. We discuss in Remark 3 concrete examples for \(f\) such that the corresponding attractor \(\mathcal{A}\) has strictly positive and finite Hausdorff dimension.
Our main statistical result is concerned with convergence rates as sample size \(N \to \infty\) for the problem of inferring an initial condition \(\theta \in \mathcal{A}\) and with
predicting \(u_\theta(\tau)\) evolving in \(\mathcal{A}\) at arbitrary times \(\tau>0\).
Theorem 1. Assume that \(\theta_0\in \mathcal{A}\) and consider i.i.d. data \((Y_i, X_i)_{i=1}^N\) obtained through (3 ) for some \(t>0\) where \(u_\theta\) solves the PDE (4 ) with known \(f \in C^3(\mathbb{R})\) satisfying (5 ). Then,
one can construct estimators \(\hat{\theta}_N = \hat{\theta}(Y_i, X_i)_{i=1}^N\) such that for every \(\tau>0\) we have as \(N \to \infty\)\[\|u_{\hat{\theta}_N}(\tau) - u_{\theta_0}(\tau)\|_{L^2(\Omega)} \lesssim \|\hat{\theta}_N - \theta_0\|^2_{L^2(\Omega)} = O_{P_{\theta_0}^\mathbb{N}}\Big( \frac{\log N}{N}\Big).\]
Inspection of the proofs shows that \(\hat{\theta}_N\) and \(u_{\hat{\theta}_N}\) themselves take values in \(\mathcal{A}\) (or in its inertial manifold
\(\mathcal{M}\)). The convergence rate for prediction of the states \(u_\theta(\tau), \tau>0,\) scales exponentially in \(\tau\), see Proposition 10. A similar remark applies to the dependence of the rate of convergence on the measurement time \(t>0\).
The proof of the preceding theorem uses that \(\mathcal{A}\) has the metric complexity of a bounded subset of a finite-dimensional space. One may then expect the convergence rate to be even faster, namely \(1/\sqrt N\), as is the case when considering finite-dimensional parameter spaces in other severely ill-posed non-linear inverse problems such as the Calderón problem, see [20], [21]. However even though \(\mathcal{A}\) is in a certain sense finite-dimensional, it is not clear whether local LAN
expansions of \(P_{\theta+h}\) in directions \(h\), and the related notion of Fisher information, make sense on \(\mathcal{A}\) due to its lack of linear
structure. Hence techniques from classical finite-dimensional statistics ([22]) do not apply straightforwardly in the present context. Instead the proof of
Theorem 1 is based on general theory from Bayesian nonparametric statistics [23] as
adapted to non-linear PDE inverse problems in [24] and [25]. This proof method requires crucially the following key stability estimate for the parameterisation \(\theta \mapsto u_\theta\) of the inverse problem underlying the regression model. It
shows that the ill-posed nature of that inverse problem disappears when the dynamics has equilibrated to the global attractor \(\mathcal{A}\).
Theorem 2. Let \(u_\theta\) be as in Theorem 1. For all \(t>0\), there exists a
positive constant \(c=c(t,f)\) depending only on \(t, f\) such that \[\|\theta-\vartheta\|_{L^2} \le c \|u_{\theta}(t)-u_{\vartheta}(t)\|_{L^2} ,\quad\quad
\forall\;\theta,\vartheta\in \mathcal{A} .\]
An exact value for the constant \(c\) is given in Theorem 7. The proof is inspired by Theorem 1B in [5] for the Navier-Stokes equations where one assumes a reverse Poincaré inequality \(\|\nabla(\theta-\vartheta)\|_{L^2} \lesssim
\|\theta-\vartheta\|_{L^2}\). For \(\theta, \vartheta\) belonging to the global attractor \(\mathcal{A}\) of the PDE (4 ), this inequality holds true as
we show in Proposition 6 below. The proof relies on certain properties of the distribution of the eigenvalues of the Laplacian and of the non-linearity \(F\) featuring in (1 ) which are specific to the two-dimensional periodic situation. Proving similar results for other dimensions or the Navier-Stokes equations remains a formidable challenge, but we
conjecture that similar theorems do hold true in these settings as well.
Remark 3 (Discussion of Condition 5 ). Because \(f\) is assumed to be continuous over \(\mathbb{R}\) it is locally bounded, and then a
straightforward calculation shows that the two-sided inequality in (5 ) equivalently rewrites in the following more interpretable form: \[-a-b_1~{\rm sign}(s)|s|^{p-1}
\le
f(s)
\le
a-b_2~{\rm sign}(s)|s|^{p-1}
,\] for all \(s\in\mathbb{R}\) and some \(a\ge 0\), \(b_i > 0\), \(p>2\). Examples of such reaction functions
\(f\) also satisfying the second condition in (5 ) include \(f(s)=\sum_{i=1}^{p-2} h_i(s) s^i + c_p s |s|^{p-1}\) with \(p \ge 3\)
and negative \(c_p<0\), for smooth bounded and Lipschitz functions \(h_i\)’s.
Let us remark, to avoid triviality, that for such \(f\) the global attractor \(\mathcal{A}\) exists (Proposition 12) and can further be seen to have non-trivial dimension: This is granted for instance for \(f\) satisfying (5 ), as soon as \(f'(0)\) is sufficiently large and \(f(0)=0\). Indeed, Theorems 9.4 and 9.5 in [26] imply that
\(\mathcal{A}\) has positive Hausdorff dimension as soon as \(f\) is of the form \(f(s)=\lambda s + h(s)\) with sufficiently large \(\lambda\) and smooth \(h\) satisfying (5 ) in place of \(f\) and such that \(h(0)=h'(0)=0\). This
can be seen by applying these results with \(f\) there given by \(-h(s)\), \(n=2\), \(\lambda\) large enough, and \(g=g_1=0\), also taking note that \(h(u)\in C^\ell(\Omega)\) for any \(\ell\in\mathbb{N}\) since \(h\) is smooth, and hence so is
\(u\) (see Theorem 11.8 in [3]). In particular, \(\mathcal{A}\) has non-trivial
dimension for \(f(s)=\sum_{i=1}^p c_i s^i\) with \(p\) odd, \(c_p<0\) and large enough \(c_1\), and more generally for
\(f(s) = \sum_{i=1}^{p-1} h_i(s) s^i + c_p s^p\) with \(p\ge 3\) odd, \(h_i\) as before, \(c_p<0\), and \(h_1'(0)\) large enough.
The intuition behind hypothesis (5 ) is, one the one hand, that the forcing \(f(u)\) behaves like \(-|u|^{p-1}\) when \(|u|\) is
large, thus ‘taking away energy in the system’ to prevent blowups. On the other hand, when \(|u|\) is close to zero, the forcing \(f(u)\) behaves like \(f'(0)u\), preventing the dynamics to die out by adding energy proportional to \(f'(0)\).
We now construct two prior distributions, one on \(\mathcal{A}\) and one on the slightly larger inertial manifold \(\mathcal{M} \supset \mathcal{A}\) (to be constructed in Section 3.2 below). We then derive contraction rates for the posterior distributions arising from measurements (3 ), following ideas in Bayesian nonparametric statistics, see [23], Chapter 7.3 in [27], or Chapter 1 in [25]. These in turn imply the existence of estimators in Theorem 1.
The global attractor \(\mathcal{A}\) from Proposition 12 has finite-dimensional \(L^2(\Omega)\)-covering numbers: \[\label{CoverEstimate}
N(\mathcal{A}, \|\cdot\|_{L^2},\varepsilon)
\lesssim
\varepsilon^{-D}
,~~
\varepsilon> 0
,\tag{6}\] for some finite \(D\); this follows from Definition 13.1 and Theorem 13.19 in [3] whose
proof applies in the periodic setting as well. Fix \(\varepsilon>0\) and an integer \(n_\varepsilon\asymp \varepsilon^{-D}\). Let \(a_1, \ldots,
a_{n_\varepsilon}\in\mathcal{A}\) be an \(\varepsilon\)-covering of \(\mathcal{A}\). Define \(\Pi_\varepsilon\) as the uniform distribution in \(L^2(\Omega)\) over \(\{a_1,\ldots, a_{n_\varepsilon}\}\). For any arbitrary real \(\tau >0\), define sequences \[\varepsilon_j
=
\frac{1}{j^{\tau}}
,\quad\quad
c_j \propto \frac{1}{j^{1+\tau}}
,\quad\quad
j\ge 1
,\] such that \(\sum_{j\ge 1} c_j = 1\). We then define the prior \(\Pi\) as a weighted combination of \(\Pi_{\varepsilon_j}\)\[\label{PriorNet}
\Pi
\equiv
\sum_{j=1}^\infty c_j \Pi_{\varepsilon_j}
.\tag{7}\] This defines a probability measure supported on \(\mathcal{A}\) and, given data from (3 ), gives rise to the posterior measure \(\Pi(\cdot|Z^{(N)})\), also supported on \(\mathcal{A}\), obtained as \[d\Pi(a|Z^{(N)})
\propto
e^{\ell_N(a)} d\Pi(a)
,\quad\quad
\ell_N(a)
=
-\frac{1}{2} \sum_{i=1}^N |Y_i-u_{a}(t, X_i)|^2
, ~~a \in \mathcal{A}.\] We can now establish the following contraction result for the posterior measure \(\Pi(\cdot| Z^{(N)})\).
Theorem 4. Assume that \(\theta_0\in \mathcal{A}\) and consider i.i.d. data \(Z^{(N)}=(Y_i, X_i)_{i=1}^N\) obtained through (3 ) for some
\(t>0\) where \(u_\theta\) solves the PDE (4 ) with known \(f \in C^3(\mathbb{R})\) satisfying (5 ).
Let \(\delta_N \equiv \sqrt{\log(N)/N}\). Then, there exists \(M>0\) such that we have in \(P_{\theta_0}^\mathbb{N}\)-probability as \(N\to \infty\)\[\Pi\Big( a\in\mathcal{A}: \|u_{a}(t)-u_{\theta_0}(t)\|_{L^2} \le M \delta_N~ |~ Z^{(N)} \Big)
\to
1
.\]
Proof of Theorem 4. We apply Theorem 1.3.2 in [25]
with parameter space \(\Theta\equiv \mathcal{A}\), constant regularization sets \(\Theta_N \equiv \mathcal{A}\) satisfying \(\Pi(\Theta_N^c)=0\) since \(\Pi(\mathcal{A})=1\), and \(\delta_N\) as in the statement. Also note that \(u_\theta(t) \in \mathcal{A}\) consists of functions uniformly bounded by a fixed
constant \(U\) in view of Proposition 12 and the Sobolev embedding \(H^2(\Omega) \hookrightarrow
L^\infty(\Omega)\). First observe that the ‘intrinsic’ semimetric \(d_\mathcal{G}\) there is given by \[d_\mathcal{G}(\theta, \theta')
\equiv
\|u_\theta(t)-u_{\theta'}(t)\|_{L^2(\Omega)}, ~t>0.\] Proposition 10 entails that the map \(\theta\mapsto u_\theta(t)\) is
\(L^2\)-Lipschitz continuous, with Lipschitz constant \(C_t\). We can thus upper-bound the covering numbers \[\label{lipest}
\log N(\mathcal{A}, d_\mathcal{G}, \delta_N)
\leq
\log N(\mathcal{A}, \|\cdot\|_{L^2}, \delta_N/C_t)
\lesssim \log(1/\delta_N)
\lesssim N\delta_N^2
,\tag{8}\] by virtue of the covering estimate (6 ). It remains to establish a small ball estimate for the measurable sets \[\mathcal{B}_N
=
\Big\{ a\in \mathcal{A}: \|u_a(t)-u_{\theta_0}(t)\|_{L^2(\Omega)} \le \delta_N \Big\}
.\] Defining the ball \[\mathcal{B}'_N
=
\Big\{ a\in \mathcal{A}: \|a-\theta_0\|_{L^2} \le \delta_N/C_t \Big\}
,\] with \(C_t\) as above yields \(\Pi(\mathcal{B}_N)
\ge \Pi(\mathcal{B}'_N)\). To lower bound \(\Pi(\mathcal{B}'_N)\) observe that, for any \(N\ge 1\) fixed and for all \(j\) large enough such
\(\varepsilon_j\le \delta_N/C_t\), the \(\varepsilon_j\)-covering \(a_1,\ldots, a_{n_{\varepsilon_j}}\) constructed previously admits at least one element
\(a_i\) in \(\mathcal{B}'_N\). Since \(\Pi_{\varepsilon_j}\) is uniformly distributed over \(n_{\varepsilon_j}\)
elements, we find \[\Pi(\mathcal{B}'_N)
\ge
\sum_{j:\varepsilon_j\le \delta_N/C_t} c_j \Pi_{\varepsilon_j}(\mathcal{B}'_N)
\ge
\sum_{j:\varepsilon_j\le \delta_N/C_t} \frac{c_j}{n_{\varepsilon_j}}
.\] Recalling that \(n_{\varepsilon_j}\asymp \varepsilon_j^{-D}\), \(\varepsilon_j = j^{-\tau}\) and \(c_j \propto j^{-(1+ \tau)}\), a straightforward
calculation provides \[\Pi(\mathcal{B}_N')
\gtrsim
\sum_{j\ge (\delta_N/C_t)^{-\frac{1}{\tau}}} \frac{1}{j^{1+\tau + D\tau}}
\gtrsim
\delta_N^{D+1}
\ge
e^{-AN\delta_N^2}
,\] for \(A>0\) large enough. Theorem 1.3.2 in [25] thus implies the conclusion of the theorem. ◻
2.1.2 Finite-dimensional parametrization and the inertial manifold↩︎
The negative Laplacian \(-\Delta\) has eigenfunctions \(e_j\) forming an orthonormal basis of \(L^2(\Omega)\) with corresponding eigenvalues \(\lambda_j\) satisfying \(0=\lambda_0 \le \ldots \le \lambda_{j-1} \le \lambda_j \asymp j\) as \(j \to \infty\). On the Sobolev space \(H^2(\Omega)\) we consider now the equivalent sequence norm \(\|u\|^2_{\tilde{H}^2} \equiv \sum_{j \ge 0}(1+\lambda_j)^2 \langle e_j, u \rangle_{L^2}^2\). Also, for any \(n\in\mathbb{N}\), let \(P_n\) denote the projection of \(u \in L^2(\Omega)\) onto the first \(n\) eigenfunctions of \(-\Delta\), i.e.\[\label{projn}
P_n[u]
\equiv
\sum_{j=0}^{n-1} \left\langle u,e_j \right\rangle_{L^2} e_j
,\tag{9}\] and \(Q_n[u]=u-P_n[u]\) the projection onto its orthogonal complement.
Let \(\rho>0\) be an \(H^2\)-bound on \(\mathcal{A}\) (see after Theorem 16) and let \(B_{H^2}(8\rho)\) be the centred \(H^2\)-ball with radius \(8\rho\). Theorem 16 below entails the existence of an \(L^2\)-Lipschitz map \(\Phi:P_n[L^2(\Omega)]\to Q_n[L^2(\Omega)]\) such that any \(u\in\mathcal{A}\) writes (uniquely) as \(u=\Phi(u_n)+u_n\) where \(u_n\equiv P_n[u]\). In addition, \(\Phi\) can be taken to be
supported on the \(n\)-dimensional ellipsoid \[\label{defDPhi}
D(\Phi)
\equiv
P_n[L^2(\Omega)]\cap B_{H^2}(8\rho)
.\tag{10}\] Next, the map \(F: D(\Phi)\to L^2(\Omega)\) defined as \[F(v)
=
v+\Phi(v)
,\quad\quad
v\in D(\Phi)\] is also \(L^2\)-\(L^2\) Lipschitz. We thus define the ‘inertial manifold’ \(\mathcal{M}\equiv F(D(\Phi))\) of the reaction-diffusion
system, which is bounded in \(H^2(\Omega)\) and contains \(\mathcal{A}\) as a subset – see also (22 ). To construct a prior on \(\mathcal{M}\), first define a random vector \(X=(X_1,\ldots, X_n)\) taking values in the ellipsoid \(D(\Phi)\) seen as a subset of \(\mathbb{R}^n\) with a continuous and (strictly) positive density with respect to the Lebesgue measure. By abuse of notation we also denote by \(X\) the corresponding random field \[X
\equiv
\sum_{j=1}^n
X_j e_j
,\] taking values in the function space \(P_n[L^2(\Omega)]\). We then take the prior \(\Pi\) as \[\label{Prior2}
\Pi
\equiv
\text{Law}(F(X))
~
\text{in}~ L^2(\Omega)
.\tag{11}\] This defines a probability measure supported on \(\mathcal{M}\) and, given data from (3 ), gives rise to the posterior measure \(\Pi(\cdot|Z^{(N)})\), also supported on \(\mathcal{M}\), obtained as \[d\Pi(a|Z^{(N)})
\propto
e^{\ell_N(a)} d\Pi(a)
,\quad\quad
\ell_N(a)
=
-\frac{1}{2} \sum_{i=1}^N |Y_i-u_{a}(t, X_i)|^2
,~~ a \in \mathcal{M}.\] We now establish the contraction result below for the posterior measure \(\Pi(\cdot| Z^{(N)})\).
Because the support of \(\Phi\), hence also that of \(F\), is contained in a compact subset of \((P_n[L^2(\Omega)], \|\cdot\|_{L^2})\) (see after Theorem
16), and since \(F\) is a continuous map into \(L^2(\Omega)\), then \(\mathcal{M}\) is also a compact subset of \(L^2(\Omega)\). We will need to make this quantitative: Since \(\mathcal{M}=F(D(\Phi))\), \(F\) is \(L^2-L^2(\Omega)\) Lipschitz, and \(D(\Phi)\) is an \(n\)-dimensional ellipsoid in \(L^2(\Omega)\), we can control the covering numbers of \(\mathcal{M}\) as \[\label{CoverEstimateManifold}
N(\mathcal{M}, \|\cdot\|_{L^2},\varepsilon)
\leq
N(D(\Phi), \|\cdot\|_{L^2}, C\varepsilon)
\lesssim
\varepsilon^{-n}
,\quad\quad
\varepsilon>0 ,\tag{12}\] for some \(C>0\), for instance by using results in Sec. 4.3.7 in [27].
We can now prove the following result.
Theorem 5. Assume that \(\theta_0\in \mathcal{A}\) and consider i.i.d. data \(Z^{(N)}=(Y_i, X_i)_{i=1}^N\) obtained through (3 ) for some
\(t>0\) where \(u_\theta\) solves the PDE (4 ) with known \(f \in C^3(\mathbb{R})\) satisfying (5 ).
Let \(\delta_N \equiv \sqrt{\log(N)/N}\). Then, there exists \(M>0\) such that we have in \(P_{\theta_0}^\mathbb{N}\)-probability as \(N\to \infty\)\[\Pi\Big( a\in\mathcal{M}: \|u_a(t)-u_{\theta_0}(t)\|_{L^2} \le M \delta_N~ |~ Z^{(N)} \Big)
\to
1
.\]
Proof of Theorem 4. We apply Theorem 1.3.2 in [25]
with parameter space \(\Theta\equiv \mathcal{M}\) given by the inertial manifold constructed in Section 3.2, constant regularization sets \(\Theta_N \equiv
\mathcal{M}\) satisfying \(\Pi(\Theta_N^c)=0\) since \(\Pi(\mathcal{M})=1\), and \(\delta_N\) as in the statement, noting also that \(u_\theta(t), \theta \in \mathcal{M},\) lies in \(\mathcal{M}\) which is bounded in \(H^2(\Omega) \hookrightarrow L^\infty(\Omega)\). First observe that the
‘intrinsic’ semimetric \(d_\mathcal{G}\) there is given by \[d_\mathcal{G}(\theta, \theta')
\equiv
\|u_\theta(t)-u_{\theta'}(t)\|_{L^2(\Omega)},~~ t>0
.\] Proposition 10 entails that the map \(\theta\mapsto u_\theta(t)\) is \(L^2\)-Lipschitz
continuous, with Lipschitz constant \(C_t>0\), so as in (8 ) we can upper-bound the covering numbers \[\log N(\mathcal{M}, d_\mathcal{G}, \delta_N)
\leq
\log N(\mathcal{M}, \|\cdot\|_{L^2}, \delta_N/C_t) \lesssim N\delta_N^2
,\] by virtue of the estimate (12 ). It remains to establish a small ball estimate for the sets \[\mathcal{B}_N
=
\Big\{ \theta\in \mathcal{M}: \|u_\theta(t)-u_{\theta_0}(t)\|_{L^2} \le \delta_N \Big\}
,\] Defining the ball \[\mathcal{B}'_N
=
\Big\{ \theta\in \mathcal{M}: \|\theta-\theta_0\|_{L^2} \le \delta_N/C_t\Big\}
,\] yields \(\Pi(\mathcal{B}_N)
\ge \Pi(\mathcal{B}'_N).\) Since \(\theta_0\in \mathcal{A}\) and the inertial manifold \(\mathcal{M}=F(D(\Phi))\) contains \(\mathcal{A}\), there
exists \(\bar \theta_0\in D(\Phi)\), \(\bar \theta_0 = \sum_{j=1}^n \beta_j e_j\) say, such that \(\theta_0 = F(\bar \theta_0)\). We now argue that \(\bar\theta_0\) is an interior point of \(D(\Phi)\). Let \(c\ge 1\) be a universal constant such that \(c^{-1} \|\cdot\|_{H^2}\le
\|\cdot\|_{\tilde{H}^2}\le c\|\cdot\|_{H^2}\). Then, \(\|\theta_0\|_{\tilde{H}^2}\le c\|\theta_0\|_{H^2}\le c\rho\) because \(\theta_0\in\mathcal{A}\). Writing \(\theta_0=\bar \theta_0 + \Phi(\bar \theta_0)\), then Parseval’s identity yields \(\|\bar \theta_0\|_{\tilde{H}^2}\le \|\theta_0\|_{\tilde{H}^2}\le c\rho\) since \(\bar
\theta_0\in P_n[L^2(\Omega)]\) and \(\Phi(\bar \theta_0)\) lies in the orthogonal complement of \(P_n[L^2(\Omega)]\). In particular, we have \(\|\bar
\theta_0\|_{H^2}\le c\|\bar \theta_0\|_{\tilde{H}^2}\le c^2 \rho\). Therefore, up to increasing the radius \(8\rho\), featuring in the definition of \(D(\Phi)\) in (10 ), to a value that is larger than \(c^2\rho\), we may assume without loss of generality that \(\bar \theta_0\) belongs to the interior of \(D(\Phi)\). Because \(F\) is \(L^2\)-Lipschitz, with Lipschitz constant \(L_F\), then \[\Pi(\mathcal{B}_N')
=
\mathbb{P}\big(\|F(X)-F(\bar \theta_0)\|_{L^2}\le C^{-1} \delta_N\Big)
\ge
\mathbb{P}\big(\|X-\bar \theta_0\|_{L^2}\le (L_F C)^{-1} \delta_N\big)
.\] Let \(p\) denote the positive continuous density of \(X\) in \(D(\Phi)\). Taking \(N\) large enough so that the
\(L^2\)-ball \(E_N\), centred at \(\bar \theta_0\) with radius \((L_F C)^{-1} \delta_N\), is contained in \(D(\Phi)\), the probability in the last display is of order \[\int_{E_N} p(x)\, dx
=
p(\bar \theta_0) \omega_n \big((L_FC)^{-1} \delta_N\big)^n
+
o(\delta_N^n)
\equiv
L\delta_N^n + o(\delta_N^n)
,\] for some \(L>0\) and where \(\omega_n\) is the volume of the unit ball of \(\mathbb{R}^n\). We thus find for all \(N\) large enough that there exists \(A>0\) such that \[\Pi(\mathcal{B}_N)
\gtrsim
\delta_N^{n}
\ge
e^{-AN\delta_N^2}
.\] Theorem 1.3.2 in [25] now yields the conclusion. ◻
We first prove the reverse Poincaré inequality for the attractor of our PDE, which is the key result.
Proposition 6. There exists a constant \(C>0\) such that \[\|\nabla(\theta-\vartheta)\|_{L^2}
\le
C \|\theta-\vartheta\|_{L^2}
,\quad\quad
\forall\;\theta,\vartheta\in \mathcal{A}
.\]
In fact, it follows from the proof that Proposition 6 also holds whenever \(\theta,\vartheta\in \mathcal{M}\).
Proof of Proposition 6. Theorem 16 entails that there exists \(n\in\mathbb{N}\) and a map \[\Phi : P_n[L^2(\Omega)] \to Q_n[L^2(\Omega)]\] such that any \(u\in\mathcal{A}\) admits the (orthogonal) decomposition \(u=u_n+\Phi(u_n)\) with \(u_n=P_n[u]\), and such that \[\label{AContH2}
\|\Phi(v)-\Phi(w)\|_{H^2}
\le
C\|v-w\|_{L^2}
,\quad\quad
\forall\;v,w\in P_n[L^2(\Omega)]
,\tag{13}\] for some universal constant \(C>0\). For all \(\theta,\vartheta\in\mathcal{A}\), let us write \(\theta = \theta_n +
\Phi(\theta_n)\) and \(\vartheta = \vartheta_n + \Phi(\vartheta_n)\), where \(\theta_n=P_n[\theta]\) and \(\vartheta_n=P_n[\vartheta]\). Then, \[\begin{align}
\|\nabla(\theta-\vartheta)\|_{L^2}
&\lesssim
\|\theta-\vartheta\|_{H^2}
\\[2mm]
&\le
\|\theta_n-\vartheta_n\|_{H^2} + \|\Phi(\theta_n)-\Phi(\vartheta_n)\|_{H^2}
\\[2mm]
&
\lesssim
(1 + \lambda_n) \|\theta_n-\vartheta_n\|_{L^2} + C \| \theta_n - \vartheta_n\|_{L^2}
,
\end{align}\] where we used (13 ) and the equivalence \(\|u\|^2_{H^2} \asymp \sum_{j\ge 0} (1+\lambda_j^2) \left\langle u,e_j \right\rangle_{L^2}^2\) in the last inequality. Because \(\|\theta_n-\vartheta_n\|_{L^2}\le \|\theta-\vartheta\|_{L^2}\), we obtain the reverse Poincaré inequality \[\label{revPoincare}
\|\nabla(\theta-\vartheta)\|_{L^2}
\lesssim
\|\theta-\vartheta\|_{L^2}
,\quad\quad
\forall\;\theta,\vartheta\in \mathcal{A}
,\tag{14}\] which concludes the proof. ◻
The proof of the following theorem (which is Theorem 2 from the introduction) is based on the preceding proposition and inspired by Theorem 1B in [5], but slightly easier since the attractor is invariant under the dynamics so that we know that the reverse Poincaré inequality from Proposition 6 holds true at all times \(t>0\). Inspection of the proof that follows shows that the conclusion holds whenever \(\theta,\vartheta\in\mathcal{M}\), where \(\mathcal{M}\) is the inertial manifold (containing \(\mathcal{A}\)).
Theorem 7. There exists a constant \(c>0\) depending only on \(f\) such that \[\|\theta-\vartheta\|_{L^2} \le e^{ct}
\|u_{\theta}(t)-u_{\vartheta}(t)\|_{L^2} ,\quad\quad \forall\;t\ge 0,\; \forall\;\theta,\vartheta\in \mathcal{A} .\]
Proof of Theorem 7. Let \(w(t)\equiv u_\theta(t)-u_\vartheta(t)\) for all \(t\ge 0\), and
observe that \(w\) satisfies the time-evolution equation \[\label{pseudlin}
\frac{dw}{dt} - \Delta w
=
f(u_\theta(t)) - f(u_\vartheta(t))
\equiv
g(t)
,\quad\quad
w(0) = u_\theta(0) - u_\vartheta(0)
.\tag{15}\] Taking the \(L^2(\Omega)\)-inner product of (15 ) with \(w\) yields \[\frac{1}{2} \frac{d}{dt}\|w(t)\|_{L^2}^2
+ \|\nabla w(t)\|_{L^2}^2 = \left\langle g(t),w(t) \right\rangle_{L^2}
.\] Recall from Proposition 12 that the global attractor \(\mathcal{A}\) is bounded in \(H^2(\Omega)\); the same holds for the inertial manifold \(\mathcal{M}\) (see after Theorem 16). By the
Sobolev embedding \(H^2(\Omega)\hookrightarrow L^\infty(\Omega)\), let \(R>0\) be a uniform \(L^\infty\)-bound of \(\mathcal{M}\) hence also \(\mathcal{A}\), by virtue of the inclusion \(\mathcal{A}\subset \mathcal{M}\) (see after Theorem 16). Also recall from (18 ) that the attractor \(\mathcal{A}\) is invariant under the dynamics so that \(u_\theta(t)\in \mathcal{A}\) and \(u_\vartheta(t)\in \mathcal{A}\) for all \(t\ge 0\) whenever \(\theta,\vartheta\in\mathcal{A}\); similarly, when \(\theta, \vartheta\in\mathcal{M}\) we have \(u_\theta(t),u_\vartheta(t)\in\mathcal{M}\) for all \(t\ge 0\) by invariance of the inertial manifold (see (23 )). In particular, we have \[\|u_\theta(t)\|_{L^\infty}
\le
R,\quad\quad
\text{and}
\quad\quad
\|u_\vartheta(t)\|_{L^\infty}\le R
,\quad\quad
\forall\;t\ge 0
.\] Since \(f\) is continuously differentiable over \(\mathbb{R}\), it is Lipschitz on the interval \([-R, R]\) to the effect that \[|g(t)|
\le
\|f'\|_{L^\infty([-R, R])} |u_\theta(t) - u_\vartheta(t)|
\equiv
c_f |w(t)|
.\] Consequently, we have \[\frac{1}{2} \frac{d}{dt}\|w(t)\|_{L^2}^2
\ge
- \|\nabla w(t)\|_{L^2}^2 - c_f \|w(t)\|_{L^2}^2
.\] The invariance property of \(\mathcal{A}\) and \(\mathcal{M}\) under the dynamics also implies that the reverse Poincaré inequality in Proposition 6 applies, so that \[\|\nabla(u_\theta(t)-u_\vartheta(t))\|_{L^2}
\le
c_P\|u_\theta(t)-u_\vartheta(t)\|_{L^2}
,\quad\quad
\forall\;t\ge 0
,\] for some universal constant \(c_P>0\) independent of \(\theta\) and \(\vartheta\). This entails that \[\frac{1}{2}
\frac{d}{dt}\|w(t)\|_{L^2}^2
\ge
-c_P^2\|w(t)\|^2_{L^2} - c_f \|w(t)\|_{L^2}^2
=
-(c_P^2+c_f)\|w(t)\|_{L^2}^2
,\quad\quad
\forall\;t\ge 0
.\]Grönwall’s inequality then implies that \[\label{keybd}
\|w(t)\|_{L^2}^2
\ge
\|w(0)\|_{L^2}^2 e^{-(c_P^2+c_f)t}
,\quad\quad
\forall\; t\ge 0,\tag{16}\] which yields the conclusion. ◻
Since the posterior measure in Theorems 4 and 5 is supported on
\(\mathcal{A}\) or \(\mathcal{M}\), respectively, the previous theorem implies:
Theorem 8. Let \(\theta_0 \in \mathcal{A}\), let \(\delta_N \equiv \sqrt{\log(N)/N}\) and let \(\Pi(\cdot|Z^{(N)})\) be the posterior
measure arising from data \(Z^{(N)}=(X_i,Y_i)_{i=1}^N\) as in (3 ) with \(t>0\) and prior \(\Pi\) given either by (7 ) or (11 ). Then there exists \(M>0\) large enough such that \[\Pi\Big( \theta : \|\theta-\theta_0\|_{L^2} \le M \delta_N~ |~ Z^{(N)} \Big)
\to
1
,\] in \(P_{\theta_0}^\mathbb{N}\)-probability as \(N\to \infty\).
Corollary 9. Let \(\theta_0\in \mathcal{A}\) and \(\delta_N = \sqrt{\log(N)/N}\). There exists a sequence of \(\mathcal{A}\)-valued
estimators \(\hat{\theta}_N\) such that \(\|\hat{\theta}_N-\theta_0\|_{L^2} = O_{P_{\theta_0}^\mathbb{N}}(\delta_N)\) as \(N\to \infty\).
Theorem 1 now follows from this corollary and Proposition 10 below.
Proof of Corollary 9. We follow the argument in Proposition 6.7 of [23] and assume \(M=1\) in Theorem 8 (otherwise one re-defines \(\delta_N\) appropriately). Let \(\Pi(\cdot|Z^{(N)})\) be the posterior measure arising from data \(Z^{(N)}=(X_i,Y_i)_{i=1}^N\) as in (3 )
and \(\varepsilon\)-net prior \(\Pi\) given by (11 ). Denote by \(B(\theta, r)\) the closed ball in the metric space \((\mathcal{A}, \|\cdot\|_{L^2})\) centred at \(\theta\in \mathcal{A}\) with radius \(r>0\). For any \(\theta\in \mathcal{A}\),
let \[r_N(\theta) = \inf\Big\{ r > 0 : \Pi(B(\theta, r) | Z^{(N)}) \ge \frac{1}{2} \Big\} .\] Theorem 8 entails that \(\Pi(B(\theta_0, \delta_N)|Z^{(N)})\to 1\) in \(P_{\theta_0}^\mathbb{N}\)-probability, so that \[r_N(\theta_0)\le \delta_N + o_{P_{\theta_0}^\mathbb{N}}(1) .\]
Among all the \(\theta\)’s in \(\mathcal{A}\) for which \(r_N(\theta)<\infty\), pick \(\hat{\theta}_N\in\mathcal{A}\)
such that the corresponding ball has minimal radius up to a \(\delta_N\)-error. In particular, we have \[r_N(\hat{\theta}_N) \le r_N(\theta_0)+\delta_N \le 2\delta_N +
o_{P_{\theta_0}^\mathbb{N}}(1) .\] Now, with \(P_{\theta_0}^\mathbb{N}\)-probability approaching \(1\) then \(B(\theta_0, r_N(\theta_0))\) and \(B(\hat{\theta}_N, r_N(\hat{\theta}_N))\) cannot be disjoint, otherwise the posterior mass of their union would be approaching \(1+1/2(>1)\) in \(P_{\theta_0}^\mathbb{N}\)-probability as \(N\to\infty\). In particular, this provides \[\|\hat{\theta}_N - \theta_0\|_{L^2} \le r_N(\theta_0) + r_N(\hat{\theta}_N) +
o_{P_{\theta_0}^\mathbb{N}}(1) \le 3\delta_N + o_{P_{\theta_0}^\mathbb{N}}(1) ,\] which concludes the proof. ◻
3 Background material and proofs of Theorem 2 and Proposition 6↩︎
3.1 Reaction-diffusion equations, absorbing sets, and the global attractor↩︎
Throughout this section, we will assume that \(f:\mathbb{R}\to \mathbb{R}\) is a function in \(C^3(\mathbb{R})\) satisfying (5 ) for some \(p>2\). The following result follows from Theorem 1.1 in [28]; see also Sections 8.3-8.4 in [3]. When \(\theta\in L^2(\Omega)\), we say that the solution is weak when (4 ) holds as an equality in
\(L^q([0,T], H^{-\kappa}(\Omega))\) for any \(\kappa > (p-2)/p\), where \(q\) is conjugate to \(p\).
Proposition 10. For all \(\theta\in L^2(\Omega)\), the system of equations (4 ) admits a unique weak solution \(u_{\theta}\in C^0([0,\infty),
L^2(\Omega))\) such that \(u_{\theta}\in L^2([0,T], H^1(\Omega))\cap L^p([0,T]\times \Omega)\) for all \(T>0\). In addition, there exists a constant \(c=c(f)>0\) such that \[\|u_{\theta}(t)-u_{\vartheta}(t)\|_{L^2}
\le
e^{ct} \|\theta-\vartheta\|_{L^2}
,\quad\quad
\forall\;t\ge 0,\;\theta,\vartheta\in L^2(\Omega)
.\] If \(\theta\in H^1(\Omega)\cap L^p(\Omega)\), then \(u_{\theta}\) is a strong solution and we have \(u_{\theta}\in C^0([0,\infty), H^1(\Omega))\)
with \(u_{\theta}\in L^2([0,T], H^2(\Omega))\cap L^\infty([0,T], L^p(\Omega))\) for all \(T>0\).
We say that the solution is strong whenever (4 ) holds in \(L^2([0,T]\times \Omega)\).
We next show that the dynamical system \(S(t)\) has an absorbing set in \(H^1\).
Proposition 11. There exists a constant \(C=C(f)>0\) such that for all \(\theta\in L^2(\Omega)\) we have \(\|u_\theta(t)\|_{H^1}\le
C\) for all \(t\ge t_0\), for some \(t_0=t_0(\|\theta\|_{L^2})\).
Proof of Proposition 11. The proof follows the same lines as that of Proposition 11.1 and 11.3 in [3] once we establish that for some constants \(b,c>0\), \[\label{eq:PDERob}
\frac{d}{dt}\|u_\theta\|^2_{L^2} + 2c\|u_\theta\|^2_{L^2} \le 2b ,\quad\quad \forall\;t>0.\tag{17}\] To show (17 ) in our periodic setting, we cannot rely on the standard Poincaré inequality as in [3]. Instead, let us use the growth condition (5 ) on \(f\) and that \(|\Omega|=1\): As in (11.6) in [3] we have \[\frac{1}{2} \frac{d}{dt}\|u_\theta\|^2_{L^2} +
\|\nabla u_\theta\|^2_{L^2} + \alpha_2 \|u_\theta\|^p_{L^p} \le k ,\quad\quad \forall\;t\ge 0 .\] Because \(p\ge 2\), Hölder’s and Young’s inequalities provide \[\|u_\theta\|^2_{L^2} \le
\|u_\theta\|^2_{L^p} \le \frac{2}{p} \|u_\theta\|^p_{L^p} + \frac{p-2}{p} ,\] using again that \(|\Omega|=1\). Combining the preceding displays we deduce that \[\frac{1}{2}
\frac{d}{dt}\|u_\theta\|^2_{L^2} + \|\nabla u_\theta\|^2_{L^2} + \frac{2\alpha_2}{p} \|u_\theta\|^2_{L^2} - \frac{\alpha_2(p-2)}{2} \le k ,\quad\quad \forall\;t\ge 0 .\] Since \(\alpha_2>0\), then dropping the
term \(\|\nabla u_\theta\|_{L^2}\) in the last display yields (17 ) hence concludes the proof. ◻
From the previous proposition we can deduce:
Proposition 12. Let \(S(t)=u_\theta(t)\) be the solution operator for (4 ) for \(f\) satisfying (5 ). Then, the
semidynamical system \((L^2(\Omega),\{S(t)\}_{t \geq 0})\) admits a global attractor \(\mathcal{A}\) as defined in Definition 10.4 in [3]; moreover \(\mathcal{A}\) is bounded in \(H^2(\Omega)\), and given by formula (2 ).
Proof. That a global attractor in the sense of Definition 10.4 in [3] exists follows from Theorem 10.5 in [3], and that it coincides with our definition follows from eq. (10.10) in the same reference. The proof of \(H^2\)-boundedness follows
as in the proof of Theorem 11.7 in [3]. ◻
It follows in particular that \(\mathcal{A}\) is invariant under the dynamics; \[\label{invA} u_\theta(t)\in \mathcal{A} ,\quad\quad \forall\;t\ge
0,\;\forall\;\theta\in\mathcal{A},\tag{18}\] and independent of the choices of the absorbing set.
While the global attractor \(\mathcal{A}\) describes the precise limiting states of our reaction diffusion equation, it has a complicated analytical structure. Instead we shall work with a slightly larger manifold \(\mathcal{M}\) that contains \(\mathcal{A}\), that shares many properties with \(\mathcal{A}\) but is analytically more tractable (once its existence is
shown).
To construct \(\mathcal{M}\) for our PDE we will follow arguments developed in [29], and this requires some auxiliary
results which we provide now.
First we will need to upgrade Proposition 11 to the effect that the semidynamical system \((H^2(\Omega), \{ S(t) : t\ge 0\})\) admits a
bounded \(H^2\)-absorbing set – note that this is a stronger requirement than the boundedness of \(\mathcal{A}\) in \(H^2(\Omega)\).
Theorem 13. There exists a constant \(C=C(f)>0\) such that for all \(\theta\in H^1(\Omega)\) we have \(\|u_\theta(t)\|_{H^2}\le
C\) for all \(t\ge \tilde{t}\), for some \(\tilde{t}=\tilde{t}(\|\theta\|_{H^1})\).
Proof of Theorem 13. Proposition 11 entails that there exists a constant
\(c_1>0\) independent of \(\theta\) such that \[\label{gradest}
\|u_\theta(t)\|_{L^2}+\|\nabla u_\theta(t)\|_{L^2} \le c_1
,\quad\quad
\forall\;t\ge t_0\equiv t_0(\|\theta\|_{H^1})
,\tag{19}\] where \(t_0\) only depends on \(\|\theta\|_{H^1}\). Taking the \(L^2\)-inner product of (4 ) with
\(-\Delta u_\theta\) yields \[\frac{1}{2} \frac{d}{dt}\|\nabla u_\theta\|^2_{L^2}
+
\|\Delta u_\theta\|^2_{L^2}
=
-\left\langle f(u_\theta),\Delta u_\theta \right\rangle_{L^2}
\lesssim
\|\nabla u_\theta\|^2_{L^2},\] arguing as on p.227f. in [3], using also the second part of the hypothesis (5 ) for
\(f\). Integrating the last display between \(t\) and \(t+1\) and using (19 ) yields \[\label{laplaceest}
\|\nabla u_\theta(t+1)\|^2_{L^2}
+
\int_{t}^{t+1} \|\Delta u_\theta(s)\|^2_{L^2}\, ds
\lesssim
c_1 + \int_{t}^{t+1} \|\nabla u_\theta(s)\|^2_{L^2}\, ds
\le 2c_1 \equiv \;c_2
,\quad\quad
\forall\;t\ge t_0\tag{20}\] for some \(c_2>0\) independent of \(\theta\). Adapting the arguments in Section 11.2.1 of [3] to our periodic situation one can prove that \[\|u_\theta(t)\|_{L^\infty}
\le
c_3
,\quad\quad
\forall\;t\ge t_1(\|\theta\|_{H^1})
,\] with \(c_3>0\) a constant independent of \(\theta\). Now taking the inner product of (4 ) with \(\Delta^2
u_\theta\) yields \[\begin{align}
\frac{1}{2} \frac{d}{dt}\|\Delta u_\theta\|^2_{L^2} -\langle \Delta u_\theta, \Delta^2 u_\theta \rangle_{L^2} &= \langle f(u_\theta), \Delta^2 u_\theta \rangle_{L^2}
\end{align}\] so that by integration by parts, and the Cauchy-Schwarz and Young inequalities, \[\frac{1}{2} \frac{d}{dt}\|\Delta u_\theta\|^2_{L^2}
+
\|\nabla \Delta u_\theta\|^2_{L^2} \\
= -\left\langle \nabla f(u_\theta),\nabla \Delta u_\theta \right\rangle_{L^2} \le
\frac{1}{2}\|\nabla (f(u_\theta))\|^2_{L^2} + \frac{1}{2} \|\nabla \Delta u_\theta\|^2_{L^2}.\] Rearranging and dropping the \(\|\nabla \Delta u_\theta\|_{L^2}\) term yields \[\frac{d}{dt}\|\Delta u_\theta\|^2_{L^2}
\le
\|\nabla(f(u_\theta))\|^2_{L^2}
.\] We now use the identity \(\nabla (f(u_\theta))=f'(u_\theta) \nabla u_\theta\) established in the proof of Lemma 14, and the \(L^\infty\) bound above on \(u_\theta(t)\), to obtain \[\|\nabla(f(u_\theta))\|_{L^2}
=
\|f'(u_\theta) \nabla u_\theta\|_{L^2}
\le
\|f'(u_\theta)\|_{L^\infty} \|\nabla u_\theta\|_{L^2}
\le
\|f'\|_{L^\infty([-c_3, c_3])} \|\nabla u_\theta\|_{L^2}
,\] for all \(t\ge t_1\). This yields \[\frac{d}{dt}\|\Delta u_\theta\|^2_{L^2}
\lesssim
\|\nabla u_\theta\|^2_{L^2}
.\] Now integrating the last display between \(s\) and \(t\), \(s \le t\), and using Fubini’s theorem provides \[\|\Delta
u_\theta(t)\|^2_{L^2}
\lesssim
\|\Delta u_\theta(s)\|^2_{L^2}
+
\int_s^t \|\nabla u_\theta(\tau)\|^2_{L^2}\, d\tau
.\] Integrating in \(s\) between \(t-1\) and \(t\) as well as using (19 ) and (20 ) yields
\[\|\Delta u_\theta(t)\|^2_{L^2}
\lesssim
\int_{t-1}^t \|\Delta u_\theta(s)\|^2_{L^2}
+
\int _{t-1}^t \|\nabla u_\theta(s)\|^2_{L^2}\,ds
\le c_1+c_2
,~~
\forall\;t\ge 1+\max\{t_0, t_1\}
.\] Combining this (19 ) entails that \[\|u_\theta(t)\|_{H^2}
\lesssim \|u_\theta(t)\|_{L^2} + \|\Delta u_\theta(t)\|_{L^2} \le c\] for all \(t\) large enough depending only on \(\|\theta\|_{H^1}\) and some \(c>0\), which yields the conclusion. ◻
Lemma 14. Let \(h\in C^2(\mathbb{R})\) and \(u\in H^2(\Omega)\), \(\Omega = [0,1]^2\). Then, \(h(u)\in
H^2(\Omega)\), and for some universal constant \(c>0\), we have \[\|h(u)\|_{H^2} \lesssim M_h(\|u\|_{H^2}) ,\] where \(M_h(s) \equiv
\|h\|_{L^\infty([-cs, cs])} + \|h'\|_{L^\infty([-cs, cs])}s + \|h''\|_{L^\infty([-cs,cs])} s^2\) for all \(s\ge 0\).
Proof of Lemma 14. First note that \(u\in H^2(\Omega)\hookrightarrow L^\infty(\Omega)\) by the Sobolev embedding, so that we can write
\[\label{Richardwantsnumbers}
|h(u)|
\le
|h(u)| 1_{\{|u|\le c \|u\|_{H^2}\}}
\le
\|h\|_{L^\infty([-c\|u\|_{H^2}, c\|u\|_{H^2}])}
<
\infty
,\tag{21}\] where \(c>0\) is any number such that \(\|u\|_{L^\infty(\Omega)}\le c \|u\|_{H^2}\), and noting that \(h\) is continuous
over \(\mathbb{R}\). It follows that \(h(u)\in L^\infty(\Omega)\) hence also \(h(u)\in L^2(\Omega)\) since \(\Omega\) is
bounded. Thus, it remains to establish that \(\Delta(h(u))\in L^2(\Omega)\). We have \(\nabla (h(u))=h'(u) (\nabla u)\) by virtue of the chain rule for weak derivatives (see, e.g.,
Proposition 9.5 in [30]). Note that for all \(v,w\in H^1(\Omega)\), we have \(v \nabla w, w
\nabla v \in L^1(\Omega)\) so that Propostion 9.4 in [30] provides \(\nabla(vw)=v \nabla w + w\nabla v\) in the
weak sense. Consequently, if we show that \(h''(u)\nabla u\in L^1(\Omega)\) and \(h'(u)\Delta u\in L^1(\Omega)\), this will yield \(\Delta (h(u))=
h''(u)|\nabla u|^2+h'(u)\Delta u\) in the weak sense. Since we also need to show that the r.h.s. of the last equality belongs to \(L^2(\Omega)\), it suffices to argue that each term in fact belongs to
\(L^2(\Omega)\hookrightarrow L^1(\Omega)\). For the first term, arguing as above implies that \(h''(u)\in L^\infty(\Omega)\), and the Sobolev embedding \(H^1(\Omega)\hookrightarrow L^4(\Omega)\) implies that \(|\nabla u|^2\in L^2(\Omega)\) since \(\nabla u\in H^1(\Omega)\). For the second term, we have \(h'(u)\in L^\infty(\Omega)\) and \(\Delta u\in L^2(\Omega)\) since \(u\in H^2(\Omega)\). We deduce that \(h(u)\in
H^2(\Omega)\), and the expression obtained for \(\Delta (h(u))\) implies that \[\begin{align}
\|h(u)\|_{H^2}
&\lesssim
\|h(u)\|_{L^2} + \|\Delta (h(u))\|_{L^2}
\\[2mm]
&\lesssim
\|h\|_{L^\infty([-c\|u\|_{H^2}, c\|u\|_{H^2}])}
+
\|h''(u)\|_{L^\infty} \|\nabla u\|^2_{L^4} + \|h'(u)\|_{L^\infty} \|\Delta u\|_{L^2}
\\[2mm]
&\lesssim
\|h\|_{L^\infty([-c\|u\|_{H^2}, c\|u\|_{H^2}])}
+
\|h'(u)\|_{L^\infty} \|u\|_{H^2}
+
\|h''(u)\|_{L^\infty} \|u\|^2_{H^2}
,
\end{align}\] by virtue of the Sobolev embedding \(L^4(\Omega)\hookrightarrow H^1(\Omega)\) applied to \(\nabla u\). Arguing as in (21 ) to bound
\(\|h''(u)\|_{L^\infty}\) and \(\|h'(u)\|_{L^\infty}\) thus yields \[\|h(u)\|_{H^2}
\lesssim
M_h(\|u\|_{H^2})
,\] with \(M_h\) as in the statement. ◻
Lemma 15. Let \(g\in C^3(\mathbb{R})\) and \(u\in H^2(\Omega)\). Then we have \[\|g'(u)v\|_{L^2}
\lesssim
M_0(\|u\|_{H^2}) \|v\|_{L^2}
,\quad\quad
v\in L^2(\Omega)
,\] where \(M_0(s)
\equiv
\|g'\|_{L^\infty([-cs, cs])}\) for all \(s\ge 0\) and some universal \(c>0\), and \[\|g'(u)v\|_{H^2}
\lesssim
M_1(\|u\|_{H^2}) \|v\|_{H^2}
,\quad\quad
v\in H^2(\Omega)
,\] where \(M_1(s)
\equiv
\|g'\|_{L^\infty([-cs, cs])}
+
\|g''\|_{L^\infty([-cs, cs])}s + \|g'''\|_{L^\infty([-cs,cs])}s^2\) for all \(s\ge 0\).
Proof of Lemma 15. For the first inequality, we bound \[\|g'(u)v\|_{L^2}
\le
\|g'(u)\|_{L^\infty} \|v\|_{L^2}
,\] and the first inequality in the proof of Lemma 14 applied to \(h\equiv g'\) yields the desired bound. For the second inequality, we use
the multiplier inequality for Sobolev norms to the effect that \[\|g'(u)v\|_{H^2}
\lesssim
\|g'(u)\|_{H^2} \|v\|_{H^2}
,\] and Lemma 14 applied to \(h\equiv g'\) concludes the proof. ◻
We can now state and prove the main theorem of this section. Recall that \(P_n\) is the \(L^2\)-projector from (9 ) and \(Q_n= {\rm Id} -
P_n\).
Theorem 16. Let \(\mathcal{A}\) be the global attractor from Proposition 12. There exists \(n\in\mathbb{N}\) and a map \(\Phi : P_n[L^2(\Omega)]\to Q_n[L^2(\Omega)]\cap H^2(\Omega)\) such that (i) any \(u\in\mathcal{A}\) writes as \(u=u_n + \Phi(u_n)\) with \(u_n=P_n[u]\), and (ii) there exists a constant \(C>0\) such that \[\label{AContH2State}
\|\Phi(v)-\Phi(w)\|_{H^2}
\le
C\|v-w\|_{L^2(\Omega)}
,\quad\quad
\forall\;v,w\in P_n[L^2(\Omega)]
.\qquad{(1)}\]
Observe that Theorem 16 does not imply that any \(u\in L^2(\Omega)\) can be decomposed as \(u = P_n[u] +
\Phi(P_n[u])\), otherwise the \(H^2\)-Lipschitz continuity in item (ii) could not hold. In fact, \(\Phi\) need not even be linear, and one can further show that \(\Phi(v)=0\) when \(\|v\|_{H^2}>8\rho\), where \(\rho\) is an \(H^2\)-bound on \(\mathcal{A}\) (see Proposition 12).
One can now define the inertial manifold as the ‘graph’ of \(\Phi\)\[\label{defInerMan} \mathcal{M} \equiv \Big\{ v + \Phi(v) : v\in
P_n[L^2(\Omega)],~ \|v\|_{H^2}\le 8\rho \Big\} .\tag{22}\] The properties of \(\Phi\) imply that \(\mathcal{M}\) contains \(\mathcal{A}\) and
is an \(L^2(\Omega)\)-Lipschitz manifold bounded in \(H^2(\Omega)\). The proof of Theorem 16 further
entails that \(\mathcal{M}\) is invariant under the solution operators \(S(t)\) of the reaction-diffusion system (4 ), i.e.\[\label{invM} u_\theta(t) \in \mathcal{M} ,\quad\quad \forall\;t\ge 0,\;\forall\;\theta\in\mathcal{M} ;\tag{23}\] this follows from the construction of \(\mathcal{M}\) (see (iii) in Section 3 of [29]).
Proof of Theorem 16. The existence of a map \(\Phi\) with the required properties follows from Theorem 3.1 in [29]. First, we rewrite the reaction-diffusion system (4 ) as \[\label{PDEA}
\frac{du}{dt} + A[u] = g(u)
,\tag{24}\] where \(A[u] \equiv (\delta {\rm Id} - \Delta)u\) with \(\delta>0\), and \(g(s) \equiv f(s) + \delta s\). Then, \(A\) is linear, self-adjoint on \(L^2(\Omega)\) and bijective from \(H^2(\Omega)\) to \(L^2(\Omega)\), and \(g\) satisfies the same assumptions as \(f\) in (5 ) for different values for the parameters \(k, \alpha_1, \alpha_2\). We will now check
the conditions of Section 2 in [29], with operator \(A\) as above, Hilbert space \(H\equiv
L^2(\Omega)\), dense domain \(D(A)\equiv H^2(\Omega)\), eigenvalues \(\sigma_j \equiv \delta + \lambda_j\) such that \(0<\sigma_0\le \sigma_1 \le \ldots
\sigma_j \asymp j\) as \(j\to\infty\) (provided \(\delta > 0(=\lambda_0)\)), and functional \(R[u]\equiv -g(u)\). Since \(g\in C^2(\mathbb{R})\), Lemma 14 provides \(g(u)\in H^2(\Omega)\) for any \(u\in
H^2(\Omega)\), to the effect that \(R\) maps \(H^2(\Omega)\) into \(H^2(\Omega)\). In addition, \(R\) is Fréchet
differentiable from \(D(A)\equiv H^2(\Omega)\) into \(D(A^{1-\beta})\equiv H^2(\Omega)\) with \(\beta=0\); see Proposition 6.4 in [11]. Since \(g\in C^3(\mathbb{R})\), Lemma 15 implies that
(2.1a) and (2.1b) in [29] hold with \(\beta=0\), and non-negative and monotone non-decreasing functions \(M_0\) and \(M_1\) featuring in Lemma 15. Theorem 13 entails that \((H^2(\Omega), \{S(t) : t\ge 0\})\) admits an absorbing ball in \(H^2(\Omega)\), i.e. there exists an \(H^2\)-ball \(B\) such that the image of any \(H^2\)-ball through \(S(t)\) is included in \(B\)
for all \(t\) large enough. The sequence \(\{\lambda_{k+1}-\lambda_k : k\ge 0\}\) is unbounded (see Section 15.4.2 in [3]) so that, for any \(C>0\), there exists \(n\in\mathbb{N}\) such that \[\sigma_{n+1}-\sigma_n
=
\lambda_{n+1}-\lambda_n
>
C
,\] which yields (3.28) in [31]. Consequently, Theorem 3.1 there yields the existence of a map \(\Phi\) as in
the statement with \(n\) as above. ◻
Acknowledgements. The authors would like to thanks Soufiane Noubir for various interesting discussions about the content of this article, and further gratefully acknowledge funding from an ERC Advanced Grant (UKRI G116786) and EPSRC
programme grant EP/V026259.
Department of Pure Mathematics & Mathematical Statistics
A. V. Babin and M. I. Vishik, Translated and revised from the 1989 Russian original by BabinAttractors of evolution equations, vol. 25. North-Holland Publishing Co., Amsterdam,
1992, p. x+532.
J. Robinson, Infinite-dimensional dynamical systems: An introduction to dissipative parabolic PDEs and the theory of global attractors. Cambridge Univ. Press,
2001.
[4]
R. Nickl, “Consistent inference for diffusions from low frequency measurements,”Ann. Statist., vol. 52, no. 2, pp. 519–549, 2024, doi: 10.1214/24-aos2357.
[5]
R. Nickl and E. Titi, “On posterior consistency of data assimilation with Gaussian process priors: The 2DNavier-Stokes
equations,”Ann. Statist., vol. 52, no. 4, pp. 1825–1844, 2024.
[6]
R. Nickl, “Bernstein-von Mises theorems for time evolution equations,”arXiv preprint arXiv:2407.14781, 2024.
[7]
R. Nickl, G. A. Pavliotis, and K. Ray, “Bayesian nonparametric inference in McKean-Vlasov models,”Ann. Statist., vol. 53, no.
1, pp. 170–193, 2025, doi: 10.1214/24-aos2459.
[8]
G. Koers, B. Szabo, and A. W. van der Vaart, “Linear methods for non-linear inverse problems,”Annals of Statistics, pp. to appear, 2025.
[9]
G. S. Alberti, D. Barnes, A. Jambhale, and R. Nickl, “On low frequency inference for diffusions without the hot spots conjecture,”Math. Stat. Learn., vol. 8, no.
3–4, pp. 305–322, 2025, doi: 10.4171/msl/53.
[10]
M. Giordano and S. Wang, “Statistical algorithms for low-frequency diffusion data: A PDE approach,”Ann. Statist., vol. 53, no. 3, pp. 1150–1175, 2025,
doi: 10.1214/25.
[11]
D. Konen, “Inverting the Fisher information operator in non-linear models,”Arxiv preprint arXiv:2601.13254, 2026.
[12]
R. Nickl and F. Seizilles, “Inferring diffusivity from killed diffusion,”Annals of Statistics, pp. to appear, 2026.
[13]
A. Magra, F. van der Meulen, and A. van der Vaart, “Semi-parametric bernstein-von mises theorem in a a parabolic PDE problem,”Arxiv preprint,
2026.
[14]
A. Castre and R. Nickl, “On gradient stability in nonlinear PDE models and inference in interacting particle systems,”arXiv preprint, 2026.
[15]
A. M. Stuart, “Inverse problems: A Bayesian perspective,”Acta Numer., vol. 19, pp. 451–559, 2010.
[16]
S. L. Cotter, M. Dashti, J. C. Robinson, and A. M. Stuart, “Bayesian inverse problems for functions and applications to fluid mechanics,”Inverse Problems, vol. 25,
no. 11, pp. 115008, 43, 2009.
[17]
S. Reich and C. Cotter, Probabilistic forecasting and Bayesian data assimilation. Cambridge University Press, New York, 2015, p. x+297.
[18]
K. Law, A. Stuart, and K. Zygalakis, Data assimilation. Springer, Cham, 2015, p. xviii+242.
[19]
D. Konen and R. Nickl, “Data assimilation with the \(2D\) navier-stokes equations: Optimal gaussian asymptotics for the posterior measure,”Annals of Statistics, pp. to appear, 2026.
[20]
K. Abraham and R. Nickl, “On statistical Calderón problems,”Math. Stat. Learn., vol. 2, no. 2, pp. 165–216, 2019.
[21]
J. Bohr, “A Bernstein–von-Mises theorem for the Calderón problem with piecewise constant conductivities,”Inverse Problems,
vol. 39, no. 1, pp. Paper No. 015002, 18, 2023.
[22]
A. W. van der Vaart, Asymptotic statistics. Cambridge: Cambridge Univ. Press, 1998.
[23]
S. Ghosal and A. van der Vaart, Fundamentals of nonparametric Bayesian inference, vol. 44. Cambridge University Press, Cambridge, 2017, p. xxiv+646.
[24]
F. Monard, R. Nickl, and G. P. Paternain, “Consistent inversion of noisy non-Abelian X-ray transforms,”Comm. Pure Appl. Math., vol. 74,
no. 5, pp. 1045–1099, 2021.
A. V. Babin and M. I. Vishik, “Attractors of evolution partial differential equations and estimates of their dimension,”Uspekhi Mat. Nauk, vol. 38, no. 4(232), pp.
133–187, 1983.
[27]
E. Giné and R. Nickl, Mathematical foundations of infinite-dimensional statistical models. New York: Cambridge University Press, 2016.
[28]
M. Marion, “Attractors for reaction–diffusion equations: Existence and estimate of their dimension,”Applicable Analysis, vol. 25, no. 1–2, pp. 101–147, 1987, doi:
10.1080/00036818708839678.
[29]
C. Foias, G. R. Sell, and E. S. Titi, “Exponential tracking and approximation of inertial manifolds for dissipative nonlinear equations,”Journal of Dynamics and
Differential Equations, vol. 1, no. 2, pp. 199–241, 1989, doi: 10.1007/BF01046904.
C. Foias, G. R. Sell, and R. Temam, “Inertial manifolds for nonlinear evolutionary equations,”Journal of Differential Equations, vol. 73, no. 2, pp. 309–353, 1988,
doi: 10.1016/0022-0396(88)90110-6.