July 13, 2026
We introduce a tangent-space multiscale manifold method for nonlinear heterogeneous elliptic problems. The method represents the fine-scale solution by a nonlinear reconstruction of a coarse state. The ideal reconstruction eliminates fine scales through a constrained variational problem, and the computable reconstruction approximates this map by localized nonlinear patch solves blended with a partition of unity. Because the approximation set is a nonlinear manifold, the coarse equation is posed with tangent multiscale test functions. We also formulate a network-interpolated variant in which only the restricted patch outputs used by the partition-of-unity blend, together with their tangent actions, are approximated by local learned maps. For heterogeneous monotone nonlinear diffusion, we record the structural monotonicity, differentiability, patch-map regularity, and conditional perturbation estimates that separate the geometric stability mechanism from localization, residual, and optional learning defects. A rigorous a priori theory for the decay of the localization defect, and the resulting convergence rates in the coarse mesh size, is deferred to a separate analysis; here these defects are controlled conditionally and their decay is demonstrated numerically.
Keywords: Multiscale problems; Numerical homogenization; Nonlinear elliptic problems; Neural networks; Rough coefficients
Multiscale elliptic problems with rapidly varying coefficients arise in porous media, composite materials, heterogeneous diffusion, and nonlinear continuum models. Direct resolution on the finest scale is often too expensive, while standard coarse finite element spaces do not contain the oscillatory response induced by the coefficient field. We focus on applications with general rough coefficient fields, without assumed periodicity, self-similarity, or scale separation. In this setting, classical analytical homogenization, which derives an effective equation from precisely such structural assumptions, does not apply. The method developed in this paper instead belongs to a class of numerical homogenization methods that construct problem-adapted coarse approximations computationally from the actual coefficient field and require no such structural assumptions. For nonlinear constitutive laws the fine-scale response is also state dependent. A useful coarse method must therefore retain the variational structure of the fine-scale problem, localize the fine-scale work, and keep the online unknowns on a low-dimensional coarse space; this coarse–fine viewpoint is closely related to variational multiscale methods, bubble-function enrichments, and residual-based stabilized finite element methods [1]–[6].
Classical multiscale finite element and heterogeneous multiscale methods incorporate local fine-scale information through adapted basis functions, cell problems, or effective coefficients [7]–[9]. A complementary principle is to eliminate unresolved scales by a variational splitting of the fine finite element space and to approximate the fine-scale part by local problems. This localization idea appears already in the adaptive variational multiscale method of Larson and Målqvist [10], where the fine-scale part is approximated by decoupled local problems in a slice space and controlled by a posteriori estimates. For linear elliptic problems, the subsequent work [11] gave the localized corrector construction and a priori localization theory that led to the localized orthogonal decomposition framework; see, for example, [12], [13]. Efficient implementations and related localized fine-scale elimination techniques are now well established for linear heterogeneous problems [14]. Related numerical-homogenization frameworks localize the fine-scale response in other ways: the constraint energy minimizing generalized multiscale finite element method builds exponentially decaying multiscale basis functions by local energy minimization [15], while operator-adapted wavelets and the associated gamblet decomposition recover the fine-scale solution through an optimal-recovery, game-theoretic formulation [16], [17].
Nonlinear problems change the structure of fine-scale elimination: the unresolved correction is no longer a linear operator applied to a coarse basis function or residual, but depends on the current coarse state. Several multiscale frameworks have been extended to nonlinear problems. Multiscale finite element methods were developed for nonlinear elliptic problems in [18], and the generalized multiscale finite element method for nonlinear elliptic equations in [19]. Within the heterogeneous multiscale method, quasilinear elliptic homogenization problems were analyzed in [20] and solved by an offline–online strategy in [21]. Within the localized orthogonal decomposition framework, corrector-based constructions have been given for semilinear elliptic [22] and parabolic [23] problems, and for genuinely nonlinear diffusion, through linearization, in [24], [25]. These approaches treat the state dependence in different ways. The nonlinear multiscale finite element method solves nonlinear local problems elementwise, as nonlinear analogues of basis maps with boundary data prescribed by the coarse function, tested against standard coarse functions, with a convergence theory set in the classical homogenization framework. The heterogeneous multiscale method solves linear cell problems with the macroscopic state frozen at quadrature points and relies on scale separation for the micro–macro coupling. The generalized multiscale finite element method employs state-parametrized linear spectral spaces computed offline. The nonlinear localized orthogonal decomposition methods compute linear correctors from a linearization of the equation in each iteration. The method developed here instead represents the complete eliminated fine scale by a single nonlinear map from coarse states to fine-scale corrections and solves the resulting coarse problem on the associated nonlinear manifold. As in the localized orthogonal decomposition framework, it requires neither periodicity nor scale separation of the coefficient.
Let \(V_h\) be the fine finite element space, let \(V_H\subset V_h\) be a coarse finite element space, let \(\Pi_H:V_h\to V_H\) be a projection, and set \(V_f=\operatorname{ker}\Pi_H\). For a coarse state \(v_H\in V_H\), the ideal fine-scale correction \(v_f(v_H)\in V_f\) is defined by the constrained residual equation, \[A(v_H+v_f(v_H);w_f)=L(w_f) \quad \forall w_f\in V_f \label{eq:introduction-global-correction}\tag{1}\] and the corresponding reconstruction map and manifold are defined by, \[M_h(v_H)=v_H+v_f(v_H),\qquad \mathcal{M}_h=M_h(V_H)\subset V_h \label{eq:introduction-global-manifold}\tag{2}\] The fine finite element solution is represented exactly by this manifold: if \(u_h\) solves the fine problem, then \(u_h=M_h(\Pi_Hu_h)\). The approximation enters only when the global map \(M_h\) is replaced by a localized computable map.
The practical method approximates \(M_h\) by localized patch maps. The use of local patches and a partition-of-unity split is reminiscent of the partition-of-unity finite element method (PUFEM) and generalized finite element method (GFEM), where local approximation spaces are combined into a global trial space [26], [27]. Here, however, the split is used primarily to localize the fine-scale equations for the eliminated correction, rather than to prescribe an enriched global trial space; this viewpoint is also related to multiscale GFEM constructions based on local approximation spaces [28]. For each partition-of-unity function \(\varphi_{H,i}\), we solve a nonlinear fine-scale problem on an oversampled patch \(\omega_i^{\ell}\). The local input is the vector \(P_i v_H\) of coarse degrees of freedom used on that patch. Since the global blend only sees the support of \(\varphi_{H,i}\), only the restricted output \(R_i v_{f,i}^{\ell}(P_i v_H)\) on \(\omega_i^0=\operatorname{int}(\operatorname{supp}\varphi_{H,i})\) is needed. The localized reconstruction is defined by, \[\begin{align} M_h^{\ell}(v_H) &=v_H+(I-\Pi_H)I_h\Bigl( \sum_{i\in\mathcal{I}}\varphi_{H,i}E_{i,0}^{\operatorname{ext}} R_i v_{f,i}^{\ell}(P_i v_H) \Bigr) \label{eq:introduction-localized-reconstruction} \end{align}\tag{3}\] where \(I_h\) is fine nodal interpolation and \(E_{i,0}^{\operatorname{ext}}\) is extension by zero from \(\omega_i^0\) to the global fine grid. The final projection \(I-\Pi_H\) returns the blended correction to the fine-scale space. Thus the patch solutions are local representatives of the complete nonlinear fine-scale correction, not nonlinear analogues of individual basis correctors.
Because \(M_h^{\ell}(V_H)\) is a nonlinear manifold, the coarse residual must be tested with tangent multiscale functions. The patch-solved method seeks \(u_H^{\ell}\in V_H\) such that, \[\begin{align} A\bigl(M_h^{\ell}(u_H^{\ell});D M_h^{\ell}(u_H^{\ell})[\delta v_H]\bigr) &=L\bigl(D M_h^{\ell}(u_H^{\ell})[\delta v_H]\bigr) \quad \forall \delta v_H\in V_H \label{eq:introduction-tangent-coarse-problem} \end{align}\tag{4}\] and then sets \(u_{\mathrm{ms}}^{\ell}=M_h^{\ell}(u_H^{\ell})\). The same tangent formulation applies at the ideal level with \(M_h\) in place of \(M_h^{\ell}\). The tangent test space is not an implementation detail; it is the natural variational test space on the nonlinear approximation manifold.
The localized patch maps are finite-dimensional nonlinear solution operators. They may be evaluated by direct patch solves, or approximated offline and reused online. We suggest an optional network-interpolated method that replaces only the restricted local output maps and their tangent actions by learned maps, giving a reconstruction \(M_{h,\theta}^{\ell}\). This differs from global operator-learning approaches, which approximate maps between full function spaces [29], [30], from model-reduction strategies in which neural networks parametrize a global low-dimensional solution manifold [31], and from classical reduced-basis methodology, which builds a global reduced space for a full parameter family [32], [33]. It is closer to learned localized numerical-homogenization surrogates [34]–[36], but here the learned quantities are the support-restricted patch outputs that enter the partition-of-unity reconstruction. The deterministic localized method remains the primary object; learning is an optional interpolation layer.
The contributions are as follows. First, we formulate an ideal nonlinear fine-scale elimination in which the coarse variable parametrizes the fine finite element solution through the manifold map \(M_h\). Second, we derive the tangent-space coarse equation associated with this nonlinear manifold. Third, we introduce the localized reconstruction \(M_h^{\ell}\), including the support-restricted patch outputs used in the partition-of-unity blend. Fourth, we give patch-solved and network-interpolated computable coarse problems in a common residual form. Fifth, for heterogeneous monotone nonlinear diffusion, we verify monotonicity, local Lipschitz continuity, tangent coercivity, finite-dimensional patch-map regularity, and establish a conditional perturbation estimate in which the chord curvature is absorbed into the stability constant while localization and reduced-residual defects remain on the right-hand side.
We emphasize that the present paper is deliberately methodological. In the error analysis, the localization defect \(\eta_{\mathrm{loc}}(H,\ell)\) and the learning defect enter only as quantities on the right-hand side, and the estimates are conditional on their smallness. A rigorous a priori theory for these defects—in particular the exponential decay of the localization error in the oversampling radius \(\ell\), and the resulting convergence rates in the coarse mesh size \(H\)—is not established here; it is the subject of a separate, more theoretical companion manuscript. In the present paper the decay of the localization defect is instead demonstrated numerically. Accordingly, the contribution of this paper is the tangent-space manifold formulation, the computable localized and network-interpolated reconstructions, and the conditional perturbation mechanism that transfers reconstruction and residual estimates to the final multiscale error, rather than the localization theory itself.
Section 2 introduces the ideal nonlinear reconstruction map, the fine-scale manifold, and the tangent-space coarse equation. Section 3 defines the localized patch construction and the partition-of-unity reconstruction. Section 4 formulates the network interpolation of restricted local patch outputs and tangent actions. Section 5 states the patch-solved and network-interpolated coarse problems in a common algebraic form. Section 6 verifies the structural assumptions for heterogeneous monotone nonlinear diffusion. Section 7 gives the geometric error split, the chord-curvature kickback, and the conditional localization and residual estimates that organize the error analysis. Section 8 presents numerical experiments for heterogeneous monotone nonlinear diffusion, covering the localization and coarse-mesh convergence of the patch-solved method, the network-interpolated solve, and a backward-Euler parabolic test.
This section defines the ideal nonlinear reconstruction map. It is global and therefore not intended for computation, but it identifies the fine-scale object that the localized method approximates. In the error analysis we later specialize to monotone equations; for the formulation of the method we keep only the assumptions needed to define the maps and tangent spaces below.
Let \(\Omega\) be a bounded polygonal or polyhedral domain. We are given two nested, shape-regular and quasi-uniform triangulations of \(\Omega\): a coarse triangulation \(\mathcal{T}_H\) with mesh-width parameter \(H\mathrel{\vcenter{:}}=\max_{T\in\mathcal{T}_H}\operatorname{diam}(T)\), and a fine triangulation \(\mathcal{T}_h\) with mesh-width parameter \(h\mathrel{\vcenter{:}}=\max_{K\in\mathcal{T}_h}\operatorname{diam}(K)\), where \(\mathcal{T}_h\) is a refinement of \(\mathcal{T}_H\) so that every coarse element is a union of fine elements and \(h\le H\). On these triangulations we use continuous piecewise-linear \(P_1\) Lagrange finite elements with homogeneous Dirichlet boundary conditions, and denote the resulting finite element spaces by \(V_H\) and \(V_h\). The nestedness of the triangulations yields the conforming inclusion \(V_H\subset V_h\) used throughout.
Let \(\|\cdot\|_V\) denote the energy norm used in the analysis. The reference problem is to find \(u_h\in V_h\) such that \[A(u_h;v_h)=L(v_h) \quad \forall v_h\in V_h \label{eq:fine-scale-reference-problem}\tag{5}\] where \(A(\cdot;\cdot)\) is nonlinear in the first argument and linear in the second, while \(L\) is linear. We assume that 5 is uniquely solvable and that the constrained fine-scale problems below are well posed. The structural conditions used for the model diffusion problem are stated in Section 6.
Let \(\Pi_H:V_h\to V_H\) be a projection onto the coarse space \(V_H\), so that \[\Pi_H v_H=v_H \quad \forall v_H\in V_H \label{eq:projection-identity-on-coarse-space}\tag{6}\] We assume that \(\Pi_H\) is stable in the energy norm, \[\|\Pi_H v_h\|_V\le C_{\Pi}\|v_h\|_V \quad \forall v_h\in V_h \label{eq:projection-stability}\tag{7}\] with \(C_{\Pi}\) independent of the mesh sizes, and compatible with the local patch restrictions used in Section 3. Concrete admissible choices of \(\Pi_H\) are discussed in the remark below. The local stability constants needed by the perturbation argument are stated in Section 7. We set \[V_f=\operatorname{ker}\Pi_H \label{eq:fine-scale-space}\tag{8}\] as the fine-scale space. Then every \(v_h\in V_h\) has the decomposition \[v_h=\Pi_H v_h +(v_h-\Pi_H v_h) \quad \Pi_H v_h\in V_H \quad v_h-\Pi_H v_h\in V_f \label{eq:coarse-fine-decomposition}\tag{9}\] where the first term is coarse and the second term belongs to \(V_f\). Thus \(V_f\) is the space of fine-scale functions that are invisible to the coarse projection.
The construction requires \(\Pi_H\) to be a projection satisfying the energy stability 7 ; the specific operator is otherwise unconstrained. Possible choices are, e.g., the \(L^2\)-orthogonal projection onto \(V_H\), whose \(H^1\)-stability on shape-regular, quasi-uniform meshes is classical [37], or the Scott–Zhang quasi-interpolation operator [38], which is a projection onto \(V_H\) that is both energy-stable and local. Similarly, the localized reconstruction of Section 3 uses only a partition of unity subordinate to the coarse mesh; natural choices are the \(P_1\) coarse vertex hat functions \(\varphi_{H,z}=\Lambda_z\), whose active regions are the nodal stars, or piecewise-constant element indicators. The numerical experiments in Section 8 use the \(L^2\)-projection together with vertex-patch hats.
For each \(v_H\in V_H\), define \(v_f(v_H)\in V_f\) as the solution of \[A(v_H+v_f(v_H);w_f)=L(w_f) \quad \forall w_f\in V_f \label{eq:global-fine-scale-correction}\tag{10}\] which enforces the residual equation on \(V_f\). The corresponding reconstruction map \(M_h:V_H\to V_h\) is \[M_h(v_H)=v_H+v_f(v_H) \label{eq:global-reconstruction-map}\tag{11}\] and its image is \[\mathcal{M}_h=M_h(V_H)\subset V_h \label{eq:global-nonlinear-manifold}\tag{12}\] which is the nonlinear fine-scale manifold. By construction, the residual of \(M_h(v_H)\) vanishes on \(V_f\) for every coarse state \(v_H\).
For every \(v_H\in V_H\), the global reconstruction satisfies \[\Pi_HM_h(v_H)=v_H \label{eq:global-coarse-parametrization}\tag{13}\] If \(M_h\) is differentiable at \(v_H\), then \[\Pi_HD M_h(v_H)[\delta v_H]=\delta v_H \quad \forall \delta v_H\in V_H \label{eq:global-tangent-coarse-parametrization}\tag{14}\] Consequently, the tangent space is a complement of \(V_f\), \[V_h=V_f\oplus T_{M_h(v_H)}\mathcal{M}_h \label{eq:global-tangent-direct-sum}\tag{15}\] where \(T_{M_h(v_H)}\mathcal{M}_h\) is defined in 22 .
Proof. Since \(v_f(v_H)\in V_f=\operatorname{ker}\Pi_H\), \[\Pi_HM_h(v_H)=\Pi_Hv_H+\Pi_Hv_f(v_H)=v_H \label{eq:global-coarse-parametrization-proof}\tag{16}\] Differentiating 13 gives 14 . If \(z_h\in V_f\cap T_{M_h(v_H)}\mathcal{M}_h\), then \(z_h=D M_h(v_H)[\delta v_H]\) for some \(\delta v_H\in V_H\), and 14 gives \[\delta v_H=\Pi_H z_h=0 \label{eq:global-direct-sum-injectivity-proof}\tag{17}\] Thus the intersection is trivial. Since \(V_h=V_H\oplus V_f\) by 9 and \(D M_h(v_H)\) is injective by 14 , the dimensions add up and 15 follows. ◻
Let \(u_h\) solve 5 and set \(u_H=\Pi_H u_h\). Then \(u_h-u_H\in V_f\), and testing 5 with \(w_f\in V_f\) gives \[A(u_H+(u_h-u_H);w_f)=L(w_f) \quad \forall w_f\in V_f \label{eq:exact-representation-fine-scale-equation}\tag{18}\] which is the constrained correction problem with coarse state \(u_H\). By uniqueness of that problem, \[u_h-u_H=v_f(u_H) \quad u_h=M_h(u_H) \label{eq:exact-global-manifold-representation}\tag{19}\] so the exact fine-scale solution belongs to \(\mathcal{M}_h\). The localized method in Section 3 approximates this map rather than the solution itself.
Assume that \(v_f\) is differentiable at \(v_H\). For a coarse direction \(\delta v_H\in V_H\), the tangent reconstruction is \[D M_h(v_H)[\delta v_H]=\delta v_H+D v_f(v_H)[\delta v_H] \label{eq:global-tangent-reconstruction}\tag{20}\] where \(D v_f(v_H)[\delta v_H]\in V_f\) solves the linearized fine-scale problem \[A'(v_H+v_f(v_H))[\delta v_H+D v_f(v_H)[\delta v_H],w_f]=0 \quad \forall w_f\in V_f \label{eq:global-tangent-correction}\tag{21}\] and \(A'(u)[z,w]\) denotes the derivative of \(A(\cdot;w)\) at \(u\) in the direction \(z\). The tangent space to \(\mathcal{M}_h\) at \(M_h(v_H)\) is \[T_{M_h(v_H)}\mathcal{M}_h=\{D M_h(v_H)[\delta v_H]:\delta v_H\in V_H\} \label{eq:global-tangent-space}\tag{22}\] By Lemma [lem:global-coarse-parametrization], this tangent space contains exactly one lift of each coarse direction and is complementary to \(V_f\).
The map \(M_h(v_H)=v_H+v_f(v_H)\) identifies the ideal manifold as a nonlinear graph over the coarse space. Equivalently, it is a section of the affine fibration induced by \(\Pi_H\), selecting one representative in each fine-scale fiber \(v_H+V_f\). Figure 1 illustrates this geometry and the associated tangent lift.
The ideal coarse problem is to find \(u_H\in V_H\) such that \[A(M_h(u_H);D M_h(u_H)[\delta v_H])=L(D M_h(u_H)[\delta v_H]) \quad \forall \delta v_H\in V_H \label{eq:global-manifold-problem}\tag{23}\] or equivalently, with \(u_h=M_h(u_H)\), \[A(u_h;z_h)=L(z_h) \quad \forall z_h\in T_{u_h}\mathcal{M}_h \label{eq:global-tangent-residual-equation}\tag{24}\]
Assume that the constrained correction problem 10 is uniquely solvable and that \(M_h\) is differentiable at the relevant coarse states. If \(\widehat u_H\in V_H\) solves 23 , then \(\widehat u_h=M_h(\widehat u_H)\) solves the fine-scale problem 5 . Conversely, if \(u_h\) solves 5 and \(u_H=\Pi_Hu_h\), then \(u_h=M_h(u_H)\) and \(u_H\) solves 23 .
Proof. Let \(\widehat u_H\) solve 23 and set \(\widehat u_h=M_h(\widehat u_H)\). The definition of \(M_h\) gives \[A(\widehat u_h;w_f)=L(w_f) \quad \forall w_f\in V_f \label{eq:ideal-equivalence-fine-residual}\tag{25}\] and 23 gives \[A(\widehat u_h;z_h)=L(z_h) \quad \forall z_h\in T_{\widehat u_h}\mathcal{M}_h \label{eq:ideal-equivalence-tangent-residual}\tag{26}\] By 15 , every \(v_h\in V_h\) can be written uniquely as \(v_h=w_f+z_h\) with \(w_f\in V_f\) and \(z_h\in T_{\widehat u_h}\mathcal{M}_h\). Since the residual is linear in the test function, \[A(\widehat u_h;v_h)=L(v_h) \quad \forall v_h\in V_h \label{eq:ideal-equivalence-full-residual}\tag{27}\] Thus \(\widehat u_h\) solves 5 . The converse follows from 19 and by testing the fine-scale equation with the tangent functions \(D M_h(u_H)[\delta v_H]\). ◻
The tangent formulation is needed because \(\mathcal{M}_h\) is not a linear space. If \(M_h\) is differentiable along the segment from \(w_H\) to \(v_H\), then \[M_h(v_H)-M_h(w_H)=\int_0^1 D M_h(w_H+s(v_H-w_H))[v_H-w_H] \,ds \label{eq:mean-value-identity}\tag{28}\] which replaces the linear-space difference argument used in standard Galerkin orthogonality.
This section replaces the global correction map \(M_h\) from Section 2 by a computable localized map. The construction solves nonlinear residual equations on oversampled patches, blends only the active parts of the local corrections by a coarse partition of unity, and projects the blended correction back to \(V_f\). At this stage the only required assumptions are well-posedness and differentiability of the local patch maps at the states where they are evaluated.
Let \(\{\varphi_{H,i}\}_{i\in\mathcal{I}}\) be a partition of unity subordinate to the coarse mesh \(\mathcal{T}_H\), with each support \(\operatorname{supp}\varphi_{H,i}\) a union of coarse elements (see the remark in Section 2 for admissible choices). Patches are built from coarse-element neighborhoods. For a subdomain \(\omega\subset\Omega\) that is a union of coarse elements, one coarse-element layer is added by \[\label{eq:element-neighborhood} \begin{align} \mathsf{N}(\omega)&=\operatorname{int}\Bigl(\bigcup\{\overline{T}: T\in\mathcal{T}_H,\;\overline{T}\cap\overline{\omega}\ne\emptyset\}\Bigr),\\ \mathsf{N}^{0}(\omega)&=\omega,\\ \mathsf{N}^{\ell}(\omega)&=\mathsf{N}\bigl(\mathsf{N}^{\ell-1}(\omega)\bigr)\;\;(\ell\ge1) \end{align}\tag{29}\] The active region of patch \(i\) is \[\omega_i^{0}=\operatorname{int}\bigl(\operatorname{supp}\varphi_{H,i}\bigr) \label{eq:partition-support-region}\tag{30}\] and the oversampled patch of order \(\ell\ge0\) is its \(\ell\)-layer coarse-element neighborhood, \[\omega_i^{\ell}=\mathsf{N}^{\ell}(\omega_i^{0}) \label{eq:oversampled-patch}\tag{31}\] Each \(\omega_i^{\ell}\) is a union of coarse elements with \(\omega_i^{0}\subset\omega_i^{\ell}\subset\omega_i^{\ell+1}\), and \(\ell\) counts the coarse-element layers added beyond the active region. We fix a basis \(\{\Phi_j\}_{j=1}^{N_H}\) of \(V_H\) and define the local coarse index set by \[\mathcal{I}_i^{\ell}=\{j\in\{1,\ldots,N_H\}:\operatorname{supp}\Phi_j\cap\omega_i^{\ell}\ne\emptyset\} \label{eq:local-coarse-index-set}\tag{32}\] If \(v_H=\sum_{j=1}^{N_H}c_j\Phi_j\), then the local restriction operator is \[P_i v_H=(c_j)_{j\in\mathcal{I}_i^{\ell}}\in\mathbb{R}^{n_{H,i}} \quad n_{H,i}=|\mathcal{I}_i^{\ell}| \label{eq:local-coarse-restriction-operator}\tag{33}\] Thus \(P_i v_H\) determines \(v_H\) on \(\omega_i^{\ell}\). The local fine-scale space is \[V_{f,i}^{\ell}=\bigl\{v_f\in V_f: \operatorname{supp} v_f\subset \overline{\omega_i^{\ell}},\;v_f=0\;\text{on}\;\partial\omega_i^{\ell}\bigr\} \label{eq:local-fine-scale-space}\tag{34}\] so that the local corrections satisfy homogeneous Dirichlet conditions on the artificial patch boundary. We write \(E_i^{\operatorname{ext}}\) for extension by zero from \(\omega_i^{\ell}\) to the global fine grid. The partition of unity is active only on the star \(\omega_i^{0}\subset\omega_i^{\ell}\) defined in 30 . Let \(R_i\) denote restriction to \(\omega_i^{0}\), and let \(W_{f,i}^{0}\) be the corresponding restricted output space, \[R_i\colon V_{f,i}^{\ell}\to W_{f,i}^{0} \quad W_{f,i}^{0}=R_i V_{f,i}^{\ell} \label{eq:restricted-output-space}\tag{35}\] We write \(E_{i,0}^{\operatorname{ext}}\) for extension by zero from \(\omega_i^{0}\) to the global fine grid.
For \(v_H\in V_H\), the local correction \(v_{f,i}^{\ell}(P_i v_H)\in V_{f,i}^{\ell}\) is defined by \[\begin{align} A_{\omega_i^{\ell}}\bigl(v_H+v_{f,i}^{\ell}(P_i v_H);w_{f,i}\bigr) &=L_{\omega_i^{\ell}}(w_{f,i}) \quad \forall w_{f,i}\in V_{f,i}^{\ell} \label{eq:patch-correction-problem} \end{align}\tag{36}\] where \(A_{\omega_i^{\ell}}\) and \(L_{\omega_i^{\ell}}\) are the restrictions of the global forms to the patch. Although \(v_H\) appears in the residual, only its patch restriction is used, and this restriction is determined by \(P_i v_H\) through 33 . This is why the local map can be viewed as a finite-dimensional map of \(n_{H,i}\) coarse variables.
Only the restriction of the patch correction to \(\omega_i^{0}\) enters the global blend. Indeed, \[\begin{align} \varphi_{H,i}E_i^{\operatorname{ext}}v_{f,i}^{\ell}(P_i v_H) &=\varphi_{H,i}E_{i,0}^{\operatorname{ext}}R_i v_{f,i}^{\ell}(P_i v_H) \quad \forall v_H\in V_H \label{eq:partition-support-restriction} \end{align}\tag{37}\] The localized reconstruction map is therefore \[\begin{align} M_h^{\ell}(v_H) &=v_H+(I-\Pi_H)I_h\Bigl(\sum_{i\in\mathcal{I}}\varphi_{H,i}E_{i,0}^{\operatorname{ext}}R_i v_{f,i}^{\ell}(P_i v_H)\Bigr) \label{eq:localized-reconstruction-map} \end{align}\tag{38}\] where \(I_h\) denotes a stable fine-grid interpolation or projection. The product \(\varphi_{H,i}E_{i,0}^{\operatorname{ext}}R_i v_{f,i}^{\ell}\) is generally not an element of \(V_h\), and \(I_h\) places the partition-of-unity blend in the fine space. The projection \(I-\Pi_H\) is essential because the blended correction need not belong to \(V_f\). The localized nonlinear trial manifold is \[\mathcal{M}_h^{\ell}=M_h^{\ell}(V_H)\subset V_h \label{eq:localized-trial-manifold}\tag{39}\] which is the computable counterpart of the ideal manifold \(\mathcal{M}_h\).
For a coarse direction \(\delta v_H\in V_H\), differentiation of 38 gives \[\begin{align} D M_h^{\ell}(v_H)[\delta v_H] &=\delta v_H+(I-\Pi_H)I_h\Bigl(\sum_{i\in\mathcal{I}}\varphi_{H,i}E_{i,0}^{\operatorname{ext}}R_i D v_{f,i}^{\ell}(P_i v_H)[P_i\delta v_H]\Bigr) \label{eq:localized-tangent-reconstruction} \end{align}\tag{40}\] The local tangent correction \(D v_{f,i}^{\ell}(P_i v_H)[P_i\delta v_H]\in V_{f,i}^{\ell}\) is obtained by differentiating 36 . It solves \[\begin{align} A_{\omega_i^{\ell}}'\bigl(v_H+v_{f,i}^{\ell}(P_i v_H)\bigr)\bigl[\delta v_H+D v_{f,i}^{\ell}(P_i v_H)[P_i\delta v_H],w_{f,i}\bigr] &=0 \quad \forall w_{f,i}\in V_{f,i}^{\ell} \label{eq:localized-tangent-patch-problem} \end{align}\tag{41}\] where the derivative is taken with respect to the first argument of \(A_{\omega_i^{\ell}}\). As for the state correction, only the restriction of this tangent correction to \(\omega_i^{0}\) is used in the global tangent reconstruction.
For every \(v_H\in V_H\), the localized reconstruction satisfies \[\Pi_HM_h^{\ell}(v_H)=v_H \label{eq:localized-coarse-parametrization}\tag{42}\] If \(M_h^{\ell}\) is differentiable at \(v_H\), then \[\Pi_HD M_h^{\ell}(v_H)[\delta v_H]=\delta v_H \quad \forall \delta v_H\in V_H \label{eq:localized-tangent-coarse-parametrization}\tag{43}\] Consequently, \[V_h=V_f\oplus T_{M_h^{\ell}(v_H)}\mathcal{M}_h^{\ell} \label{eq:localized-tangent-direct-sum}\tag{44}\] whenever the tangent space is defined.
Proof. The added correction in 38 lies in \(V_f\) because of the factor \(I-\Pi_H\). Therefore, \[\Pi_HM_h^{\ell}(v_H)=\Pi_Hv_H=v_H \label{eq:localized-coarse-parametrization-proof}\tag{45}\] Differentiating 42 gives 43 . The direct-sum statement follows as in Lemma [lem:global-coarse-parametrization]. ◻
The deterministic localized reconstruction can be evaluated patchwise. For a given coarse state \(v_H\), one solves 36 independently on all patches, restricts the local outputs to \(\omega_i^0\), blends the restricted outputs with \(\varphi_{H,i}\), applies \(I_h\), and finally projects with \(I-\Pi_H\). For tangent actions, one reuses the local state solutions, solves the linearized patch problems 41 for the required local coarse directions, and assembles the result by 40 .
The localized map \(M_h^{\ell}\) and its tangent action \(D M_h^{\ell}\) define the deterministic computable method. The corresponding tangent-space coarse equation is stated in Section 5; the only change relative to the ideal equation in Section 2 is the replacement of \(M_h\) by \(M_h^{\ell}\).
The localized patch-solved method from Section 3 is already computable. This section describes an optional acceleration in which the restricted local patch-output operators are replaced by learned interpolants. The restriction is important: since the global reconstruction only uses the correction on the active support \(\omega_i^0\), no network output is required on the oversampling layer \(\omega_i^{\ell}\setminus\omega_i^0\). Because the coarse equation uses tangent test functions, the relevant approximation target is the patch map together with its derivative.
For patch \(i\), let \(R_i\) and \(W_{f,i}^{0}\) be defined as in Section 3. After choosing a local basis \(\{\psi_{i,r}^{0}\}_{r=1}^{n_{0,i}}\) of \(W_{f,i}^{0}\), we identify the restricted output space with coefficient vectors by \[w_i^0=\sum_{r=1}^{n_{0,i}}\gamma_{i,r}\psi_{i,r}^{0} \quad \longleftrightarrow \quad \gamma_i=(\gamma_{i,1},\ldots,\gamma_{i,n_{0,i}})\in\mathbb{R}^{n_{0,i}} \label{eq:restricted-output-coordinates}\tag{46}\] The operator needed by the global reconstruction is \[\begin{align} \mathcal{G}_{f,i}^{\ell}&\colon \mathbb{R}^{n_{H,i}}\to W_{f,i}^{0} \tag{47}\\ \mathcal{G}_{f,i}^{\ell}(P_i v_H)&=R_i v_{f,i}^{\ell}(P_i v_H) \quad \forall v_H\in V_H \tag{48} \end{align}\] Thus the full patch problem on \(\omega_i^{\ell}\) may be used to generate the local output, but values outside \(\omega_i^{0}\) are not part of the interpolated data.
For each patch, we replace \(\mathcal{G}_{f,i}^{\ell}\) by a learned interpolant with parameter vector \(\theta_i\), \[\mathcal{N}_{f,i}^{\ell,\theta_i}\colon \mathbb{R}^{n_{H,i}}\to \mathbb{R}^{n_{0,i}} \label{eq:network-patch-map}\tag{49}\] where the output is interpreted as nodal values of a fine finite element function on \(\omega_i^{0}\). We define the learned restricted correction by \[g_{f,i}^{\ell,\theta_i}(P_i v_H)=\mathcal{N}_{f,i}^{\ell,\theta_i}(P_i v_H) \label{eq:learned-patch-correction}\tag{50}\] With \(\theta=(\theta_i)_{i\in\mathcal{I}}\), the learned global reconstruction is \[\begin{align} M_{h,\theta}^{\ell}(v_H) &=v_H+(I-\Pi_H)I_h \Bigl(\sum_{i\in\mathcal{I}}\varphi_{H,i}E_{i,0}^{\operatorname{ext}} g_{f,i}^{\ell,\theta_i}(P_i v_H)\Bigr) \label{eq:learned-reconstruction-map} \end{align}\tag{51}\] It is a network-interpolated approximation of \(M_h^{\ell}\), while \(M_h^{\ell}\) is the localized approximation of the ideal map \(M_h\). As in Lemma [lem:localized-coarse-parametrization], the projection \(I-\Pi_H\) gives \[\Pi_HM_{h,\theta}^{\ell}(v_H)=v_H \quad \forall v_H\in V_H \label{eq:learned-coarse-parametrization}\tag{52}\] In particular, the network output is not required to lie in \(W_{f,i}^{0}=R_iV_{f,i}^{\ell}\): arbitrary nodal values on \(\omega_i^{0}\) are admissible, since the projection \(I-\Pi_H\) restores membership in \(V_f\) after the partition-of-unity blend.
The tangent map is obtained by differentiating the learned restricted patch maps. For \(\delta v_H\in V_H\), \[\begin{align} D M_{h,\theta}^{\ell}(v_H)[\delta v_H] &=\delta v_H+(I-\Pi_H)I_h \Bigl(\sum_{i\in\mathcal{I}}\varphi_{H,i}E_{i,0}^{\operatorname{ext}} D\mathcal{N}_{f,i}^{\ell,\theta_i}(P_i v_H)[P_i\delta v_H]\Bigr) \label{eq:learned-tangent-reconstruction} \end{align}\tag{53}\] This gives tangent test functions without solving linearized patch problems online. The derivatives of the network maps can be evaluated by automatic differentiation. The derivative is not merely an implementation detail: tangent-space consistency requires accuracy of both \(\mathcal{N}_{f,i}^{\ell,\theta_i}\) and \(D\mathcal{N}_{f,i}^{\ell,\theta_i}\) on the admissible local parameter set.
A minimal local data set consists of local coarse states and the corresponding restricted patch corrections, \[\mathcal{D}_i^{\ell}=\bigl\{(P_i v_H^m,R_i v_{f,i}^{\ell}(P_i v_H^m))\bigr\}_{m=1}^{N_{\operatorname{train}}} \label{eq:patch-training-data}\tag{54}\] When tangent data are available, let \(\rho_{i,q}\in\mathbb{R}^{n_{H,i}}\), \(q=1,\ldots,Q_i\), denote local coarse directions used to sample derivative information. A derivative-enhanced local loss is \[\begin{align} \mathcal{L}_i(\theta_i) &=\sum_{m=1}^{N_{\operatorname{train}}}\bigl\|\mathcal{N}_{f,i}^{\ell,\theta_i}(P_i v_H^m)-\mathcal{G}_{f,i}^{\ell}(P_i v_H^m)\bigr\|_{H^1(\omega_i^{0})}^2 \tag{55}\\ &\quad+\lambda\sum_{m=1}^{N_{\operatorname{train}}}\sum_{q=1}^{Q_i}\bigl\|D\mathcal{N}_{f,i}^{\ell,\theta_i}(P_i v_H^m)[\rho_{i,q}]-D\mathcal{G}_{f,i}^{\ell}(P_i v_H^m)[\rho_{i,q}]\bigr\|_{H^1(\omega_i^{0})}^2 \tag{56} \end{align}\] with \(\lambda\ge0\). Choosing \(\lambda=0\) gives the state-only loss. Section 7 does not depend on a particular training algorithm. It uses either the realized reconstruction error of the learned map or the reduced residual obtained after solving the learned coarse problem; tangent approximation errors enter when one compares learned and deterministic residuals or uses a tangent surrogate that is not the derivative of the same learned reconstruction.
This section gives the finite-dimensional problems solved online. Sections 3 and 4 define the reconstruction maps; here we state the tangent-space coarse equation, the residual used in computation, and the work required for one residual evaluation.
We use \(\widetilde{M}_h^{\ell}\) for either the patch-solved or the network-interpolated reconstruction, \[\widetilde{M}_h^{\ell}=M_h^{\ell} \quad \text{or} \quad \widetilde{M}_h^{\ell}=M_{h,\theta}^{\ell} \label{eq:generic-reconstruction-choices}\tag{57}\] The corresponding restricted local output is denoted by \(\widetilde{g}_i^{\ell}\). For \(\mu\in\mathbb{R}^{n_{H,i}}\), \[\begin{align} \widetilde{g}_i^{\ell}(\mu) &=\begin{cases} R_i v_{f,i}^{\ell}(\mu), & \text{for the patch-solved method}\\ \mathcal{N}_{f,i}^{\ell,\theta_i}(\mu), & \text{for the network-interpolated method} \end{cases} \label{eq:generic-restricted-output} \end{align}\tag{58}\] With this notation both reconstructions have the common form \[\begin{align} \widetilde{M}_h^{\ell}(v_H) &=v_H+(I-\Pi_H)I_h\Bigl(\sum_{i\in\mathcal{I}}\varphi_{H,i}E_{i,0}^{\operatorname{ext}}\widetilde{g}_i^{\ell}(P_i v_H)\Bigr) \label{eq:generic-reconstruction-map} \end{align}\tag{59}\] and the tangent action is \[\begin{align} D\widetilde{M}_h^{\ell}(v_H)[\delta v_H] &=\delta v_H+(I-\Pi_H)I_h\Bigl(\sum_{i\in\mathcal{I}}\varphi_{H,i}E_{i,0}^{\operatorname{ext}}D\widetilde{g}_i^{\ell}(P_i v_H)[P_i\delta v_H]\Bigr) \label{eq:generic-tangent-map} \end{align}\tag{60}\] For the patch-solved method, \(D\widetilde{g}_i^{\ell}\) is obtained from the linearized patch problem 41 . For the network-interpolated method, \(D\widetilde{g}_i^{\ell}\) is the derivative of the learned local map. In both cases, \[\Pi_H\widetilde{M}_h^{\ell}(v_H)=v_H \quad \Pi_HD\widetilde{M}_h^{\ell}(v_H)[\delta v_H]=\delta v_H \label{eq:generic-coarse-parametrization}\tag{61}\] whenever the derivative is defined.
The computable coarse problem is the tangent-space Galerkin equation on the chosen localized manifold. It reads: find \(\widetilde{u}_H^{\ell}\in V_H\) such that \[\begin{align} A\bigl(\widetilde{M}_h^{\ell}(\widetilde{u}_H^{\ell});D\widetilde{M}_h^{\ell}(\widetilde{u}_H^{\ell})[\delta v_H]\bigr) &=L\bigl(D\widetilde{M}_h^{\ell}(\widetilde{u}_H^{\ell})[\delta v_H]\bigr) \quad \forall \delta v_H\in V_H \label{eq:computable-coarse-problem} \end{align}\tag{62}\] The final fine-scale approximation is \[\widetilde{u}_{\mathrm{ms}}^{\ell}=\widetilde{M}_h^{\ell}(\widetilde{u}_H^{\ell}) \label{eq:generic-fine-scale-approximation}\tag{63}\] Thus \(\widetilde{M}_h^{\ell}=M_h^{\ell}\) gives the deterministic patch-solved method, while \(\widetilde{M}_h^{\ell}=M_{h,\theta}^{\ell}\) gives the optional network-interpolated method. In the analysis and numerical experiments the nonlinear algebraic problem is considered on an admissible set where the reconstruction and tangent maps are defined.
If the fine-scale problem is the Euler–Lagrange equation of an energy \(\mathcal{E}_h:V_h\to\mathbb{R}\), with \[D\mathcal{E}_h(u_h)[v_h]=A(u_h;v_h)-L(v_h) \quad \forall u_h,v_h\in V_h \label{eq:energy-derivative-assumption}\tag{64}\] then 62 is the stationarity condition for the reduced energy \[\mathcal{J}_H^{\ell}(v_H)=\mathcal{E}_h(\widetilde{M}_h^{\ell}(v_H)) \quad v_H\in V_H \label{eq:reduced-energy-functional}\tag{65}\] Indeed, the chain rule gives \[\begin{align} D\mathcal{J}_H^{\ell}(v_H)[\delta v_H] &=D\mathcal{E}_h(\widetilde{M}_h^{\ell}(v_H))[D\widetilde{M}_h^{\ell}(v_H)[\delta v_H]] \tag{66}\\ &=A\bigl(\widetilde{M}_h^{\ell}(v_H);D\widetilde{M}_h^{\ell}(v_H)[\delta v_H]\bigr)-L\bigl(D\widetilde{M}_h^{\ell}(v_H)[\delta v_H]\bigr) \tag{67} \end{align}\] This interpretation is used for the model diffusion problem in Section 6 and is useful for designing nonlinear solvers.
Let \(\{\Phi_j\}_{j=1}^{N_H}\) be the basis of \(V_H\) used in 33 . For a coefficient vector \(c\in\mathbb{R}^{N_H}\), define \[v_H(c)=\sum_{k=1}^{N_H}c_k\Phi_k \label{eq:coarse-vector-to-function}\tag{68}\] and the local parameter vector \[\mu_i(c)=P_i v_H(c) \quad \forall i\in\mathcal{I} \label{eq:local-parameter-vector}\tag{69}\] A residual evaluation consists of four steps. First, compute the restricted local outputs \(\widetilde{g}_i^{\ell}(\mu_i(c))\) independently for all patches. Second, assemble the global state \(\widetilde{M}_h^{\ell}(v_H(c))\) by 59 . Third, for each coarse basis direction \(\Phi_j\), evaluate the local tangent outputs \(D\widetilde{g}_i^{\ell}(\mu_i(c))[P_i\Phi_j]\). The block locality is \[P_i\Phi_j=0 \quad \text{and hence} \quad D\widetilde{g}_i^{\ell}(\mu_i(c))[P_i\Phi_j]=0 \quad \text{if } j\notin\mathcal{I}_i^{\ell} \label{eq:local-tangent-block-sparsity}\tag{70}\] Fourth, assemble the tangent functions \(D\widetilde{M}_h^{\ell}(v_H(c))[\Phi_j]\) by 60 and test the residual. For the patch-solved method, the linearized patch operator can be assembled once per patch and reused for all local tangent right-hand sides associated with \(j\in\mathcal{I}_i^{\ell}\).
The resulting coarse residual has components \[\begin{align} \mathcal{R}_j(c) &=A\bigl(\widetilde{M}_h^{\ell}(v_H(c));D\widetilde{M}_h^{\ell}(v_H(c))[\Phi_j]\bigr) -L\bigl(D\widetilde{M}_h^{\ell}(v_H(c))[\Phi_j]\bigr), \quad j=1,\ldots,N_H \label{eq:coarse-residual-vector} \end{align}\tag{71}\] and the online solve is the nonlinear algebraic system \[\mathcal{R}(c)=0 \label{eq:coarse-residual-system}\tag{72}\] This is the form used in implementation and in the error discussion.
Algorithm 2 summarizes the corresponding patch-solved online procedure. The matrix \(J_H\) denotes the Jacobian, a Jacobian approximation, or a matrix-free linearization used by the chosen nonlinear coarse solver. In the experiments of Section 8 and in the extended algorithm of Appendix 9 it is the Galerkin Jacobian \(J_H=(D\widetilde{M}_h^\ell)^{\top}A'(\widetilde{u}_{\mathrm{ms}}^\ell)\,D\widetilde{M}_h^\ell\), obtained by dropping the curvature term \(\langle\mathcal{R},D^2\widetilde{M}_h^\ell\rangle\) from the exact outer derivative; this inexact choice is derived and motivated in Section 8.
The formulation of \(\mathcal{R}\) requires only first derivatives of the local patch maps. An exact Jacobian for Newton’s method would differentiate 71 once more and therefore involves second derivatives of the reconstruction. In practice one may instead use damped Newton with finite-difference Jacobian actions, quasi-Newton updates, or matrix-free Newton–Krylov iterations based on residual evaluations. In the patch-solved method, the online cost is dominated by independent nonlinear and linearized patch solves. In the network-interpolated method, those online solves are replaced by evaluations of \(\mathcal{N}_{f,i}^{\ell,\theta_i}\) and its derivative; the patch solves are moved to the offline data-generation stage.
This section fixes the model problem used to verify the structural assumptions and to guide the numerical experiments. The construction in Sections 2–5 applies to more general nonlinear elliptic problems; here we specialize to a heterogeneous monotone diffusion law.
We use the domain \(\Omega\) and the nested \(P_1\) Lagrange spaces \(V_H\subset V_h\) introduced in Section 2; nonhomogeneous Dirichlet data can be treated by a standard lifting. These spaces are \(H^1\)-conforming, with \(V_h\subset H_0^1(\Omega)\) and energy norm \[\|v_h\|_V=\|\nabla v_h\|_{\Omega} \label{eq:model-discrete-spaces}\tag{73}\] All structural statements below are used at the discrete level. Since \(V_h\) and the patch spaces are finite dimensional, differentiability and local Lipschitz continuity are understood on admissible subsets of these discrete spaces. Let \(f\in L^2(\Omega)\), let \(\alpha\ge 0\), and assume that \(a_{\varepsilon}\in L^{\infty}(\Omega)\) satisfies \[0<a_0\le a_{\varepsilon}(x)\le a_1<\infty \quad \operatorname{a.e.}\;x\in\Omega \label{eq:model-coefficient-bounds}\tag{74}\]
The discrete energy is \[\begin{align} \mathcal{E}_h(v_h) &=\int_{\Omega} a_{\varepsilon}(x) \biggl( \frac{1}{2}|\nabla v_h|^2+ \frac{\alpha}{4}|\nabla v_h|^4 \biggr)\,dx -\int_{\Omega}fv_h\,dx \label{eq:model-energy} \end{align}\tag{75}\] The nonlinear form and the load functional are \[\begin{align} A_h(u_h;v_h) &=\int_{\Omega} a_{\varepsilon}(x) \bigl(1+\alpha|\nabla u_h|^2\bigr) \nabla u_h\cdot\nabla v_h\,dx \tag{76}\\ F_h(v_h) &=\int_{\Omega}fv_h\,dx \tag{77} \end{align}\] The fine-scale reference problem is to find \(u_h\in V_h\) such that \[A_h(u_h;v_h)=F_h(v_h) \quad \forall v_h\in V_h \label{eq:model-fine-scale-problem}\tag{78}\] This problem is the Euler–Lagrange equation for \(\mathcal{E}_h\), that is, \[D\mathcal{E}_h(u_h)[v_h]=A_h(u_h;v_h)-F_h(v_h) \quad \forall u_h,v_h\in V_h \label{eq:model-energy-derivative}\tag{79}\]
The following proposition records the model-specific properties needed by the formulation and the abstract error estimate.
Assume 74 . Then \(A_h\) is strongly monotone, locally Lipschitz continuous, and differentiable in its first argument. More precisely, for all \(v_h,w_h\in V_h\), \[A_h(v_h;v_h-w_h)-A_h(w_h;v_h-w_h) \ge a_0\|v_h-w_h\|_V^2 \label{eq:model-strong-monotonicity}\tag{80}\] For every \(K>0\), define the admissible set \[\mathcal{B}_K=\bigl\{v_h\in V_h:\|\nabla v_h\|_{L^{\infty}(\Omega)}\le K\bigr\} \label{eq:model-admissible-set}\tag{81}\] For all \(v_h,w_h\in\mathcal{B}_K\), \[\|A_h(v_h;\cdot)-A_h(w_h;\cdot)\|_{V_h'} \le a_1\bigl(1+C\alpha K^2\bigr)\|v_h-w_h\|_V \label{eq:model-local-lipschitz-continuity}\tag{82}\] where \(C\) depends only on the dimension and shape regularity of the finite element mesh. The derivative is given by 83 –84 , \[\begin{align} A_h'(u_h)[z_h,r_h] &=\int_{\Omega} a_{\varepsilon}(x) \bigl(1+\alpha|\nabla u_h|^2\bigr) \nabla z_h\cdot\nabla r_h\,dx \tag{83}\\ &\quad+2\alpha\int_{\Omega} a_{\varepsilon}(x) (\nabla u_h\cdot\nabla z_h)(\nabla u_h\cdot\nabla r_h)\,dx \tag{84} \end{align}\] and it satisfies, for all \(u_h,z_h\in V_h\), \[A_h'(u_h)[z_h,z_h] \ge a_0\|z_h\|_V^2 \label{eq:model-tangent-coercivity}\tag{85}\]
Proof. Set \[g(\boldsymbol{\xi})=\bigl(1+\alpha|\boldsymbol{\xi}|^2\bigr)\boldsymbol{\xi} \label{eq:model-flux-map}\tag{86}\] Then, for all \(\boldsymbol{\xi},\boldsymbol{z}\in\mathbb{R}^d\), \[Dg(\boldsymbol{\xi})\boldsymbol{z}\cdot\boldsymbol{z} =\bigl(1+\alpha|\boldsymbol{\xi}|^2\bigr)|\boldsymbol{z}|^2 +2\alpha(\boldsymbol{\xi}\cdot\boldsymbol{z})^2 \ge |\boldsymbol{z}|^2 \label{eq:model-flux-jacobian-coercivity}\tag{87}\] Hence, for all \(\boldsymbol{\xi},\boldsymbol{\eta}\in\mathbb{R}^d\), \[\begin{align} \bigl(g(\boldsymbol{\xi})-g(\boldsymbol{\eta})\bigr)\cdot(\boldsymbol{\xi}-\boldsymbol{\eta}) &=\int_0^1 Dg\bigl(\boldsymbol{\eta}+s(\boldsymbol{\xi}-\boldsymbol{\eta})\bigr)(\boldsymbol{\xi}-\boldsymbol{\eta})\cdot(\boldsymbol{\xi}-\boldsymbol{\eta})\,ds \tag{88}\\ &\ge |\boldsymbol{\xi}-\boldsymbol{\eta}|^2 \tag{89} \end{align}\] After multiplication by \(a_{\varepsilon}\) and integration over \(\Omega\), this gives 80 . The pointwise estimate \[|g(\boldsymbol{\xi})-g(\boldsymbol{\eta})| \le \bigl(1+C\alpha(|\boldsymbol{\xi}|^2+|\boldsymbol{\eta}|^2)\bigr)|\boldsymbol{\xi}-\boldsymbol{\eta}| \quad \forall \boldsymbol{\xi},\boldsymbol{\eta}\in\mathbb{R}^d \label{eq:model-pointwise-lipschitz}\tag{90}\] gives 82 on \(\mathcal{B}_K\). Differentiating 76 with respect to the first argument gives 83 –84 , and 85 follows from 87 . ◻
For \(v_H\in V_H\) and \(i\in\mathcal{I}\), the localized correction \(v_{f,i}^{\ell}(P_i v_H)\in V_{f,i}^{\ell}\) from Section 3 is obtained by the same constitutive law on the patch. Define \[u_i^{\ell}=v_H+v_{f,i}^{\ell}(P_i v_H) \label{eq:model-local-state-definition}\tag{91}\] Then the nonlinear patch problem is \[\begin{align} \int_{\omega_i^{\ell}} a_{\varepsilon}(x) \bigl(1+\alpha|\nabla u_i^{\ell}|^2\bigr) \nabla u_i^{\ell}\cdot\nabla w_{f,i}\,dx &=\int_{\omega_i^{\ell}}fw_{f,i}\,dx \quad \forall w_{f,i}\in V_{f,i}^{\ell} \label{eq:model-local-patch-problem} \end{align}\tag{92}\] Only the restriction of \(v_H\) to \(\omega_i^{\ell}\), represented by \(P_i v_H\), enters 92 . For a direction \(\delta v_H\in V_H\), the patch tangent bilinear form is given by 93 –94 , \[\begin{align} B_i^{\ell}(z_h,w_{f,i}) &=\int_{\omega_i^{\ell}} a_{\varepsilon}(x) \bigl(1+\alpha|\nabla u_i^{\ell}|^2\bigr) \nabla z_h\cdot\nabla w_{f,i}\,dx \tag{93}\\ &\quad+2\alpha\int_{\omega_i^{\ell}} a_{\varepsilon}(x) (\nabla u_i^{\ell}\cdot\nabla z_h) (\nabla u_i^{\ell}\cdot\nabla w_{f,i})\,dx \tag{94} \end{align}\] The tangent correction \(q_{f,i}^{\ell}\in V_{f,i}^{\ell}\) in the direction \(\delta v_H\) is defined by \[B_i^{\ell}(\delta v_H+q_{f,i}^{\ell},w_{f,i})=0 \quad \forall w_{f,i}\in V_{f,i}^{\ell} \label{eq:model-local-tangent-patch-problem}\tag{95}\] and uniqueness gives the identity \[q_{f,i}^{\ell}=D v_{f,i}^{\ell}(P_i v_H)[P_i\delta v_H] \label{eq:model-local-tangent-correction-identity}\tag{96}\] The coercivity bound 85 holds on each patch with the same lower constant \(a_0\), since functions in \(V_{f,i}^{\ell}\) satisfy homogeneous patch boundary conditions. Thus the tangent patch problems are uniformly solvable.
After choosing a basis of \(V_{f,i}^{\ell}\), 92 is a finite-dimensional nonlinear algebraic system with parameter \(P_i v_H\). The Jacobian of this system with respect to the fine-scale unknown is the tangent patch operator in 95 . Strong monotonicity gives uniqueness of the nonlinear patch correction, tangent coercivity gives invertibility of the tangent patch operator, and the implicit-function theorem gives differentiability of the local patch map on admissible parameter sets. The restricted maps \(\mathcal{G}_{f,i}^{\ell}=R_i v_{f,i}^{\ell}\) are therefore well-defined \(C^1\) targets for the optional network interpolation in Section 4.
This section gives a discrete perturbation framework for the localized manifold method. All estimates are relative to the fine-scale Galerkin solution \(u_h\). The finite element discretization error between \(u_h\) and the exact weak solution is not included. The main point is to separate the geometric stability mechanism from the localization and residual defects. In particular, the chord error generated by the curvature of the reconstructed manifold is absorbed into the monotonicity constant, while residual defects from inexact nonlinear solves or learned surrogates remain on the right-hand side.
For \(z_h\in V_h\), define the residual functional by \[\operatorname{Res}(z_h)(w_h)=L(w_h)-A(z_h;w_h) \quad \forall w_h\in V_h \label{eq:error-residual-functional}\tag{97}\] The fine-scale solution satisfies \(\operatorname{Res}(u_h)(w_h)=0\) for all \(w_h\in V_h\).
The argument is local in the nonlinear state. We assume that the relevant fine-scale states belong to an admissible set \(\mathcal{U}_h\subset V_h\), and that the relevant coarse states belong to an admissible set \(\mathcal{U}_H\subset V_H\). On \(\mathcal{U}_h\), the fine-scale operator is strongly monotone and locally Lipschitz continuous, \[m_A\|v_h-w_h\|_V^2\le A(v_h;v_h-w_h)-A(w_h;v_h-w_h) \quad \forall v_h,w_h\in\mathcal{U}_h \label{eq:error-strong-monotonicity-assumption}\tag{98}\] and \[\|A(v_h;\cdot)-A(w_h;\cdot)\|_{V_h'} \le L_A\|v_h-w_h\|_V \quad \forall v_h,w_h\in\mathcal{U}_h \label{eq:error-lipschitz-continuity-assumption}\tag{99}\] with constants \(m_A>0\) and \(L_A>0\). Section 6 verifies these assumptions for the heterogeneous monotone diffusion model.
The ideal global reconstruction \(M_h\) and the computable reconstructions used below are sections of the coarse projection. Thus, \[\begin{align} \Pi_H M_h(v_H)&=v_H \quad \forall v_H\in\mathcal{U}_H \tag{100}\\ \Pi_H\widetilde{M}_h^\ell(v_H)&=v_H \quad \forall v_H\in\mathcal{U}_H \tag{101} \end{align}\] The first identity implies the exact representation of the fine-scale Galerkin solution, \[u_h=M_h(\Pi_Hu_h) \label{eq:error-exact-global-manifold-representation}\tag{102}\] We also assume that \(\Pi_H\) is stable in the \(V\)-norm, \[\|\Pi_Hv_h\|_V\le C_\Pi\|v_h\|_V \quad \forall v_h\in V_h \label{eq:error-projection-stability}\tag{103}\]
Let \(\widetilde{M}_h^\ell\) denote either the patch-solved reconstruction or the network-interpolated reconstruction, \[\widetilde{M}_h^\ell=M_h^\ell \quad \text{or} \quad \widetilde{M}_h^\ell=M_{h,\theta}^\ell \label{eq:error-generic-reconstruction}\tag{104}\] and let \(\widetilde{u}_H^\ell\in V_H\) be the computed coarse state. We set \[\widetilde{u}_{\mathrm{ms}}^\ell=\widetilde{M}_h^\ell(\widetilde{u}_H^\ell) \label{eq:error-generic-ms-solution}\tag{105}\] and compare it with the same reconstruction evaluated at the exact coarse coordinate, \[\overline{u}_h^\ell=\widetilde{M}_h^\ell(\Pi_Hu_h) \label{eq:error-comparison-state}\tag{106}\] The coarse displacement and the coarse segment used below are \[\begin{align} \delta u_H^\ell&=\Pi_Hu_h-\widetilde{u}_H^\ell \tag{107}\\ v_H(s)&=\widetilde{u}_H^\ell+s\delta u_H^\ell \quad 0\le s\le 1 \tag{108} \end{align}\] We assume that \(v_H(s)\in\mathcal{U}_H\) for \(0\le s\le1\), and that \(u_h\), \(\widetilde{u}_{\mathrm{ms}}^\ell\), and \(\overline{u}_h^\ell\) belong to \(\mathcal{U}_h\). We also assume that \(\widetilde{M}_h^\ell\) is differentiable on \(\mathcal{U}_H\).
The localized and network-interpolated reconstructions assemble restricted patch outputs on \(\omega_i^0\). We use the following stability bounds. For every family \((w_i)_{i\in\mathcal{I}}\), with \(w_i\in W_{f,i}^0\), \[\biggl\| (I-\Pi_H)I_h \biggl( \sum_{i\in\mathcal{I}}\varphi_{H,i}E_{i,0}^{\operatorname{ext}}w_i \biggr) \biggr\|_V \le C_{\operatorname{pu}} \biggl(\sum_{i\in\mathcal{I}}\|w_i\|_{H^1(\omega_i^0)}^2\biggr)^{1/2} \label{eq:partition-of-unity-stability}\tag{109}\] Moreover, for each local coarse restriction \(P_i\), there is a constant \(C_{P,i}\) such that \[\|P_i\delta v_H\|_{\mathbb{R}^{n_{H,i}}} \le C_{P,i}\|\delta v_H\|_V \quad \forall \delta v_H\in V_H \label{eq:local-coarse-restriction-stability}\tag{110}\] The Euclidean norm in 110 is understood after the local coordinate scaling used for the patch problems and training sets. With this convention, the constants are stable under quasi-uniform refinement.
The perturbation estimate is obtained by testing the residual at the computed multiscale state with the true error. The key point is that this error can be decomposed into three geometrically distinct contributions. First, we compare the ideal and localized reconstructions at the exact coarse coordinate. Second, we measure the difference between the chord on the localized manifold and its tangent approximation at the computed coarse state. Third, we isolate the tangent vector generated by the coarse-scale error, which is an admissible test function in the coarse problem, \[\begin{align} u_h-\widetilde{u}_{\mathrm{ms}}^{\ell} &= M_h(\Pi_Hu_h)- \widetilde{M}_h^{\ell}(\widetilde{u}_H^\ell) \tag{111}\\ &=M_h(\Pi_Hu_h) - \widetilde{M}_h^{\ell}(\Pi_Hu_h) + \widetilde{M}_h^{\ell}(\Pi_Hu_h) - \widetilde{M}_h^{\ell}(\widetilde{u}_H^\ell) \tag{112}\\ &=M_h(\Pi_Hu_h) - \widetilde{M}_h^{\ell}(\Pi_Hu_h) \tag{113}\\ &\quad+ \widetilde{M}_h^{\ell}(\Pi_Hu_h) - \Big( \widetilde{M}_h^{\ell}(\widetilde{u}_H^\ell) + D\widetilde{M}_h^{\ell}(\widetilde{u}_H^{\ell}) [\Pi_Hu_h-\widetilde{u}_H^{\ell}] \Big) \tag{114}\\ &\quad+ D\widetilde{M}_h^{\ell}(\widetilde{u}_H^{\ell}) [\Pi_Hu_h-\widetilde{u}_H^{\ell}] \tag{115} \end{align}\] We denote the three contributions in the final decomposition by \[\begin{align} e_{\mathrm{loc}}^\ell &\mathrel{\vcenter{:}}= M_h(\Pi_Hu_h) - \widetilde{M}_h^{\ell}(\Pi_Hu_h) \tag{116}\\ e_{\mathrm{ch}}^\ell &\mathrel{\vcenter{:}}= \widetilde{M}_h^{\ell}(\Pi_Hu_h) - \Big( \widetilde{M}_h^{\ell}(\widetilde{u}_H^\ell) + D\widetilde{M}_h^{\ell}(\widetilde{u}_H^{\ell}) [\Pi_Hu_h-\widetilde{u}_H^{\ell}] \Big) \tag{117}\\ z_{\mathrm{tan}}^\ell &\mathrel{\vcenter{:}}= D\widetilde{M}_h^{\ell}(\widetilde{u}_H^{\ell}) [\Pi_Hu_h-\widetilde{u}_H^{\ell}] \tag{118} \end{align}\] Thus, \[u_h-\widetilde{u}_{\mathrm{ms}}^{\ell} = e_{\mathrm{loc}}^\ell + e_{\mathrm{ch}}^\ell + z_{\mathrm{tan}}^\ell \label{eq:error-split-compact}\tag{119}\]
The term \(e_{\mathrm{loc}}^\ell\) is the localization error, evaluated at the exact coarse coordinate \(\Pi_Hu_h\). The term \(e_{\mathrm{ch}}^\ell\) is the chord error, measuring the nonlinear defect between the localized chord and the tangent approximation at \(\widetilde{u}_H^\ell\). Finally, \(z_{\mathrm{tan}}^\ell\) is the image of the coarse-scale error \(\Pi_Hu_h-\widetilde{u}_H^\ell\) under the localized tangent map. In particular, \(z_{\mathrm{tan}}^\ell\) belongs to the admissible tangent test space used in the coarse problem.
The estimates close because the coarse displacement is the coarse projection of the full graph error. Set \[e_h\mathrel{\vcenter{:}}= u_h-\widetilde{u}_{\mathrm{ms}}^\ell \label{eq:error-total-error-vector}\tag{120}\] Using 100 and 101 , we obtain \[\begin{align} \delta u_H^\ell &=\Pi_Hu_h-\Pi_H\widetilde{u}_{\mathrm{ms}}^\ell \tag{121}\\ &=\Pi_H(u_h-\widetilde{u}_{\mathrm{ms}}^\ell) \tag{122}\\ &=\Pi_He_h \tag{123} \end{align}\] Therefore, by 103 , \[\|\delta u_H^\ell\|_V\le C_\Pi\|e_h\|_V \label{eq:error-coarse-error-projection-bound}\tag{124}\] This observation is used twice below. Quadratic dependence on \(\delta u_H^\ell\), as in the chord error, is absorbed into the monotonicity constant. Linear dependence on \(\delta u_H^\ell\), as in a reduced residual or a tangent-surrogate defect, remains as a right-hand-side perturbation.
The chord error is a curvature term for the reconstruction actually used in the coarse problem. It is not a separate localization defect. Under a local Lipschitz bound for the tangent reconstruction, the chord error is quadratic in the coarse displacement, so the residual it contributes is cubic in the multiscale error and can be kicked back into the monotonicity estimate once that error is small.
Assume that \(D\widetilde{M}_h^\ell\) is Lipschitz on \(\mathcal{U}_H\), in the sense that \[\begin{align} \bigl\| \bigl( D\widetilde{M}_h^\ell(v_H) - D\widetilde{M}_h^\ell(w_H) \bigr)[\xi_H] \bigr\|_V &\le L_{\widetilde{M},1}^\ell \|v_H-w_H\|_V \|\xi_H\|_V \label{eq:error-localized-reconstruction-derivative-lipschitz}\\ &\quad \forall v_H,w_H\in\mathcal{U}_H \quad \forall \xi_H\in V_H \notag \end{align}\tag{125}\] Then the chord residual satisfies \[\bigl| \operatorname{Res}(\widetilde{u}_{\mathrm{ms}}^\ell)(e_{\mathrm{ch}}^\ell) \bigr| \le \gamma_{\mathrm{ch}}^\ell \|e_h\|_V^3 \label{eq:error-chord-residual-kickback-bound}\tag{126}\] with \[\gamma_{\mathrm{ch}}^\ell \mathrel{\vcenter{:}}= \frac{1}{2} L_A L_{\widetilde{M},1}^\ell C_\Pi^2 \label{eq:error-chord-kickback-constant}\tag{127}\]
Proof. By the definition of \(e_{\mathrm{ch}}^\ell\) and the mean-value identity along the segment 108 , \[\begin{align} e_{\mathrm{ch}}^\ell &= \widetilde{M}_h^\ell(\widetilde{u}_H^\ell+\delta u_H^\ell) - \widetilde{M}_h^\ell(\widetilde{u}_H^\ell) - D\widetilde{M}_h^\ell(\widetilde{u}_H^\ell)[\delta u_H^\ell] \tag{128}\\ &= \int_0^1 \bigl( D\widetilde{M}_h^\ell(v_H(s)) - D\widetilde{M}_h^\ell(\widetilde{u}_H^\ell) \bigr) [\delta u_H^\ell] \,ds \tag{129} \end{align}\] Using 125 , we get \[\begin{align} \|e_{\mathrm{ch}}^\ell\|_V &\le \int_0^1 L_{\widetilde{M},1}^\ell \|v_H(s)-\widetilde{u}_H^\ell\|_V \|\delta u_H^\ell\|_V \,ds \tag{130}\\ &= \int_0^1 sL_{\widetilde{M},1}^\ell \|\delta u_H^\ell\|_V^2 \,ds \tag{131}\\ &= \frac{1}{2} L_{\widetilde{M},1}^\ell \|\delta u_H^\ell\|_V^2 \tag{132} \end{align}\] Since \(\operatorname{Res}(u_h)=0\), the Lipschitz continuity 99 bounds the residual at the computed state by the error, \[\|\operatorname{Res}(\widetilde{u}_{\mathrm{ms}}^\ell)\|_{V_h'} =\|A(u_h;\cdot)-A(\widetilde{u}_{\mathrm{ms}}^\ell;\cdot)\|_{V_h'} \le L_A\|e_h\|_V \label{eq:error-residual-at-computed-state}\tag{133}\] Combining this with the chord bound 132 and the projection bound 124 gives \[\begin{align} \bigl| \operatorname{Res}(\widetilde{u}_{\mathrm{ms}}^\ell)(e_{\mathrm{ch}}^\ell) \bigr| &\le \|\operatorname{Res}(\widetilde{u}_{\mathrm{ms}}^\ell)\|_{V_h'}\|e_{\mathrm{ch}}^\ell\|_V \tag{134}\\ &\le \frac{1}{2} L_A L_{\widetilde{M},1}^\ell \|e_h\|_V\|\delta u_H^\ell\|_V^2 \tag{135}\\ &\le \frac{1}{2} L_A L_{\widetilde{M},1}^\ell C_\Pi^2 \|e_h\|_V^3 \tag{136} \end{align}\] This is 126 with \(\gamma_{\mathrm{ch}}^\ell\) defined by 127 . ◻
We first consider the clean case in which the chosen localized or learned coarse problem is solved exactly. In this case the tangent term \(z_{\mathrm{tan}}^\ell\) in 119 makes no contribution to the residual. This is not because \(z_{\mathrm{tan}}^\ell\) vanishes—in general it does not—but because the residual tested against it vanishes, \(\operatorname{Res}(\widetilde{u}_{\mathrm{ms}}^\ell)(z_{\mathrm{tan}}^\ell)=0\): the vector \(z_{\mathrm{tan}}^\ell\) is an admissible tangent test function, and the exactly solved coarse equation 137 forces the residual to vanish on the entire tangent test space (the nonlinear analogue of Galerkin orthogonality). The only remaining right-hand-side defect is then the localization or reconstruction error \(e_{\mathrm{loc}}^\ell\).
Let \(\widetilde{u}_{\mathrm{ms}}^\ell=\widetilde{M}_h^\ell(\widetilde{u}_H^\ell)\) and assume 98 , 99 , and the hypotheses of Lemma [lem:chord-curvature-kickback]. Assume further that the coarse state solves the tangent-space coarse problem exactly, \[\operatorname{Res}(\widetilde{u}_{\mathrm{ms}}^\ell) \bigl( D\widetilde{M}_h^\ell(\widetilde{u}_H^\ell)[\xi_H] \bigr) =0 \quad \forall \xi_H\in V_H \label{eq:error-exact-coarse-problem}\tag{137}\] If the error is sufficiently small that \(\gamma_{\mathrm{ch}}^\ell\|e_h\|_V\le\tfrac12 m_A\), then \[\|u_h-\widetilde{u}_{\mathrm{ms}}^\ell\|_V \le \frac{2L_A}{m_A} \|e_{\mathrm{loc}}^\ell\|_V \label{eq:error-exact-coarse-solve-bound}\tag{138}\]
Proof. Set \(e_h=u_h-\widetilde{u}_{\mathrm{ms}}^\ell\). Strong monotonicity and the fine-scale equation give \[\begin{align} m_A\|e_h\|_V^2 &\le A(u_h;e_h)-A(\widetilde{u}_{\mathrm{ms}}^\ell;e_h) \tag{139}\\ &= L(e_h)-A(\widetilde{u}_{\mathrm{ms}}^\ell;e_h) \tag{140}\\ &= \operatorname{Res}(\widetilde{u}_{\mathrm{ms}}^\ell)(e_h) \tag{141} \end{align}\] Using 119 , the triangle inequality, and 137 with \(\xi_H=\delta u_H^\ell\), we obtain \[\begin{align} m_A\|e_h\|_V^2 &\le \bigl| \operatorname{Res}(\widetilde{u}_{\mathrm{ms}}^\ell)(e_{\mathrm{loc}}^\ell) \bigr| + \bigl| \operatorname{Res}(\widetilde{u}_{\mathrm{ms}}^\ell)(e_{\mathrm{ch}}^\ell) \bigr| \label{eq:error-exact-proof-split-bound} \end{align}\tag{142}\] The localization term is controlled by Lipschitz continuity. Since \(u_h\) solves the fine-scale problem, \[\begin{align} \bigl| \operatorname{Res}(\widetilde{u}_{\mathrm{ms}}^\ell)(e_{\mathrm{loc}}^\ell) \bigr| &= \bigl| A(u_h;e_{\mathrm{loc}}^\ell) - A(\widetilde{u}_{\mathrm{ms}}^\ell;e_{\mathrm{loc}}^\ell) \bigr| \tag{143}\\ &\le L_A\|e_h\|_V\|e_{\mathrm{loc}}^\ell\|_V \tag{144} \end{align}\] The chord term is bounded by Lemma [lem:chord-curvature-kickback]. Hence, \[m_A\|e_h\|_V^2 \le L_A\|e_{\mathrm{loc}}^\ell\|_V\|e_h\|_V + \gamma_{\mathrm{ch}}^\ell\|e_h\|_V^3 \label{eq:error-exact-proof-before-kickback}\tag{145}\] Under the smallness assumption \(\gamma_{\mathrm{ch}}^\ell\|e_h\|_V\le\tfrac12 m_A\), the chord term is at most \(\tfrac12 m_A\|e_h\|_V^2\). Moving it to the left-hand side gives \[\tfrac12 m_A\|e_h\|_V^2 \le L_A\|e_{\mathrm{loc}}^\ell\|_V\|e_h\|_V \label{eq:error-exact-proof-after-kickback}\tag{146}\] If \(e_h\neq0\), division by \(\|e_h\|_V\) gives 138 . If \(e_h=0\), the estimate is immediate. ◻
For the patch-solved method, the computable map is \(M_h^\ell\). Its deterministic state reconstruction error is measured relative to the ideal global map \(M_h\). On the admissible coarse set \(\mathcal{U}_H\), define \[\eta_{\mathrm{loc}}(H,\ell) \mathrel{\vcenter{:}}= \sup_{v_H\in\mathcal{U}_H} \|M_h(v_H)-M_h^\ell(v_H)\|_V \label{eq:deterministic-localization-state-defect}\tag{147}\] Equivalently, one may write \(\eta_{\mathrm{loc}}^0=\eta_{\mathrm{loc}}(H,\ell)\). Since \(u_h=M_h(\Pi_Hu_h)\), the localization vector for the patch-solved method satisfies \[\|e_{\mathrm{loc}}^\ell\|_V = \|M_h(\Pi_Hu_h)-M_h^\ell(\Pi_Hu_h)\|_V \le \eta_{\mathrm{loc}}(H,\ell) \label{eq:localized-state-defect-bound}\tag{148}\] Combining 148 with Proposition [prop:exact-localized-coarse-solve] gives the following deterministic consequence.
Let \(u_{\mathrm{ms}}^\ell=M_h^\ell(u_H^\ell)\) be the patch-solved localized solution, and assume that \(u_H^\ell\) solves the localized tangent-space coarse problem exactly. If the assumptions of Proposition [prop:exact-localized-coarse-solve] hold with \(\widetilde{M}_h^\ell=M_h^\ell\), then \[\|u_h-u_{\mathrm{ms}}^\ell\|_V \le \frac{2L_A}{m_A} \eta_{\mathrm{loc}}(H,\ell) \label{eq:patch-solved-localization-bound}\tag{149}\] provided the multiscale error is sufficiently small that \(\gamma_{\mathrm{ch}}^\ell\|u_h-u_{\mathrm{ms}}^\ell\|_V\le\tfrac12 m_A\).
The present paper uses \(\eta_{\mathrm{loc}}(H,\ell)\) as the deterministic localization quantity. A proof of decay in terms of \(H\), the oversampling parameter \(\ell\), coefficient contrast, and nonlinear stability constants is deferred to a forthcoming localization analysis. Thus Proposition [prop:patch-solved-localization-bound] should be read as a stability transfer result: any estimate for \(\eta_{\mathrm{loc}}(H,\ell)\) immediately yields a corresponding multiscale error estimate.
We now allow the tangent-space coarse problem to be solved inexactly. The third term in 119 is then controlled by the reduced residual of the coarse problem.
For a reconstruction \(\widetilde{M}_h^\ell\), define the reduced residual by \[\mathcal{R}_H^\ell(w_H)(\xi_H) \mathrel{\vcenter{:}}= \operatorname{Res}\bigl(\widetilde{M}_h^\ell(w_H)\bigr) \bigl( D\widetilde{M}_h^\ell(w_H)[\xi_H] \bigr) \quad \forall \xi_H\in V_H \label{eq:error-reduced-residual}\tag{150}\] Its dual norm is \[\|\mathcal{R}_H^\ell(w_H)\|_{V_H'} \mathrel{\vcenter{:}}= \sup_{\xi_H\in V_H\setminus\{0\}} \frac{ |\mathcal{R}_H^\ell(w_H)(\xi_H)| }{ \|\xi_H\|_V } \label{eq:error-reduced-residual-dual-norm}\tag{151}\] By the definition of \(z_{\mathrm{tan}}^\ell\), the tangent residual contribution is \[\operatorname{Res}(\widetilde{u}_{\mathrm{ms}}^\ell)(z_{\mathrm{tan}}^\ell) = \mathcal{R}_H^\ell(\widetilde{u}_H^\ell)(\delta u_H^\ell) \label{eq:error-algebraic-term-as-reduced-residual}\tag{152}\] Consequently, 124 gives \[\begin{align} \bigl| \operatorname{Res}(\widetilde{u}_{\mathrm{ms}}^\ell)(z_{\mathrm{tan}}^\ell) \bigr| &\le \|\mathcal{R}_H^\ell(\widetilde{u}_H^\ell)\|_{V_H'} \|\delta u_H^\ell\|_V \tag{153}\\ &\le C_\Pi \|\mathcal{R}_H^\ell(\widetilde{u}_H^\ell)\|_{V_H'} \|e_h\|_V \tag{154} \end{align}\]
Let \(\widetilde{u}_{\mathrm{ms}}^\ell=\widetilde{M}_h^\ell(\widetilde{u}_H^\ell)\). Assume 98 , 99 , and the hypotheses of Lemma [lem:chord-curvature-kickback]. If the error is sufficiently small that \(\gamma_{\mathrm{ch}}^\ell\|e_h\|_V\le\tfrac12 m_A\), then \[\|u_h-\widetilde{u}_{\mathrm{ms}}^\ell\|_V \le \frac{2}{m_A} \bigl( L_A\|e_{\mathrm{loc}}^\ell\|_V + C_\Pi\|\mathcal{R}_H^\ell(\widetilde{u}_H^\ell)\|_{V_H'} \bigr) \label{eq:error-inexact-coarse-solve-bound}\tag{155}\]
Proof. The proof is identical to the proof of Proposition [prop:exact-localized-coarse-solve], except that the tangent residual term is not zero. Using 119 , Lemma [lem:chord-curvature-kickback], 144 , and 154 , we obtain \[\begin{align} m_A\|e_h\|_V^2 &\le L_A\|e_{\mathrm{loc}}^\ell\|_V\|e_h\|_V + \gamma_{\mathrm{ch}}^\ell\|e_h\|_V^3 + C_\Pi\|\mathcal{R}_H^\ell(\widetilde{u}_H^\ell)\|_{V_H'}\|e_h\|_V \label{eq:error-inexact-proof-before-kickback} \end{align}\tag{156}\] Under the smallness assumption \(\gamma_{\mathrm{ch}}^\ell\|e_h\|_V\le\tfrac12 m_A\), the chord term is at most \(\tfrac12 m_A\|e_h\|_V^2\). Moving it to the left-hand side gives \[\tfrac12 m_A\|e_h\|_V^2 \le \bigl( L_A\|e_{\mathrm{loc}}^\ell\|_V + C_\Pi\|\mathcal{R}_H^\ell(\widetilde{u}_H^\ell)\|_{V_H'} \bigr) \|e_h\|_V \label{eq:error-inexact-proof-after-kickback}\tag{157}\] If \(e_h\neq0\), division by \(\|e_h\|_V\) gives 155 . If \(e_h=0\), the estimate is immediate. ◻
For Newton-type reduced solves, the dual norm of the reduced residual is the natural stopping quantity. If the iteration is stopped when \[\|\mathcal{R}_H^\ell(\widetilde{u}_H^\ell)\|_{V_H'} \le \tau_{\mathrm{alg}} \label{eq:error-newton-stopping-criterion}\tag{158}\] then 155 gives \[\|u_h-\widetilde{u}_{\mathrm{ms}}^\ell\|_V \le \frac{2}{m_A} \bigl( L_A\|e_{\mathrm{loc}}^\ell\|_V + C_\Pi\tau_{\mathrm{alg}} \bigr) \label{eq:error-newton-stopping-error-bound}\tag{159}\] Thus algebraic error enters additively through the reduced residual norm, while the chord curvature is absorbed by monotonicity under the smallness assumption.
After choosing local bases, the correction problem on a patch \(\omega_i^\ell\) is a finite-dimensional system \(F_i(c_H,c_f)=0\), whose solution defines the implicit map \(c_f=\Psi_i(c_H)\). If \(F_i\) is \(C^r\) and the fine-scale Jacobian \(\partial_{c_f}F_i\) is uniformly invertible along an admissible parameter set—precisely the tangent coercivity that makes the tangent patch solves well posed—then the implicit function theorem yields \(\Psi_i\in C^r\) with \[D\Psi_i(c_H) =-\partial_{c_f}F_i(c_H,\Psi_i(c_H))^{-1}\,\partial_{c_H}F_i(c_H,\Psi_i(c_H)) \label{eq:tangent-patch-map-coefficients}\tag{160}\] whose variational form is the tangent patch problem of Sections 3 and 6. The restricted output map \(\mathcal{G}_{f,i}^\ell\) of Section 4 is the composition of \(\Psi_i\) with a linear restriction and therefore inherits the same regularity. This regularity shows that the restricted patch outputs and their tangents are well-defined smooth targets for the network interpolation and for the chord-curvature bound.
The network-interpolated method is treated as a perturbation of the deterministic localized reconstruction. The present paper only uses realized approximation quantities and residual norms. Explicit a priori estimates for these quantities, including possible Barron-type rates for the restricted patch maps, are deferred to the forthcoming localization and learning analysis.
For admissible \(v_H\), assume that the restricted learned patch outputs satisfy \[\|R_i v_{f,i}^\ell(P_i v_H)-g_{f,i}^{\ell,\theta_i}(P_i v_H)\|_{H^1(\omega_i^0)} \le \varepsilon_i^0 \label{eq:local-network-state-error}\tag{161}\] The partition-of-unity stability estimate 109 gives the corresponding global state perturbation, \[\begin{align} \|M_h^\ell(v_H)-M_{h,\theta}^\ell(v_H)\|_V &\le C_{\operatorname{pu}} \biggl(\sum_{i\in\mathcal{I}}(\varepsilon_i^0)^2\biggr)^{1/2} \tag{162}\\ \eta_{\mathrm{net}}^0 &\mathrel{\vcenter{:}}= C_{\operatorname{pu}} \biggl(\sum_{i\in\mathcal{I}}(\varepsilon_i^0)^2\biggr)^{1/2} \tag{163} \end{align}\] If derivative information is also used, the analogous active-region tangent error can be measured by \[\eta_{\mathrm{net}}^1 \mathrel{\vcenter{:}}= C_{\operatorname{pu}} \biggl(\sum_{i\in\mathcal{I}}(C_{P,i}\varepsilon_i^1)^2\biggr)^{1/2} \label{eq:network-tangent-error-definition}\tag{164}\] where \(\varepsilon_i^1\) denotes the local derivative error in the norm of 110 . This tangent quantity is not needed as a separate additive term when the coarse problem is defined by \(M_{h,\theta}^\ell\) and the tangent action is computed consistently as \(DM_{h,\theta}^\ell\). It becomes relevant when a learned tangent action is used as an independent surrogate or when one compares learned and deterministic reduced residuals.
When the learned reconstruction \(M_{h,\theta}^\ell\) defines the actual coarse model, the localization vector becomes \[e_{\mathrm{loc},\theta}^\ell \mathrel{\vcenter{:}}= M_h(\Pi_Hu_h)-M_{h,\theta}^\ell(\Pi_Hu_h) \label{eq:network-localization-vector}\tag{165}\] By the triangle inequality and 162 , \[\begin{align} \|e_{\mathrm{loc},\theta}^\ell\|_V &\le \|M_h(\Pi_Hu_h)-M_h^\ell(\Pi_Hu_h)\|_V + \|M_h^\ell(\Pi_Hu_h)-M_{h,\theta}^\ell(\Pi_Hu_h)\|_V \tag{166}\\ &\le \eta_{\mathrm{loc}}(H,\ell)+\eta_{\mathrm{net}}^0 \tag{167} \end{align}\] Thus Propositions [prop:exact-localized-coarse-solve] and [prop:inexact-coarse-solve] apply with \(\widetilde{M}_h^\ell=M_{h,\theta}^\ell\). In particular, if the learned coarse problem is solved inexactly and the learned error is sufficiently small that \(\gamma_{\mathrm{ch},\theta}^\ell\|u_h-\widetilde{u}_{\mathrm{ms},\theta}^\ell\|_V\le\tfrac12 m_A\), then \[\|u_h-\widetilde{u}_{\mathrm{ms},\theta}^\ell\|_V \le \frac{2}{m_A} \bigl( L_A(\eta_{\mathrm{loc}}(H,\ell)+\eta_{\mathrm{net}}^0) + C_\Pi\|\mathcal{R}_{H,\theta}^\ell(\widetilde{u}_{H,\theta}^\ell)\|_{V_H'} \bigr) \label{eq:network-inexact-coarse-solve-bound}\tag{168}\] Here \(\mathcal{R}_{H,\theta}^\ell\) is the reduced residual from 150 with \(\widetilde{M}_h^\ell=M_{h,\theta}^\ell\). There are two complementary ways to account for the learning defect in this estimate. In the bound 168 the learned reconstruction is the coarse model: one sets \(\widetilde{M}_h^\ell=M_{h,\theta}^\ell\), and learning enters a priori through the localization term \(\eta_{\mathrm{net}}^0\) of 167 , through the learned reduced residual \(\mathcal{R}_{H,\theta}^\ell\), and through the learned curvature constant \(\gamma_{\mathrm{ch},\theta}^\ell\). Alternatively, a learned state may be viewed as a candidate for the deterministic localized problem: one applies Proposition [prop:inexact-coarse-solve] with the deterministic map \(\widetilde{M}_h^\ell=M_h^\ell\) and its deterministic constants, evaluated at the learned coarse state \(\widetilde{u}_{H,\theta}^\ell\). Then there is no separate \(\eta_{\mathrm{net}}^0\) term; the entire learning defect is folded into the single a posteriori, computable deterministic reduced residual \(\|\mathcal{R}_H^\ell(\widetilde{u}_{H,\theta}^\ell)\|_{V_H'}\), which plays exactly the same role as the algebraic residual of an early-stopped outer Newton iteration. The two viewpoints mirror the a priori/a posteriori alternative discussed in Section 7.8.
The estimates identify the quantities that should be measured independently. Oversampling studies probe the deterministic localization quantity \(\eta_{\mathrm{loc}}(H,\ell)\). Nonlinear-solver studies probe the reduced residual norm \(\|\mathcal{R}_H^\ell(\widetilde{u}_H^\ell)\|_{V_H'}\). Network studies can either report the same reduced residual as an a posteriori learning indicator, or report active-region state and tangent errors such as \(\eta_{\mathrm{net}}^0\) and \(\eta_{\mathrm{net}}^1\). The chord term is not a separate convergence defect in the final estimate; it enters through the smallness condition \(\gamma_{\mathrm{ch}}^\ell\|e_h\|_V\le\tfrac12 m_A\), which expresses that once the multiscale error is small enough, the local curvature of the reconstructed manifold is absorbed by monotonicity.
In this section we numerically test the proposed tangent-space multiscale manifold method by applying it to the monotone model problem in Section 6 with a heterogeneous coefficient. We first verify the patch-solved method, where the per-patch fine-scale corrections are computed by an inner Newton iteration. We then replace the restricted patch outputs by neural-network interpolants and solve the full learned coarse problem. Finally, we test the same stationary reconstruction in a backward-Euler discretization of a parabolic problem. We consider two strongly heterogeneous coefficients: a smooth oscillatory medium and a non-periodic discontinuous random checkerboard.
We use the heterogeneous monotone nonlinear diffusion model of Section 6 on \(\Omega=(0,1)^2\) with homogeneous Dirichlet boundary conditions: the energy \(\mathcal{E}_h\) 75 , the nonlinear form \(A=A_h\) 76 , and the load \(L=F_h\) 77 , with \(a_\varepsilon\in L^\infty(\Omega)\), \(0<a_0\le a_\varepsilon\le a_1\), and nonlinearity strength \(\alpha\ge0\). The error is measured in the energy (\(H^1\)) seminorm \(\|\nabla v\|_{\Omega}\).
The domain is triangulated on two nested structured meshes of \(P_1\) Lagrange elements. We fix a single fine reference mesh with \(h=1/64\), that is, \(64\) elements per side, throughout. Most of the experiments use a coarse mesh of \(N=8\) elements per side, so \(H=1/8\). We report the relative energy-norm error of the localized multiscale solution against the fine reference, \[\frac{\|\nabla(u_h-u_{\mathrm{ms}}^\ell)\|_{\Omega}}{\|\nabla u_h\|_{\Omega}} \label{eq:numerics-relative-error-l-to-ideal}\tag{169}\]
The coarse space uses homogeneous Dirichlet data and \(\Pi_H\) is the \(L^2\)-projection realized through the cross-mesh mass constraint. All reported results use vertex patches, one per coarse vertex, with \(\varphi_{H,z}=\Lambda_z\) the coarse hat function.
We report two heterogeneous media, shown in Figure 5. The smooth oscillatory coefficient is \[a_\varepsilon(x,y)=\frac{1}{2}\bigl(2+\beta\sin(2\pi x/\varepsilon)\sin(2\pi y/\varepsilon)\bigr)\quad \varepsilon=\frac{1}{8},\;\beta=1.8 \label{eq:numerics-sine-coefficient}\tag{170}\] Thus \(a_\varepsilon\in[0.1,1.9]\), a contrast ratio of \(19\). The second medium is a non-periodic random checkerboard: the unit square is partitioned into cells of size \(1/32\), and on each cell \(a_\varepsilon\) is drawn independently and uniformly from \([1,20]\) with a fixed seed. This gives a piecewise-constant field with contrast ratio up to \(20\) and no periodic structure. The fine mesh \(h=1/64\) resolves each checkerboard cell with \(2\times2\) elements. Throughout, \(f\equiv1\) and \(\alpha=1\). The non-periodicity of the checkerboard is a deliberate stress on the method, which makes no periodicity assumption.
Both the patch-solved and the network-interpolated coarse problems are solved by the same outer Newton iteration, with the coarse manifold residual \(r_H(u_H)[\delta v_H]=A(\widetilde{M}_h^\ell(u_H);D \widetilde{M}_h^\ell(u_H)[\delta v_H])-L(D \widetilde{M}_h^\ell(u_H)[\delta v_H])\). As noted in Section 5, its exact derivative in a direction \(w_H\) involves the second derivative of the reconstruction; it splits into two contributions, \[\begin{align} \mathfrak J_H(u_H)[w_H,\delta v_H] &=\mathfrak J_{H,\operatorname{gal}}^\ell(u_H)[w_H,\delta v_H] +\mathfrak J_{H,\operatorname{curv}}^\ell(u_H)[w_H,\delta v_H] \tag{171}\\ \mathfrak J_{H,\operatorname{gal}}^\ell(u_H)[w_H,\delta v_H] &=A'\bigl(\widetilde{M}_h^\ell(u_H)\bigr)\bigl[D \widetilde{M}_h^\ell(u_H)[w_H],D \widetilde{M}_h^\ell(u_H)[\delta v_H]\bigr] \tag{172}\\ \mathfrak J_{H,\operatorname{curv}}^\ell(u_H)[w_H,\delta v_H] &=\bigl\langle\mathcal{R}(\widetilde{M}_h^\ell(u_H)),D^2\widetilde{M}_h^\ell(u_H)[w_H,\delta v_H]\bigr\rangle \tag{173} \end{align}\] where \(\mathcal{R}(u_h)=A(u_h;\cdot)-L(\cdot)\) is the fine-scale residual. The Galerkin Jacobian keeps only 172 , so the outer iteration is an inexact Newton method of Gauss–Newton type. The motivation is threefold. First, \(J_H\) is symmetric positive definite whenever \(A'\) is. Second, the dropped term is controlled by duality, \[\bigl|\mathfrak J_{H,\operatorname{curv}}^\ell(u_H)[w_H,\delta v_H]\bigr| \le\bigl\|\mathcal{R}(\widetilde{M}_h^\ell(u_H))\bigr\|_{V_h'}\, \bigl\|D^2\widetilde{M}_h^\ell(u_H)[w_H,\delta v_H]\bigr\|_V \label{eq:numerics-curvature-duality-bound}\tag{174}\] where the second factor is bounded through the tangent regularity constant \(L_{\widetilde{M},1}^\ell\) of Section 7. For the ideal manifold the residual factor vanishes at the solution. For a localized or learned manifold, only the tangential part of the residual is driven to zero by the outer iteration, and the full fine residual at the computed solution remains of the size of the localization and learning defects, \(\|\mathcal{R}(\widetilde{u}_{\mathrm{ms}}^\ell)\|_{V_h'}\le L_A\|e_h\|_V\) by 133 . The neglected curvature contribution is therefore small of the order of the reconstruction defect. Third, assembling \(D^2\widetilde{M}_h^\ell\) would require differentiating the per-patch tangent solves a second time and, in the learned variant of Section 8.3, using second derivatives of the networks. In all experiments below the outer iteration with the Galerkin Jacobian converges in \(2\)–\(4\) steps.
Here the per-patch fine-scale corrections are solved by the inner Newton iteration; no surrogate is involved. The fine-scale solution \(u_h\) in 78 , obtained by full \(P_1\) Newton on the fine mesh, is used as a reference solution. Recall that, by the exactness of the global map, the ideal (no localization) multiscale solution equals \(u_h\).
We solve 62 on a fixed coarse mesh of size \(H=1/8\) with the localization parameter ranging from \(\ell=1,\ldots,6\). The relative energy-norm error is reported in Figure 6 and Table 1, which show exponential decay in \(\ell\). The decay rate is approximately \(1.0\) per layer for both coefficients. The random checkerboard is only marginally harder than the smooth medium, by a constant factor of about \(1.5\) and a slightly smaller rate, despite its piecewise-constant non-periodic structure. These results provide numerical evidence that the localization defect \(\eta_{\mathrm{loc}}(H,\ell)\) appearing in Proposition [prop:patch-solved-localization-bound] decays exponentially in \(\ell\). A rigorous a priori proof of this decay is not attempted here; it is deferred to the companion analysis manuscript.
| \(\ell\) | \(\sin_\varepsilon\) | Random checkerboard |
|---|---|---|
| \(1\) | \(2.18\times10^{-2}\) | \(3.21\times10^{-2}\) |
| \(2\) | \(8.41\times10^{-3}\) | \(1.05\times10^{-2}\) |
| \(3\) | \(3.23\times10^{-3}\) | \(4.38\times10^{-3}\) |
| \(4\) | \(1.14\times10^{-3}\) | \(1.69\times10^{-3}\) |
| \(5\) | \(3.64\times10^{-4}\) | \(6.59\times10^{-4}\) |
| \(6\) | \(1.17\times10^{-4}\) | \(2.03\times10^{-4}\) |
| Fit \(\beta\) | \(\approx1.05\) | \(\approx0.99\) |
We next fix the localization radius and refine the coarse mesh. Keeping the fine reference at \(h=1/64\), we sweep \(H=1/N\in\{1/4,1/8,1/16,1/32\}\), for \(\ell=1,\ldots,4\), with patch-solved fine-scale corrections. The relative energy-norm error against \(u_h\) is reported in Figure 7 and Table 2. We note that as \(H\) decreases, the patch radius \(\ell\) needs to grow to reduce the error. The dashed envelope marks the graph for \(\ell \sim \log(1/H)\). Following this envelope, the error decreases at the linear rate \(\mathcal{O}(H)\) in the energy (\(H^1\)) norm generally expected for such methods; this rate is observed numerically here, while its rigorous justification is deferred to the companion analysis manuscript. The same behaviour holds for both coefficients.
Furthermore, the black curve is plain coarse \(P_1\) FEM for comparison. It sits far above every localized curve and does not converge while \(H\) fails to resolve the fine scale. For the \(\sin_\varepsilon\) medium, the relative FEM error exceeds \(100\%\) for \(H\ge\varepsilon=1/8\) and only begins to fall once \(H<\varepsilon\). By contrast, the localized method reaches about \(1\%\) already at \(H=1/8\) with a few fine-scale correction layers.
| \(H=1/4\) | \(H=1/8\) | \(H=1/16\) | \(H=1/32\) | |
|---|---|---|---|---|
| \(\sin_\varepsilon\) | ||||
| Pure FEM | \(1.28\) | \(1.29\) | \(3.37\times10^{-1}\) | \(2.21\times10^{-1}\) |
| \(\ell=1\) | \(3.66\times10^{-2}\) | \(2.18\times10^{-2}\) | \(1.05\times10^{-1}\) | \(6.30\times10^{-2}\) |
| \(\ell=2\) | \(9.88\times10^{-3}\) | \(8.41\times10^{-3}\) | \(4.07\times10^{-2}\) | \(3.11\times10^{-2}\) |
| \(\ell=3\) | \(4.51\times10^{-3}\) | \(3.23\times10^{-3}\) | \(1.56\times10^{-2}\) | \(1.68\times10^{-2}\) |
| \(\ell=4\) | \(1.71\times10^{-3}\) | \(1.14\times10^{-3}\) | \(5.98\times10^{-3}\) | \(6.59\times10^{-3}\) |
| Random checkerboard | ||||
| Pure FEM | \(5.94\times10^{-1}\) | \(4.71\times10^{-1}\) | \(3.91\times10^{-1}\) | \(2.02\times10^{-1}\) |
| \(\ell=1\) | \(4.01\times10^{-2}\) | \(3.21\times10^{-2}\) | \(4.90\times10^{-2}\) | \(6.83\times10^{-2}\) |
| \(\ell=2\) | \(1.12\times10^{-2}\) | \(1.05\times10^{-2}\) | \(1.75\times10^{-2}\) | \(2.52\times10^{-2}\) |
| \(\ell=3\) | \(5.23\times10^{-3}\) | \(4.38\times10^{-3}\) | \(6.78\times10^{-3}\) | \(1.07\times10^{-2}\) |
| \(\ell=4\) | \(2.06\times10^{-3}\) | \(1.69\times10^{-3}\) | \(2.74\times10^{-3}\) | \(4.44\times10^{-3}\) |
To inspect localization at the level of a single fine-scale correction, fix \(u_H=\Pi_Hu_h\) and consider, on one patch \(i\), the difference between the exact global fine-scale correction \(v_f(u_H)=u_h-u_H\) restricted to the patch and the patch-solved local fine-scale correction \(v_{f,i}^\ell(P_i u_H)\), \[d_i=v_f(\Pi_Hu_h)\big|_{\omega_i^\ell}-v_{f,i}^\ell(P_i\Pi_Hu_h) \label{eq:numerics-single-patch-corrector-error}\tag{175}\] The local fine-scale correction is pinned to zero on the artificial patch boundary \(\partial\omega_i^\ell\), while the global one is not. Therefore, \(d_i\) forms a boundary layer at \(\partial\omega_i^\ell\) and decays toward the patch seed, as shown in Figure 8. As \(\ell\) grows, the patch spreads over more coarse elements and the error layer moves outward.
The physically meaningful measure is the error over the active support \(\omega_i^0\), the star of the vertex \(z\), since the partition-of-unity weight \(\varphi_{H,i}\) annihilates the fine-scale correction outside \(\omega_i^0\) before it enters 38 . The support-restricted seminorm \(\|\nabla d_i\|_{L^2(\omega_i^0)}\) decays exponentially in \(\ell\), with rate \(\beta\approx0.6\). By contrast, the seminorm over the full growing patch \(\|\nabla d_i\|_{L^2(\omega_i^\ell)}\) increases with \(\ell\), because the integration domain grows. These results motivate the support-restricted contributions in the assembled reconstruction map \(M^\ell_h\).
We replace each restricted patch-output map by a small neural network \(\mathcal{N}_{f,i}^{\ell,\theta_i}\), giving the network-interpolated method.
Each patch uses a shallow multilayer perceptron with a single hidden layer and \(\tanh\) activation, where \(\widehat x\) is the affinely normalized input. The input \(x=P_i u_H\) is the coarse function restricted to the patch’s coarse degrees of freedom. Coarse degrees of freedom on the global Dirichlet boundary \(\partial\Omega\) are pinned to zero and carry zero sensitivity. The output is the fine-scale correction on the free (non-Dirichlet) fine nodes of the patch. The hidden width is denoted by \(n_{\mathrm{hid}}\), with \(n_{\mathrm{hid}}=128\) in the production runs. All weights \(\{W_1,b_1,W_2,b_2\}\) are trained. The tangent action required by the outer Newton solve is the analytic one-layer Jacobian, \[D\mathcal{N}_{f,i}^{\ell,\theta_i}=\operatorname{diag}(\sigma_y)W_2\operatorname{diag}(1-\tanh^2)W_1\operatorname{diag}(1/\sigma_x) \label{eq:numerics-network-analytic-jacobian}\tag{177}\] so that, thanks to the shallow single-hidden-layer architecture, no automatic differentiation is needed at solve time.
For each patch, the inputs \(\{u_H^m\}\) are Sobol space-filling points in a box around a centre \(u_H^{\mathrm c}\), perturbing only the active local coarse degrees of freedom. The box half-width is \(0.5|u_H^{\mathrm c}|\), that is, relative radius \(0.5\), per active degree of freedom. Each target \(v_{f,i}^\ell(P_i u_H^m)\) is produced by the inner-Newton patch solve, so generating the dataset is the dominant offline cost. The production runs use \(N_{\mathrm{samp}}=512\) Sobol samples per patch.
Choosing the box centre \(u_H^{\mathrm c}\) is essential since each network is accurate only near the coarse states it was sampled at, so \(u_H^{\mathrm c}\) must lie where the outer Newton iteration 62 actually evaluates the networks at solve time. We choose \(u_H^{\mathrm c}=u_H^\ell\), the coarse state of the deterministic patch-solved multiscale method, which is precisely the state the learned outer solve converges to. Obtaining \(u_H^\ell\) requires one offline solve of the untrained patch-solved problem, which may seem an unnatural prerequisite for a single solve. Its cost, however, is marginal: the offline stage is dominated by generating the training data, which already runs the deterministic patch solver \(N_{\mathrm{samp}}\) times per patch, so the one extra coarse solve that fixes the centre is negligible. The real justification is amortization: once trained, the networks are reused across many online solves with no further patch solves. This is exploited in Section 8.4, where the networks are re-used at every step of a time-dependent problem.
The primary loss is the relative \(H^1(\omega_i^\ell)\) energy-seminorm of the output residual in the patch stiffness metric. We add Jacobian, or Sobolev, supervision that matches the network’s analytic input Jacobian to
the patch-solved fine-scale correction tangent \(D v_{f,i}^\ell\), available from the patch tangent block. This term is measured in the same \(H^1\) patch-energy metric \(A\) as the value loss: writing the Cholesky factorization \(A=LL^\top\), we penalize \(\|L^\top(D\mathcal{N}_{f,i}^{\ell,\theta_i}-D v_{f,i}^\ell)\|^2\), so that
a plain least-squares loss equals the energy-norm error of the tangent mismatch. It uses weight \(\lambda=1\). This directly trains the tangent \(D M_{h,\theta}^\ell\) used by the outer
Newton solve. Optimization uses Adam with learning rate \(10^{-3}\), full-batch training for \(2000\) epochs, an \(80/20\) train-validation split, and
best-validation-error checkpointing every \(50\) epochs. Training uses JAX/optax; the deployed network, including the forward pass and analytic tangent, uses only numpy/scipy.
The hidden width, sample count, and sampling-box radius were fixed by preliminary single-patch tuning: the out-of-sample error—measured on a fresh Sobol set drawn with an independent seed and disjoint from the training samples—plateaus beyond a moderate width and training-set size, and a tighter box lowers the in-distribution error at the cost of poor out-of-box extrapolation, which motivates the generous radius \(0.5\). On out-of-sample coarse states the learned fine-scale correction reproduces the patch-solved one well, and the \(L^2\)-orthogonality constraint is approximately inherited even though it is not enforced.
We compare value-only (\(\lambda=0\)) and value-plus-Jacobian (\(\lambda=1\)) supervision in the production setting: support-restricted networks at \(h=1/64\), for both coefficients, with the out-of-sample error measured over the support and centred at the multiscale coarse state \(u_H^\ell\) (Table 3, median over the interior patches). Jacobian supervision lowers both the value error and, more importantly, the tangent error that the outer Newton consumes: by about \(3.4\times\) (value) and \(3.9\times\) (tangent) for the smooth \(\sin_\varepsilon\) medium, and \(5.5\times\) and \(5.9\times\) for the checkerboard. The piecewise-constant fine-scale corrections remain markedly easier to learn than the smooth oscillatory ones at both settings.
| Value Relative \(H^1\) | Tangent Relative Error | |||
|---|---|---|---|---|
| Coefficient | \(\lambda=0\) | \(\lambda=1\) | \(\lambda=0\) | \(\lambda=1\) |
| \(\sin_\varepsilon\), smooth oscillatory | \(\sim9.1\%\) | \(\sim2.7\%\) | \(\sim24\%\) | \(\sim6.0\%\) |
| Checkerboard, piecewise constant | \(\sim6.2\%\) | \(\sim1.1\%\) | \(\sim16\%\) | \(\sim2.8\%\) |
We run the complete network-interpolated method: one network is trained per active patch, and the outer Newton problem 62 is solved with the learned reconstruction \(M_{h,\theta}^\ell\) and its analytic tangent \(D M_{h,\theta}^\ell\). The resulting solution is \[u_{\mathrm{ms},\theta}^\ell=M_{h,\theta}^\ell(u_{H,\theta}^\ell) \label{eq:numerics-learned-solution}\tag{178}\] We report the deviation from the patch-solved solution and, for context, the error of both methods against the fine reference. All \(81\) vertex patches are trained with \(\ell=2\), hidden width \(n_{\mathrm{hid}}=128\), \(N_{\mathrm{samp}}=512\), and \(\lambda=1\).
Only the restriction of each fine-scale correction to \(\operatorname{supp}\varphi_{H,i}=\omega_i^0\) enters 38 , since \(\varphi_{H,i}\) annihilates every node outside \(\omega_i^0\). We therefore shrink each network output layer to the support nodes. At \(h=1/64\), this means about \(136\) active output nodes out of about \(1030\) patch nodes on average. Besides cutting the parameter count of the dominant \(W_2\) block to \(1.70\times10^6\) over all patches, this concentrates capacity on the nodes that survive multiplication by \(\varphi_{H,i}\) and improves accuracy. All results below use these support-restricted networks.
With \(u_H^\ell\)-centred, support-restricted networks, the learned method reproduces the patch-solved multiscale solution to \(0.55\%\) for the \(\sin_\varepsilon\) medium and \(0.12\%\) for the checkerboard. The outer Newton solve with the learned reconstruction converges in \(3\)–\(4\) iterations, and the recovered coarse state is within \(0.02\)–\(0.03\%\) of the patch-solved one. End-to-end against the fine reference, the learned solve barely degrades the patch-solved accuracy, from \(0.84\%\) to \(1.03\%\) for the \(\sin_\varepsilon\) medium and from \(1.05\%\) to \(1.06\%\) for the checkerboard; see Table 4 and Figure 9.
| Coefficient | L/MS | MS/F | L/F | Coarse | Its |
|---|---|---|---|---|---|
| \(\sin_\varepsilon\) | \(5.49\times10^{-3}\) | \(0.84\times10^{-2}\) | \(1.03\times10^{-2}\) | \(2.4\times10^{-4}\) | \(4\) |
| Checkerboard | \(1.23\times10^{-3}\) | \(1.05\times10^{-2}\) | \(1.06\times10^{-2}\) | \(2.8\times10^{-4}\) | \(3\) |
Beyond the assembled solve, we validate the trained restricted networks at the level of individual patches. For each trained patch we compare the learned fine-scale correction to the patch-solved Newton one on a fresh out-of-sample Sobol set centred at the multiscale coarse state \(u_H^\ell\), measuring the relative \(H^1\) error over the support \(\omega_i^0\) — the only part of each correction that survives the partition-of-unity blend in 38 . The per-patch out-of-sample median error reaches about \(5\%\) for the \(\sin_\varepsilon\) medium (overall median \(\approx2\%\)) and stays below \(1.6\%\) for the checkerboard (overall median \(\approx0.7\%\)). The heatmaps and seed cross-sections confirm that the network reproduces the exact correction on the support, and the per-patch error bars (Figures 10 and 11) show the piecewise-constant corrections are easier to learn, in agreement with the aggregate accuracy above.
As a more demanding test, we add a time derivative. On \(\Omega=(0,1)^2\times(0,T]\), we seek \(u(\cdot,t)\) with \[\partial_t u-\nabla\cdot\bigl(a_\varepsilon(1+\alpha|\nabla u|^2)\nabla u\bigr)=f\quad \text{in }\Omega,\quad u=0\quad \text{on }\partial\Omega,\quad u(\cdot,0)=u_0 \label{eq:numerics-parabolic-problem}\tag{179}\] using the same monotone nonlinear diffusion as in Section 8.1, with \(f\equiv1\) and \(\alpha=1\). We use either a smooth bump \(u_0=x(x-1)y(y-1)\) or \(u_0\equiv0\).
Backward Euler with step \(\tau\) on the fine space reads: given \(u_h^{n-1}\in V_h\), find \(u_h^n\in V_h\) with \[\frac{1}{\tau}(u_h^n-u_h^{n-1},v_h)_\Omega + A(u_h^n;v_h) = (f,v_h)_\Omega\quad \forall v_h\in V_h \label{eq:numerics-backward-euler-fine}\tag{180}\] This problem is solved using a Newton iteration at each step, and the sequence \(\{u_h^n\}\) is the reference solution.
For the multiscale approach, the fine-scale corrections—and hence \(M_h^\ell\) and \(D M_h^\ell\)—are the time-independent stationary ones. The mass term and the history in the parabolic problem enter only the coarse manifold equation. This means that the learned corrections are computed once in an offline stage and can then be reused in every time step, giving the fine-scale corrections by a simple forward pass.
For the multiscale solution each step solves the problem: find \(u_H^n\in V_H\) such that, for all \(\delta v_H\in V_H\), \[\begin{align} &\frac{1}{\tau}\bigl(M_h^\ell(u_H^n)-u_{\mathrm{ms}}^{n-1},D M_h^\ell(u_H^n)[\delta v_H]\bigr)_\Omega \notag\\ &\qquad +A\bigl(M_h^\ell(u_H^n);D M_h^\ell(u_H^n)[\delta v_H]\bigr) =\bigl(f,D M_h^\ell(u_H^n)[\delta v_H]\bigr)_\Omega \quad \forall \delta v_H\in V_H \label{eq:numerics-backward-euler-coarse} \end{align}\tag{181}\] and the reconstructed solution is \(u^n_{\mathrm{ms}} = M^\ell_h(u^n_H)\). In matrix form, with \(\mathbf{M}_h\) the fine mass matrix, \(A_h(\cdot)\) the nonlinear operator action, \(b_h\) the load vector, and history \(u_{\mathrm{ms}}^{n-1}=M_h^\ell(u_H^{n-1})\), the equation is \[D M_h^\ell(u_H^n)^\top\biggl[\frac{1}{\tau}\mathbf{M}_h\bigl(M_h^\ell(u_H^n)-u_{\mathrm{ms}}^{n-1}\bigr)+A_h(M_h^\ell(u_H^n))-b_h\biggr]=0 \label{eq:numerics-backward-euler-coarse-matrix-form}\tag{182}\] It is solved by Newton with the Galerkin Jacobian \(J_H(v_H)=(D M_h^\ell(v_H))^\top(\tau^{-1}\mathbf{M}_h+A'(M_h^\ell(v_H)))D M_h^\ell(v_H)\). Relative to the stationary solve 62 , the only additions are the mass-history term in the residual and \(\tau^{-1}\mathbf{M}_h\) in the Jacobian. As in the stationary case, the omitted curvature term pairs the full backward-Euler fine residual, now including the mass-history term, with \(D^2M_h^\ell\); it is of the size of the reconstruction defect at convergence, and dropping it does not affect the converged time steps. At each step, the network-multiscale solution \(u_{\mathrm{ms},\theta}^n\) is compared to the patch-solved multiscale solution \(u_{\mathrm{ms}}^n\) and to the fine reference \(u_h^n\).
With \(\tau=0.02\) over \(10\) steps, Figures 12 and 13 show that the network-multiscale solution reproduces the patch-solved multiscale solution to its stationary accuracy once the trajectory lies in the networks’ training neighborhood. The error is approximately \(0.5\%\) for the \(\sin_\varepsilon\) medium and \(0.12\%\) for the checkerboard. The solution tracks the fine reference to approximately \(1.0\%\) for the \(\sin_\varepsilon\) medium and \(1.06\%\) for the checkerboard, close to the patch-solved multiscale discretization error (\(0.84\%\) and \(1.05\%\), respectively), so the learned interpolation adds little on top of the discretization floor. The outer Newton solve converges to tolerance \(10^{-8}\) at every step, requiring \(1\)–\(3\) iterations after relaxation.
The first-step accuracy is governed by how far the initial coarse state lies from the steady state, which is the networks’ training centre. This distance is measured by \(\|u_H^n-u_H^{\mathrm{stat}}\|/\|u_H^{\mathrm{stat}}\|\) in the right panels of Figures 12 and 13. Which initial condition is benign is coefficient-dependent. For the \(\sin_\varepsilon\) medium, the bump starts close to the steady state, at distance \(0.16\), and the network error is uniform at about \(0.5\%\). The zero initial condition starts farther away, at distance \(0.72\), and the network error spikes to about \(6.7\%\) at the first step before recovering. For the checkerboard, the situation reverses: the zero initial state starts close, at distance \(0.21\), while the bump starts far, at distance \(1.38\), and spikes to about \(4.4\%\) before recovering. This is an accuracy effect, not a solver effect; the mass term \(\tau^{-1}\mathbf{M}_h\) keeps the coarse Jacobian well conditioned.


Figure 12: \(\sin_\varepsilon\), backward Euler with \(\tau=0.02\) and restricted networks. Top: bump initial data, where the coarse state starts at distance \(0.16\) from steady state and the network error is uniform at about \(0.5\%\). Bottom: zero initial data, where the first step starts at distance \(0.72\) and the network error spikes to about \(6.7\%\) before recovering..


Figure 13: Random checkerboard, backward Euler with \(\tau=0.02\) and restricted networks. Top: bump initial data, where the trajectory starts at distance \(1.38\) from steady state and the network error spikes to about \(4.4\%\) before recovering. Bottom: zero initial data, where the trajectory starts at distance \(0.21\) and the network error stays between about \(0.12\%\) and \(0.36\%\)..
This experiment is small and we do not claim a wall-clock speed-up here. Its purpose is instead to demonstrate the potential of learning the corrections, and to indicate why one would consider learning in the first place. Because the reconstruction is stationary, the same networks are reused unchanged at every time step, so once trained they are amortized over the whole trajectory; the more time steps, the more favourable the offline–online trade-off becomes. Moreover, the offline stage can be made considerably more efficient than the present per-patch training: one may train a single network that is reused on every patch, or on classes of similar patches, rather than one network per patch. One may also replace the data-driven training by an unsupervised approach that minimizes the patch energy or residual directly, thereby requiring no offline data generation. These directions are left for future work.
All experiments share the fixed fine reference mesh \(h=1/64\). The code uses numpy/scipy except for network training, which uses JAX/optax. The runtime network, including the forward pass and
analytic tangent, is JAX-free.
This appendix gives the implementation-level patch-solved algorithm used for the stationary numerical experiments. It expands Algorithm 2 by including the cross-mesh constraint, patch caches, and reuse of patch factorizations for tangent assembly. The listing is kept non-floating so that it can break across pages.
Algorithm. Tangent-space multiscale manifold method with outer manifold Newton and inner per-patch Newton.
Precompute mesh, operators, and patches Build nested coarse and fine meshes, and the inclusion \(\iota_H\colon V_H\hookrightarrow V_h\) Assemble the cross-mesh mass matrix \(M_{HF}\) representing the \(V_f=\ker\Pi_H\) constraint Assemble the fine load vector \(b_h\) Build the \(\ell\)-ring patch \(\omega_i^\ell\) and the free (non-Dirichlet) node set \(\operatorname{free}_i\) Restrict the constraint matrix to obtain \(B_i\), using the rows of all coarse vertices of \(\overline{\omega_i^\ell}\) interior to \(\Omega\) and the columns of the free fine nodes Store the restricted inclusion, load, and partition-of-unity weights on \(\omega_i^\ell\) Outer Newton iteration on the coarse manifold equation Set \(u_H\gets0\) on the interior coarse degrees of freedom, with boundary values pinned to zero \(u_h\gets\Call{Reconstruct}{u_H}\) \(DM\gets\Call{AssembleTangent}{u_H}\) \(R_h\gets\operatorname{residual}(u_h)-b_h\) \(r_H\gets(DM^\top R_h)|_{\operatorname{int}}\) \(u_H,u_h\) \(A_h'\gets\operatorname{tangent\_matrix}(u_h)\) \(J_H\gets(DM^\top A_h'DM)|_{\operatorname{int}}\) Solve \(J_H\delta u_H=-r_H\) on the interior coarse degrees of freedom Update \(u_H\gets u_H+\delta u_H\) \(u_h\gets\iota_Hu_H\) \(v_{f,i}\gets\Call{PatchEvaluate}{i,u_H}\) Scatter-add \(\varphi_{H,i}v_{f,i}\) to the fine vector on \(\omega_i^\ell\) Project the blended correction with \(I-\Pi_H\), as in 38 \(u_h\) \(g_i\gets\iota_Hu_H|_{\omega_i^\ell}\) Initialize \(v\) from the patch cache or set \(v\gets0\) \(u\gets g_i+v\) \(r\gets(\operatorname{residual}_{\omega_i^\ell}(u)-b_i)|_{\operatorname{free}_i}\) break \(K\gets\operatorname{tangent\_matrix}_{\omega_i^\ell}(u)\) Solve the constrained patch system with block matrix \(\bigl[\begin{smallmatrix}K_{aa}&B_i^\top\\ B_i&0\end{smallmatrix}\bigr]\) Update the free correction degrees of freedom \(v[\operatorname{free}_i]\gets v[\operatorname{free}_i]+dq\) \(K^\ast\gets\operatorname{tangent\_matrix}_{\omega_i^\ell}(g_i+v)\) Factorize the final constrained patch matrix \(S_i\gets\operatorname{Factor}(K^\ast,B_i)\) Cache \((v,K^\ast,S_i)\) for the current coarse state \(u_H\) \(v\) \(DM\gets\iota_H\) Retrieve \((v,K^\ast,S_i)\) from PatchEvaluate Let \(Z_i\) be the coarse basis functions whose support intersects \(\omega_i^\ell\) Solve the linearized patch problem 41 for all directions in \(Z_i\) using the cached factorization \(S_i\) Scatter-add the partition-of-unity weighted tangent blocks to \(DM\) Project the blended tangent correction with \(I-\Pi_H\), as in 40 \(DM\)
Mats G. Larson was supported in part by the Swedish Research Council, Grants Nos. 2021-04925 and 2025-05562, the Knut and Alice Wallenberg Foundation, Grant No. KAW 2025.0277, and the Swedish Research Programme Essence. Anna Persson acknowledges support from the Swedish Research Council Grant no. 2022-03543.
Generative AI tools, including OpenAI’s ChatGPT and Anthropic’s Claude, were used to assist with language editing, LaTeX formatting, figure-caption drafting, reference suggestions, and code generation. All AI-assisted material was critically reviewed, verified, and edited by the authors. The authors take full responsibility for the mathematical content, numerical results, and conclusions of the manuscript.
Authors’ Addresses:
Mats G. Larson, Department of Mathematics and Mathematical Statistics, Umeå University, Sweden
mats.larson@umu.se
Anna Persson, Department of Information Technology, Uppsala University, Sweden
apersson@it.uu.se