Born Discrete, Made Smooth:
Variational Formulation of Shallow Neural Networks

Matej Benko
Institute of Mathematics, Faculty of Mechanical Engineering, Brno University of Technology
Technická \(2896/2\), \(616\,69\), Brno, Czech Republic
Matej.Benko@vutbr.cz Pierre Bousquet
Université de Toulouse, INSA Toulouse, CNRS, IMT, F-31062 Toulouse Cedex 9, France
pierre.bousquet@math.univ-toulouse.fr Iwona Chlebicka
University of Warsaw
ul. Banacha 2, 02-097 Warsaw, Poland
i.chlebicka@mimuw.edu.pl
Błażej Miasojedow
University of Warsaw
ul. Banacha 2, 02-097 Warsaw, Poland
b.miasojedow@mimuw.edu.pl


Abstract

Although neural networks are remarkably effective, their underlying optimization principles remain theoretically elusive, often characterized by non-convex landscapes and stochastic heuristics. In this work, we propose a paradigm shift by replacing the discrete training problem of shallow neural networks with a well-posed continuum variational surrogate. We identify a family of \(\lambda\)-convex functionals over parameter densities in weighted Sobolev spaces and prove that these variational problems are globally well-posed, stable, and exhibit unexpected almost \(C^3\) regularity.

Unlike existing Wasserstein-based or Mean-Field approaches, which often face limited regularity and discretization challenges, our formulation provides direct access to elliptic regularity and convex analysis. This allows us to prove that the optimal parameter density can be obtained by solving a single linear system, bypassing iterative optimization entirely. We establish explicit generalization error controls at a rate of \(1/\alpha\) relative to the regularization parameter, and prove that finite-width networks of size \(N\) achieve the continuum optimum at an \(O(1/N)\) rate. This perspective bridges the gap between the Neural Tangent Kernel (NTK) and feature-learning regimes, providing a principled framework for understanding over-parameterization through the lens of variational calculus.

1 Introduction↩︎

Understanding the surprising empirical effectiveness of gradient-based optimization on highly overparameterized, nonconvex objectives remains one of the central challenges in deep learning theory. Explanation of non-overfitting in such problems is in a short supply. A widely studied approach focuses on representing the network output as an integral over a distribution of hidden units with \[\label{eq-def-fN-intro} f_N(x):=\tfrac1N\textstyle\sum_{i=1}^N w_i h(\theta_i,x)\,, \quad \text{ where \((w_i,\theta_i)_{i=1}^N\) stand for the unknown parameters,}\tag{1}\] and studies the minimization of a risk functional over the space of measures, with gradient descent interpreted as a flow on that space. This mean-field perspective has revealed deep connections to interacting particle systems, optimal transport, and Wasserstein gradient flows [1][7], and has led to convergence guarantees, insights into feature learning, and PDE-based analyses of training dynamics. In parallel, the neural tangent kernel (NTK) framework [8] and related overparameterization results [9], [10] explain why gradient descent succeeds in a linearized infinite-width regime. More recent efforts extend mean-field PDE analyses to deeper, residual architectures [11], sharpen particle and Langevin approximations [12], [13].[14] study existence of minimizers in the finite-width landscape via a closure of the search space; [15] establish nearly optimal approximation rates for shallow \(\mathrm{ReLU}^k\) networks on Sobolev classes. Both works operate at finite networks, whereas we concentrate on the regularity of the continuum minimizer. Finally, [16] add a Sobolev penalty on the network output over the input domain \(D\) to improve optimization conditioning. In contrast, our regularization acts on the parameter density over \(\Omega\) to enforce variational well-posedness and regularity. This makes the two approaches complementary yet structurally distinct.

Despite this progress, two principal frameworks dominate the field, each with characteristic trade-offs. Mean-field/Wasserstein approaches support a particle interpretation of gradient descent but offer only displacement convexity and require non-trivial approximation to pass to finite-width models. The NTK framework achieves a truly convex objective but only by linearizing around initialization, limiting guarantees to the lazy-training regime. Both leave partly open a different set of questions about the variational problem itself: Does the infinite-width objective admit a unique, stable minimizer? What regularity do optimal parameter distributions possess? Can one obtain global convexity for a nonlinear continuum model without linearization?

In this paper we propose a third route that addresses these questions directly. We formulate shallow neural network training as a variational problem over parameter densities in a weighted Sobolev space \(\mathcal{W} = W^{1,2}(\Omega)\cap L^2_\omega(\Omega)\) forming a linear Hilbert-space setting that retains the infinite-width mean-field spirit while gaining direct access to convex analysis, elliptic PDE theory, and quantitative gradient flow estimates. The regularized objective on \({\mathcal{W}}\) which is given by \[{\mathscr{F}}^{(f)}_{\alpha,\beta}(u) = \mathcal{R}(f, u\,\mathrm{d}\theta) + \alpha\|u\|_{L^2_\omega}^2 + \beta\|\nabla u\|_{L^2}^2\] combines a squared-risk term measuring data fidelity with \(L^2_\omega\) and Sobolev regularization. The \(L^2_\omega\) term induces coercivity and \(\lambda\)-convexity for \(\lambda>0\); the Sobolev term promotes smoothness and plays a role analogous to entropy regularization in the well-established, yet computationally hard, Wasserstein formulations [17][19]. In turn, the minimizer \(u^*_f\) to \({\mathscr{F}}^{(f)}_{\alpha,\beta}\) can be interpreted as the implicit target of NN in the overparameterized regime, while the regularization parameters \(\alpha,\beta\) play the role of weight decay and smoothness bias.

Our Hilbert-space framework is built on the classical calculus of variations staying fully compatible with standard shallow network parametrizations. Although our analysis does not directly model stochastic gradient descent, the exponential convergence rate \(e^{-\lambda t}\) of the gradient flow associated to our variational problem and the structure of the Euler–Lagrange equation suggest that the \(L^2_\omega\) geometry may serve as a useful proxy for understanding the implicit bias in overparameterized networks. In particular, the \(\lambda\)-convex structure implies that any discretization of the gradient flow inherits stability consistent with the empirical robustness of SGD near overparameterized minima [1], [2]. In our study, the minimization of \({\mathscr{F}}_{\alpha,\beta}^{(f)}\) can be reduced to the optimization of a quadratic functional depending on the coordinates of the competitors in a chosen basis of \(L^2_\omega(\Omega)\). Since the projection of the minimizer onto a finite dimensional subspace of \(L^2_\omega(\Omega)\) can be obtained by merely solving a linear system of equations, the optimization is a fast and easy procedure, as it does not require any iteration or gradient descent arguments.

The variational formulation we propose yields a cluster of remarkable properties.

No Lavrentiev gap. A key consistency result shows that the infimum of the risk is the same whether one optimizes over atomic measures (finite-width networks), Sobolev densities, or smooth compactly supported functions. No approximation error is introduced by passing to the continuum, and finite-width networks of width \(N\) achieve the continuum optimum up to an \(O(1/N)\) error. This places the continuum theory as an exact relaxation of the finite discrete training problem. The absence of Lavrentiev’s phenomenon is a typical key step in inferring regularity of minimizers [20][23].

\(\lambda\)-convexity. The functional \({\mathscr{F}}_{\alpha,\beta}^{(f)}\) is \(\lambda\)-convex on \(L^2_\omega\) with \(\lambda = 2\alpha\geq 0\), which gives the exponential convergence of the gradient flow, and it is \(2\min(\alpha,\beta)\)-convex on \({\mathcal{W}}\), which leads to the existence and the uniqueness of the minimizer, as well as its stability with respect to the data. These are global properties of the model, not consequences of linearization, equipping with strong tools, which are typically not available in the Wasserstein approach, [24]. The \(L^2_\omega\)-gradient flow of \(\mathcal{F}^{(f)}_{\alpha,\beta}\), the continuous-time analogue of gradient descent, converges exponentially fast to \(u^*\), cf. [25]. A time-discrete implicit Euler scheme provides an approximation with explicit error, combining the exponential rate with a \(\sqrt{\tau}\) term, where \(\tau\) is the discretization step.

Near-\(C^3\) smoothness. The minimizer \(u^*\) of \({\mathscr{F}}_{\alpha,\beta}^{(f)}\) solves an explicit Euler–Lagrange equation. Moreover, it is \(C^{2,s}\) for every \(s<1\), which gives a major structural advantage: it prevents pathological concentrations of mass in parameter space and opens the door to higher-order numerical schemes via classical PDE methods. In contrast, in the Wasserstein approach, regularity of minimizers is inherited from the activation function, whereas here we get surprising smoothness even in the case of merely locally Lipschitz activation functions, see [1]. This suggests that the predictions of the neural network concentrate around graphs of almost \(C^3\)-functions, which illustrates the non-overfitting phenomenon, robustness of training, and suggests that the neural networks are attracted to the lower-dimensional manifolds, cf. [26][28]. The stability results we infer provide a concrete form of implicit bias toward smooth solutions and explain why the variational formulation yields stable and well-generalizing models.

Contributions. Our framework provides an entirely variational formulation for infinite-width shallow networks. Replacing the discrete training process with a regularized surrogate, we identify a class of globally well-posed and stable objectives. This approach follows three steps: (i) transforming the \(N\)-neuron optimization into a continuum problem over weighted Sobolev spaces; (ii) exploiting the structure to ensure a unique and almost \(C^3\) regular minimizer; (iii) bridging to finite-width networks via atomic measures with \(O(1/N)\) consistency. This offers novel insight into the implicit bias of neural networks, as detailed below.

  • A variational theory of shallow network training. We formulate shallow neural network training as a Sobolev-space variational problem (Section 2) and concentrate on the properties of solutions. We prove that this continuum formulation is exact: there is no Lavrentiev gap, so atomic, Sobolev, and smooth minimization classes coincide (Theorem 1). We show that this infinite-width problem is quantitatively faithful to finite networks via an \(O(1/N)\) approximation bound (Proposition 2).

  • Global convexity and well-posedness beyond existing frameworks. We introduce a regularized functional \(\mathcal{F}_{\alpha,\beta}^{(f)}\) that is globally \(2\min(\alpha, \beta)\)-convex on \({\mathcal{W}}\) yielding existence, uniqueness, and stability of minimizers. We establish that the associated \(L^2_\omega\) gradient flow converges exponentially fast to equilibrium at rate \(e^{-2\alpha t}\) (Corollary 3). In particular, this provides an explicit analytical mechanism for non-overfitting in the infinite-width regime.

  • Implications to ML. The minimizers to the regularized functional admit an Euler–Lagrange characterization and exhibit near-\(C^3\) regularity via elliptic PDE arguments (Theorem 3), a level of smoothness not accessible in standard infinite-width analyses. Propositions 4 and 5, together with Corollary 2, show quantitative dependence of the minimizer under perturbations providing a link between regularization, stability, and generalization.

Shallowness. Despite the dominance of deep architectures, the single-hidden-layer setting already captures the core variational complexity of neural network training: a nonlinear, non-convex objective over an infinite-dimensional space. Resolving these obstructions in this setting isolates the essential structure and provides a foundation for multilayer extensions [11].

Limitations. The current formulation is for shallow (one-hidden-layer) networks. Moving from one to more layers complicates the structure of the optimization problem, which becomes strongly nonlinear, see Section 3. In fact, for two hidden layers the parameter density lives on \(\Omega_1 \times \Omega_2\) and the risk functional becomes a composition, making \({\mathscr{R}}\) possibly nonconvex even after Sobolev regularization. The Euler–Lagrange system becomes a coupled nonlinear PDE rather than a linear equation, so existence of smooth minimizers requires additional structural assumptions such as small initial condition or a perturbative regime around a known solution. The gradient flow no longer reduces to a linear semigroup and convergence is elusive.

Summary. Our approach is related to the mean-field literature [1][3], [12], [13], but operates in a different analytical regime. The connection to Sobolev-type objectives recently explored in [16] is suggestive, though the focus there is on conditioning of the optimization problem rather than on variational well-posedness or regularity of minimizers. Our results also complement convex and function-space formulations of infinite-width networks [29], [30], which concentrate on description of the learning process; we focus instead on the variational structure of the objective and the regularity of its minimizers. Finally, while the NTK framework [8] achieves convexity through linearization and is well-suited to the lazy-training regime [9], [10], our \(\lambda\)-convexity holds for the same nonlinear problem that arises in the infinite-width limit of shallow network training, without any linearization assumption.

2 Variational formulation↩︎

We develop a continuum approach to shallow networks based on parameter distributions.

2.1 Exact variational characterization of neural network training↩︎

Let \(D\subset {\mathbb{R}}^d\) be the input (a bounded Borel set equipped with the Lebesgue measure), let \(f\in L^2(D)\) and let \(\Omega\subset {\mathbb{R}}^{d+1}\) be the parameter space of hidden units, assumed to be open. Let \(\theta=(\theta_0,\theta')\in\Omega\) parametrize the hidden layer with \(\theta_0\in {\mathbb{R}}\) being interpreted as the bias, \(\sigma\) be an activation function that we assume to be continuous with the typical choices of ReLU or sigmoid functions. We also consider the function \(h: \Omega\times D\to {\mathbb{R}}\) given by \[\label{eq:h-def}h(\theta,x):=\sigma(\theta_0+\theta'\cdot x)\,\tag{2}\] and we require that there exists \(C_h>0\) such that for every \((\theta,x)\in \Omega\times D\), \[\label{eq:h-ass} |h(\theta,x)|\leq C_h(1+|x|)(1+|\theta|)\,.\tag{3}\] This condition holds true in the typical choices of ReLU and sigmoid functions. In this framework training a shallow neural network consists in approximating \(f\) with functions \(f_N\) that are given by 1 .

The learning process amounts to the minimization of \(\int_D (f(x)-f_N(x))^2{\,{\rm d}x}\) over \((w_i,\theta_i)_{i=1}^N\). Each function \(f_N\) can be rewritten as \[\label{eq:fN-muN} f_N(x)=\int_\Omega h(\theta,x){\,{\rm d}\mu}_N(\theta)\quad\text{for }\;\mu_N = \frac{1}{N} \sum_{i=1}^N w_i \delta_{\theta_i}\in{{\mathscr{M}}^{\rm at}}(\Omega)\,.\tag{4}\] Observe that \(\mu_N\) is a finite atomic signed measure (with finite number of atoms). Denote by \({\mathscr{M}}_a(\Omega)\) the set of those finite signed Borel measures \(m\) such that \(\int_{\Omega}|\theta|^a{\,{\rm d}}|m|(\theta)<\infty.\) Then, we generalize 4 for \(m\in{\mathscr{M}}_1(\Omega)\) by introducing the function \[\label{eq:fm} f_m(x)=\int_\Omega h(\theta,x){\,{\rm d}}m(\theta)\,.\tag{5}\] In view of 3 and the fact that \(m\in {\mathscr{M}}_1(\Omega)\), the function \(f_m\) is well-defined and belongs to \(L^{2}(D)\). We can thus consider the population risk functional \[\label{eq:risk} {\mathscr{R}}(f,m):=\int_D (f(x)-f_m(x))^2{\,{\rm d}x}\,.\tag{6}\] The training problem becomes the minimization of \({\mathscr{R}}(f,m)\) over \(m\in{\mathscr{M}}_1(\Omega)\), as suggested in [3]. To obtain a Hilbert-space formulation, we restrict attention to measures that admit a density, writing \({\,{\rm d}}m(\theta)=u(\theta){\,{\rm d}\theta}\) for some Sobolev function \(u\). Since \(\Omega\) is typically unbounded, we work in a weighted space that controls integrability, so that all quantities under consideration are well-posed. Let \(\omega:\Omega\to[1,\infty)\) be a smooth weight such that \[\label{eq:weight-assumptions} c_\omega:=\textstyle\int_\Omega {(1+|\theta|)^4}\tfrac{1}{\omega(\theta)}{\,{\rm d}\theta}<\infty\,.\tag{7}\] Moreover, we define \[L^2_\omega(\Omega) = \left\{ u:\textstyle\int_\Omega u(\theta)^2\omega(\theta){\,{\rm d}\theta}<\infty \right\} \qquad\text{and}\qquad {\mathcal{W}}:= W^{1,2}(\Omega)\cap L^2_\omega(\Omega)\,.\] Condition 7 and the Schwarz inequality imply that for every \(v\in L^{2}_\omega(\Omega)\), the measure \(v{\,{\rm d}x}\) belongs to \({\mathscr{M}}_1(\Omega)\). Our main result establishes a striking equivalence: neural network training admits an exact variational characterization. In particular, the continuum formulation is not merely an approximation, but an exact relaxation of the finite discrete problem, which reveals its intrinsic variational structure.

Theorem 1. Under the conditions of this section, for every \(f\in L^2(D)\), one has: \[\label{eq:R-no-Lavr} \inf_{\mu \in {{\mathscr{M}}^{\rm at}}(\Omega)} {\mathscr{R}}(f,\mu)= \inf_{m \in {\mathscr{M}}_1 (\Omega)} {\mathscr{R}}(f,m) = \inf_{v \in {\mathcal{W}}} {\mathscr{R}}(f,v{\,{\rm d}x})= \inf_{v \in C_c^\infty(\Omega)} {\mathscr{R}}(f,v{\,{\rm d}x})\,.\tag{8}\]

From a variational perspective, this shows that no Lavrentiev gap occurs for \({\mathscr{R}}(f,\cdot)\) across measures, Sobolev maps, and smooth functions. In particular, minimizing sequences can be transferred between these classes without loss of optimality, which is a fundamental structural property for regularity theory. This bridges the discrete and continuum regimes at a deeper level, ensuring that analytical and numerical approximations faithfully capture the true variational behavior of the problem, and places neural network training firmly within the scope of variational analysis.

Let us point out the rate of convergence of the risk with respect to the width of the neural network in the spirit of [1]. Observe that \(L^{2}_\omega(\Omega)\) is a natural space to consider \({\mathscr{R}}\) and further regularity results will be provided for the problem on the right-hand side below. In the next proposition, we denote by \({\mathscr{M}}^{\rm at}_N(\Omega)\) the set of those purely atomic signed measures on \(\Omega\) that involve at most \(N\) atoms. Observe that such measures can be written as in 4 for suitable coefficients \(w_i\).

Proposition 2. Let \({\mathscr{R}}\) be given by 6 . Then, there exists \(C>0\) depending only on \(D, h, \omega\) such that for every \(f\in L^{2}(D)\), for every \(N\geq 2\), \[\inf_{\mu\in {\mathscr{M}}_{N}^{\rm at}(\Omega)}{\mathscr{R}}(f,\mu) \leq \inf_{v\in L^{2}_\omega(\Omega)} \left({\mathscr{R}}(f,v{\,{\rm d}x})+ \tfrac{C}{N}\|v\|_{L^{2}_\omega (\Omega)}^2\right)\,.\] As a consequence, if there exists a bounded minimizing sequence for \({\mathscr{R}}(f,\cdot)\) in \(L^{2}_\omega(\Omega)\), then \[\inf_{\mu\in {\mathscr{M}}_{N}^{\rm at}(\Omega)}{\mathscr{R}}(f,\mu) \leq \inf_{v\in L^{2}_\omega(\Omega)}{\mathscr{R}}(f,v{\,{\rm d}x})+O\left(\tfrac{1}{N}\right) \,.\]

2.2 Towards justification of non-overfitting: stability of training↩︎

By expanding the risk term and using the Fubini theorem, we obtain for \(f\in L^{2}(D)\) and \(m\in {\mathscr{M}}_1(\Omega)\), that \[\label{eq:quadratic-decomposition} {\mathscr{R}}(f,m) = \|f\|_{L^2(D)}^2 -2\int_\Omega Q(\theta){\,{\rm d}}m(\theta) +\int_\Omega\int_\Omega K(\theta,{\vartheta}){\,{\rm d}}m(\theta){\,{\rm d}}m({\vartheta})\,,\tag{9}\] where \[\label{eq:bK} Q(\theta):=\int_D f(x)h(\theta,x){\,{\rm d}x}\qquad\text{and}\qquad K(\theta,{\vartheta}):=\int_D h(\theta,x)h({\vartheta},x){\,{\rm d}x}\,.\tag{10}\] In view of 3 for every \(\theta, {\vartheta}\in \Omega\), for \(C_{h}':=C_{h}\|1+|x|\|_{L^{2}(D)}\), one has: \[\label{def-CKCQ} |Q(\theta)|\leq C_{h}\,(1+|\theta|)\int_{D}(1+|x|)|f(x)|{\,{\rm d}x}\qquad\text{and} \qquad |K(\theta,{\vartheta})|\leq (C_{h}')^2\,(1+|\theta|)\,(1+| {\vartheta}|)\,.\tag{11}\] The functional \({\mathscr{R}}(f,\cdot)\) is convex on \({\mathscr{M}}_1(\Omega)\) (see Lemma 3 below) but not necessarily strictly convex. We endow the Hilbert space \({\mathcal{W}}\) with the norm \(\|u\|_{{\mathcal{W}}}^2 := \|u\|_{L^{2}_\omega(\Omega)}^2 + \|\nabla u\|_{L^{2}(\Omega)}^2\). We are thus led to introduce for every \(\alpha, \beta\geq 0\) the natural regularized objective \[\label{eq:regularized-functional} {\mathscr{F}}^{(f)}_{\alpha,\beta}(u) := \begin{cases} {\mathscr{R}}(f,u {\,{\rm d}x}) +\alpha\|u\|_{L^2_\omega}^2 +\beta\|\nabla u\|_{L^2}^2\,, & u\in \mathcal{W}\,,\\ +\infty\,, & u\in L^{2}_\omega(\Omega)\setminus {\mathcal{W}}\,. \end{cases}\tag{12}\] The objective combines a term measuring data fidelity with Sobolev and \(L^2_\omega\) regularization terms that provide coercivity and enforce regularity in the parameter space. In particular, the Sobolev term \(\|\nabla u\|_{L^2}^2\) induces a diffusive effect in the associated gradient flow, analogous to entropy regularization in Wasserstein formulations well-designed for the study of the learning process, cf. [17][19]. We avoid it as, in contrast to the computationally challenging Wasserstein approach, the \(L^2_\omega\) setting is naturally compatible with the study of the solutions to standard shallow neural network parametrizations. This enables the use of convex-analytic techniques together with classical regularity theory.

We notice that, when restricted to its domain \({\mathcal{W}}\), the regularized functional from 12 is Fréchet differentiable and \(2\min(\alpha, \beta)\)-convex, i.e., the functional \(u\mapsto{\mathscr{F}}^{(f)}_{\alpha,\beta}(u)-\min(\alpha, \beta)\|u\|_{{\mathcal{W}}}^2\) is convex on \({\mathcal{W}}\). Moreover, on \(L^{2}_\omega(\Omega)\), the functional \({\mathscr{F}}^{(f)}_{\alpha,\beta}\) is \(2\alpha\)-convex.

Some of our results require to assume that \(h\) is locally Lipschitz with respect to the first variable, in the sense that for every \(R>0\), there exists \(c_{h,R}>0\) such that \[\label{eq-Phi-loc-Lip} |h(\theta, x) - h({\vartheta},x)|\leq c_{h,R}|\theta-{\vartheta}|(1+|x|)\quad \text{for every \(\theta, {\vartheta}\in \Omega\cap B_R\), for every \(x\in D\)}\,.\tag{13}\] The regularized functional enjoys strong properties due to the convex structure of \({\mathscr{F}}^{(f)}_{\alpha,\beta}\). It ensures global existence and uniqueness of the minimizer \(u_f^*\) together with an explicit Euler–Lagrange PDE. The almost \(C^{3}\) regularity of \(u_f^*\) is a major advantage: it lifts the optimal density far beyond generic solutions, opening the way to prove sharp asymptotic analysis, stability estimates, and a justification of particle approximation tools that are essential yet scarce in mean-field neural network theory.

Theorem 3. Under the conditions of this section, let \(\alpha,\beta>0\) and \(h\) be a continuous function satisfying 3 and 13 . Then \({\mathscr{F}}^{(f)}_{\alpha,\beta}\) is \(2\min(\alpha, \beta)\)-convex on \({\mathcal{W}}\) and \(2\alpha\)-convex on \(L^{2}_\omega(\Omega)\). Moreover, it admits a unique minimizer \(u^*\in\mathcal{W}\), which is a classical solution to the Euler–Lagrange equation: \[\label{eq:euler-lagrange} -\beta \Delta u^*(\theta) + \alpha \omega(\theta) u^*(\theta) -Q(\theta)+\int_{\Omega}K(\theta,{\vartheta})u^*({\vartheta}){\,{\rm d}}{\vartheta} = 0\,,\tag{14}\] and it holds that \(u^* \in C^{2,s}_{\mathrm{loc}}(\Omega)\) for every \(s\in(0,1)\).

We stress that due to convexity, our variational problem is stable with respect to the target.

Proposition 4. Under the conditions of this section, let \(\alpha, \beta>0\) and \(h\) be a continuous function satisfying 3 . Then, the unique minimizer \(u_f^*\in {\mathcal{W}}\) of \({\mathscr{F}}_{\alpha, \beta}^{(f)}\) depends linearly on \(f\) and there exists \(C=C(\Omega,D,\omega,h)>0\) such that \[\|u_f^*\|_{L^{2}_\omega(\Omega)} \leq \tfrac{C}{\alpha} \|f\|_{L^{2}(D)}\, \quad \text{and} \quad \|\nabla u_f^*\|_{L^{2}(\Omega)} \leq \tfrac{C}{\sqrt{\alpha\beta}} \|f\|_{L^{2}(D)}\,.\]

As a consequence, given \(f_1, f_2\in L^{2}(D)\), the corresponding minimizers \(u_{f_1}^*, u_{f_2}^{*}\) satisfy \[\|u_{f_1}^* - u_{f_2}^*\|_{L^2_\omega(\Omega)} \le \tfrac{C}{\alpha} \|f_1 - f_2\|_{L^2(D)} \, \quad \text{and} \quad \|\nabla u_{f_1}^*-\nabla u_{f_2}^*\|_{L^{2}(\Omega)} \leq \tfrac{C}{\sqrt{\alpha\beta}} \|f_1-f_2\|_{L^{2}(D)}\,.\] The continuous dependence of the minimizers with respect to the targets \(f\) entails a similar property for the minimal values of the functionals \({\mathscr{F}}^{(f)}_{\alpha, \beta}\):

Corollary 1. Under the conditions of this section, let \(\alpha, \beta>0\) and \(h\) be a continuous function satisfying 3 . Given \(f_1, f_2\in L^{2}(D)\), let \(u_{f_1}^*\) and \(u_{f_2}^*\) be the minimizers of \({\mathscr{F}}_{\alpha, \beta}^{(f_1)}\) and \({\mathscr{F}}_{\alpha, \beta}^{(f_2)}\). Then, there exists \(C=C(\Omega,D,\omega,h)>0\) such that \[|{\mathscr{F}}^{(f_1)}_{\alpha, \beta}(u_{f_1}^*)-{\mathscr{F}}^{(f_2)}_{\alpha, \beta}(u_{f_2}^*)|\leq \tfrac{C}{\alpha}(\|f_1\|_{L^{2}(D)}+\|f_2\|_{L^{2}(D)})\|f_1-f_2\|_{L^{2}(D)}\,.\]

Let us also emphasize the local uniform stability property. We say that \(u \in C^2(\overline{\Omega})\) if \(u,\; \nabla u,\; D^2 u\) admit continuous extensions to \(\overline{\Omega}\) and \(\|u\|_{C^2(\overline{\Omega})} := \|u\|_{L^\infty(\Omega)} + \|\nabla u\|_{L^\infty(\Omega)} + \|D^2u\|_{L^\infty(\Omega)}<\infty.\)

Proposition 5. Under the conditions of this section, let \(\alpha,\beta>0\) and \(h\) be a Borel function satisfying 3 and 13 . For \(f\in L^{2}(D)\), let \(u_{f}^*\in {\mathcal{W}}\) be the unique minimizer of \({\mathscr{F}}_{\alpha, \beta}^{(f)}\). Then, for every \(\Omega' \Subset \Omega\), there exists \(C=C(\Omega', \Omega, \alpha, \beta, D, \omega, h)>0\) such that \(\|u_{f}^*\|_{C^{2}(\overline{\Omega'})}\leq C\|f\|_{L^{2}(D)}\).

It follows that for every \(f_1, f_2\in L^{2}(D)\) and every \(\Omega'\Subset \Omega\), the corresponding minimizers \(u^*_{f_1}\) and \(u^*_{f_2}\) satisfy \[\|u^*_{f_1}-u^*_{f_2}\|_{C^{2}(\overline{\Omega'})} \leq C \|f_1-f_2\|_{L^{2}(D)}\,.\]

As a consequence of Proposition 4, Proposition 5 and the Fubini theorem, we get:

Corollary 2. Under the conditions of this section, let \(f \in L^2(D)\), \(\varsigma>0\), and let \(f_{\varepsilon}= f + \varepsilon\), where \(\varepsilon\) is a mean-zero random error with \(\mathbb{E}\big[\|\varepsilon\|_{L^2(D)}^2\big] \le \varsigma^2\). If \(u_{f}^*,u_{f_{\varepsilon}}^* \in {\mathcal{W}}\) are the unique minimizers of \({\mathscr{F}}_{\alpha,\beta}^{(f)}\) and \({\mathscr{F}}_{\alpha,\beta}^{(f_{\varepsilon})}\), then \(\mathbb{E}\big[\|u_{f_{\varepsilon}}^* - u_f^*\|_{L^2_\omega(\Omega)}^2\big] \le \tfrac{c^2}{\alpha^2}\,\varsigma^2\) for some \(c=c(\Omega,D,\omega,h)>0\). Moreover, for every \(\Omega' \Subset \Omega\), there exists \(C = C(\Omega', \Omega,\alpha,\beta,D,\omega,h)>0\) such that \[\mathbb{E}\Big[ \|u_{f_{\varepsilon}}^* - u_f^*\|_{C^{2}(\overline{\Omega'})}^2 \Big] \le C^2 \varsigma^2.\]

Propositions 4 and 5, together with Corollary 2, establish Lipschitz dependence of the minimizer on the target, with generalization error scaling linearly in the noise level \(\varsigma\) and controlled by \(1/\alpha\). The near-\(C^3\) regularity ensures that the local geometry of the learned representation varies smoothly with the data, providing a concrete form of implicit bias toward smooth, well-generalizing solutions.

In order to describe the target of the neural network related to \({\mathscr{F}}_{\alpha,\beta}^{(f)}\), we consider the gradient flow of this functional in the Hilbert space \(L^2_\omega(\Omega)\), following [25]. This approach is natural in the analysis of overparametrized neural networks because the gradient flow corresponds to the continuous-time limit of gradient descent (the steepest descent curve of \({\mathscr{F}}^{(f)}_{\alpha,\beta}\) in the \(L^2_\omega\)-geometry). It provides a powerful analytic framework for our novel study of convergence, implicit bias, and generalization properties that are often difficult to capture in discrete time.

Since \({\mathscr{F}}^{(f)}_{\alpha, \beta}\) is quadratic, the existence of the gradient flow follows from the Hille–Yosida theory. We introduce the unbounded linear operator \(A:D(A)\subset L^{2}_\omega(\Omega)\to L^{2}_\omega(\Omega)\) defined by its domain \[D(A):=\left\lbrace u\in {\mathcal{W}}: \exists p=p(u)\in L^{2}_{\omega}(\Omega) \textrm{ such that } \forall v\in {\mathcal{W}}, \int_{\Omega}\nabla u \cdot \nabla v {\,{\rm d}\theta}= -\int_{\Omega}pv\omega{\,{\rm d}\theta}\right\rbrace\,\] and \(A u:=2\alpha u -2\beta p(u)+\frac{2}{\omega}\int_{\Omega}K(\theta, \cdot)u(\theta){\,{\rm d}\theta}\). Note that here the function \(p(u)\) is \(\Delta u/\omega\).

It follows from the Hille–Yosida theorem complemented by the Brézis–Komura theorem that:

Proposition 6. Under the conditions of this section, for every \(u_0\in L^{2}_{\omega}(\Omega)\), there exists a unique \(u\in C^{0}([0,\infty[;L^{2}_\omega(\Omega))\cap C^1((0,\infty);L^{2}_\omega(\Omega))\cap C^0((0,\infty);D(A))\) such that \[\label{eq13070} \begin{cases} \frac{du}{dt}= -A u+2\frac{Q}{\omega}\qquad \forall t>0\,,\\ u(0)=u_0\,. \end{cases}\qquad{(1)}\] Moreover, there exists a continuous semigroup of contractions \((S_t)_{t\geq 0}\) such that \(u(t)=S_t u(0)\), and \[\forall v_1, v_2\in L^{2}_\omega(\Omega), \qquad \|S_t v_1 -S_t v_2\|_{L^{2}_\omega(\Omega)} \leq e^{-2\alpha t}\|v_1-v_2\|_{L^{2}_\omega(\Omega)}\,.\] Finally, \(t\mapsto \|u'(t)\|_{L^{2}_\omega(\Omega)}\) is nonincreasing on \((0,\infty)\) and \[{\mathscr{F}}^{(f)}_{\alpha, \beta}(u(t))\leq \inf_{v\in {\mathcal{W}}}\left({\mathscr{F}}^{(f)}_{\alpha, \beta}(v)+\frac{\alpha}{e^{2\alpha t}-1} \|u(0)-v\|_{L^{2}_\omega(\Omega)}\right)\,.\]

Corollary 3. Under the conditions of this section, let \(v\in L^{2}_\omega(\Omega)\) and \(u_f^*\) be a minimizer of \({\mathscr{F}}^{(f)}_{\alpha,\beta}\). Then \(\|S_t v - u_f^*\|_{L^2_\omega} \le e^{-2\alpha t} \|v - u_f^*\|_{L^2_\omega}\).

In the case of shallow networks, one can approximate the minimizer to \({\mathscr{F}}^{(f)}_{\alpha,\beta}\) by considering the restriction of the functional to a finite dimensional space \(E_\ell := \mathrm{span}\{e_0,\dots,e_\ell\} \subset \mathcal{W}\), where \(\{e_i\}_{i\ge 0}\) forms a complete basis of the Hilbert space \(L^2_\omega(\Omega)\). The problem thus reduces to solving a linear system, see Section 3. In order to pass to the multi-layer case, the problem starts to be strongly nonlinear and analytically infeasible. Hence, another learning algorithm is necessary. One of the ideas could be to adapt the implicit Euler scheme associated to the gradient flow, see e.g. [25]. Given \(\tau>0\) and an initial condition \(u\in {\mathcal{W}}\), we define \[\label{eq:jko} \tilde{\varrho}^0 = u, \qquad \tilde{\varrho}^{k\tau} = \arg\min_{v \in L^2_\omega(\Omega)} \left\{ {\mathscr{F}}_{\alpha,\beta}^{(f)}(v) + \tfrac{1}{2\tau} \|v - \tilde{\varrho}^{(k-1)\tau}\|_{L^2_\omega}^2 \right\}\,.\tag{15}\] The scheme admits minimizers due to the convexity and the lower semicontinuity of \({\mathscr{F}}_{\alpha,\beta}^{(f)}\). Defining the piecewise constant interpolation \(\tilde{\varrho}_\tau(t)\) by 37 , one obtains convergence toward the gradient flow as \(\tau \to 0\), cf. Theorem 11. Combining this with Corollary 3 yields the error estimate:

Corollary 4. Under the conditions of this section, let \(u \in \mathcal{W}\) and \(u_{f}^*\) be the minimizer of \({\mathscr{F}}^{(f)}_{\alpha,\beta}\). Then, for any \(\tau>0\) and \(t\geq 0\), \[\|\tilde{\varrho}_\tau(t) - u_{f}^*\|_{L^2_\omega} \le e^{-2\alpha t}\|u - u_{f}^*\|_{L^2_\omega} + 2(\sqrt{2}+1) \sqrt{\tau \, {\mathscr{F}}^{(f)}_{\alpha,\beta}(u)}\,.\]

3 Numerical experiments↩︎

The experiments below illustrate the key properties established in the theoretical analysis: no Lavrentiev gap, \(\lambda\)-convexity, regularity of the optimal parameter density. We note that our numerical findings show that estimations of \({\mathscr{F}}^{(f)}_{\alpha,\beta}\) are both accurate and robust to noise and outliers. All simulations are done on standard personal computer and can be found in the supplementary material. Reduction to a linear system. We approximate the parameter density \(u \in {\mathcal{W}}\) by a polynomial ansatz \(u(\theta) \approx \sum_{i=1}^M a_i\,\widehat{u}_i(\theta).\) We have used three different types of basis functions \(\widehat{u}_i\): polynomials, orthonormal cosine basis and Legendre polynomials. We minimize \({\mathscr{F}}_{\alpha,\beta}^{(f)}\) over the coefficient vector \(\vec{a}\in{\mathbb{R}}^M\). Since the functional \({\mathscr{F}}_{\alpha,\beta}^{(f)}\) involves expectations with respect to the (unknown) data distribution, it cannot be evaluated directly. We therefore work with its empirical version, obtained by replacing the population quantities with averages over the observed data. Given data points \((x_k, f(x_k))_{k=1}^{N_D}\) , the functional reduces to the sum of a quadratic form and an affine term: \[\widehat{{\mathscr{F}}}_{\alpha,\beta}^{\;(f)}(a) = C_D\left( \vec{f}\,^\top \vec{f} - 2\vec{f}\,^\top U\vec{a} + \vec{a}^\top U^\top U \,\vec{a}\right)+{\vec{a}}^\top(\alpha V + \beta W)\,\vec{a}\,,\] where \(C_D=\frac{\mathcal{L}^d(D)}{N_D}\) (here, \(\mathcal{L}^d(D)\) is the Lebesgue measure of \(D\) ), \(\vec{f} = (f(x_1),\ldots,f(x_{N_D}))^\top\), \(\vec{a}=(a_1,\ldots,a_M)^\top\), and the matrices \[U_{ki} = \int_\Omega h(\theta, x_k)\,\widehat{u}_i(\theta){\,{\rm d}}\theta, \quad V_{ij} = \langle \widehat{u}_i , \widehat u_j \rangle_{L^2_\omega (\Omega)}, \quad W_{ij} = \langle \nabla \widehat{u}_i, \nabla \widehat{u}_j \rangle _{L^2 (\Omega)}\] are determined by the given data, the parameter domain \(\Omega\), and the weight \(\omega(\theta)=1+|\theta|^{2d+4}\). Crucially, no iterative optimization or gradient descent is required. We can find the minimizer directly by solving the linear equation: \[\nabla \widehat{{\mathscr{F}}}^{\;(f)}_{\alpha, \beta} (\vec{a}) = 0 \implies \vec{a} = \left(U^\top U +\frac{\alpha}{C_D} V+\frac{\beta}{C_D} W\right)^{-1}\left(U^\top \vec{f}\right) \, .\] The matrix \(U^\top U + \frac{\alpha}{C_D} V + \frac{\beta}{C_D} W\) is positive semidefinite (and moreover positive definite for \(\alpha>0\) in the case the basis functions \(\{\widehat{u}_i\}_{i=1}^{M}\) are linearly independent in \(L^2_\omega(\Omega)\) or \(\beta>0\) and the gradients of the basis functions are linearly independent in \(L^{2}(\Omega)\)) and the minimizer is found via pseudoinverse (or inverse); the problem reduces to ridge regression.

Let us compare: (i) the unregularized case \(\alpha=\beta=0\), which corresponds to the least-squares projection onto the polynomial ansatz; and (ii) a single-hidden-layer network \(f_N\) with \(N=10{\,}000\) ReLU neurons trained by Adam optimizer [31], representing the finite-width benchmark. In Example 4 NN was trained via SGD due to the instability of Adam.

3.0.0.1 Example 1: Sinus function for \(\boldsymbol{d=1}\).

We generate \(N_D = 50\) observations of the function \(\sin(7x)\) plus small Gaussian noise on the interval \((-1,1)\). The minimum of \({\mathscr{F}}^{(f)}_{\alpha, \beta}\) is approximated by a polynomial 38 of order \(s=15\), thus \(M = 136\). We have chosen \(\Omega=(-R,R)\times (-L,L)\) with \(R=L=7\). See the results on Figure 1 (a) and compare the prediction for the single-layer neural network with \(N=10\,000\) neurons (NN), our prediction for \(\alpha=\beta=0\) and then with small penalties \(\alpha=8.8\times10^{-12}\), \(\beta=8.8\times 10^{-10}\).

a
b

Figure 1: Prediction for functions \(\mathrm{sin}\) and perturbed \(\mathrm{sign}\) in \(d=1\). a — Example 1, b — Example 2

Figure 1 (a) shows that the regularized solution tracks the target accurately and smoothly, while the unregularized ansatz overfits and the neural network baseline is noisier. The Sobolev penalty makes the optimal density far from oscillatory, producing a function concentrated near a smooth manifold reflecting near-\(C^3\) regularity of Theorem 3 and the implicit bias toward smooth solutions. In addition, surprisingly, the regularized version is able to recover almost fully the next period of the \(\sin\) function.

3.0.0.2 Example 2: Discontinuous target — sign function (\(d=1\))

We generate \(N_D = 50\) observations of \(\mathrm{sign}(x)\) on \((-1,1)\) with small Gaussian noise, including one outlier, and use trigonometric polynomials \(\{\widehat{u}_i\}_{i=1}^{M}\) 39 up to the coefficient \(s=60\) (\(M=1{,}891\)). We have chosen \(\Omega=(-R,R)\times (-L,L)\) with \(R=5\), \(L = 5.1\). For the regularized functional, we have taken \(\alpha = 4\times 10^{-8}\) and \(\beta = 2\times 10^{-7}\).

Figure 1 (b) illustrates the stability stated in Propositions 45: the minimizer varies continuously with the data (Corollary 2), avoiding the large excursions of the unregularized case, while the no-Lavrentiev-gap property (Theorem 1) guarantees that the polynomial approximation incurs no hidden error. Additional pictures are given in Appendix, see Figures 2 and 3.

3.0.0.3 Examples 3–4: Benchmark datasets (\(d>1\))

We validate the approach on two standard regression benchmarks, comparing against the single-hidden-layer network with \(\mathrm{ReLU}\) activation function with \(N=10{\,}000\) neurons. In both examples we have split the dataset into the train and test with ratio of test to be \(0.2\).

In Example 3, we have used Diabetes dataset [32] containing \(d=10\) features with \(N_D=442\) observations. We have taken \(\Omega=B_{1}^{2}\times (-1,1)\) with \(B_{1}^2\) the unit ball in \({\mathbb{R}}^2\), and \(s=5\) implying \(M=4{,}368\). The sparsity of the matrix \(U\) is \(64\,\%\). We have chosen very small coefficients \(\alpha/C_D = \beta/C_D =10^{-10}\).

In Example 4, we have selected California housing dataset [33]. We take \(\Omega=B_{1}^{8}\times (-1,1)\) (where \(B_{1}^{8}\) is the unit ball in \({\mathbb{R}}^8\)), \(\alpha/C_D = \beta/C_D =10^{-10}\), polynomials of order \(s = 6\), so that \(M = 5{,}005\). The sparsity of the matrix \(U\) is \(34\,\%\). We have selected single category of ocean proximity (near bay) leading to \(N_D = 1{,}860\) observations with \(d=8\) features.

Table 1: Results of approximation in Examples 3 and 4. We present Root Mean Square Error (RMSE), Mean Absolute Error (MAE) and \({R}^2\) coefficient.
Example 3 Example 4
Metric \(\widehat{\mathscr{F}}^{\;(f)}_{0, 0}\) \(\widehat{\mathscr{F}}^{\;(f)}_{\alpha, \beta}\) NN \(\widehat{\mathscr{F}}^{\;(f)}_{0, 0}\) \(\widehat{\mathscr{F}}^{\;(f)}_{\alpha, \beta}\) NN
RMSE \(51.29\) \(52.00\) \(63.28\) \(0.38\) \(0.55\) \(0.29\)
MAE \(40.54\) \(40.39\) \(48.43\) \(0.62\) \(0.39\) \(0.54\)
\(R^2\) \(0.50\) \(0.49\) \(0.24\) \(0.73\) \(0.78\) \(0.79\)

The results presented in Table 1 confirm that the regularized functional achieves competitive or superior accuracy, aligning with Theorem 1. This yields strong finite-sample performance.

4 Conclusion↩︎

The right geometry unlocks the right tools. Casting shallow network training as a variational problem over parameter densities in a weighted Sobolev space gives simultaneous access to global \(\lambda\)-convexity, elliptic regularity, and gradient flow theory, which is a combination unavailable in either the NTK or Wasserstein frameworks. The outcome is a continuum model that is analytically controlled, dynamically stable, and exactly tied to finite-width networks via an explicit \(O(1/N)\) bound: no approximation error, no linearization, no weak regularity.

The near-\(C^3\) smoothness of optimal parameter densities is perhaps the most surprising consequence. It suggests that overparameterized networks are implicitly attracted to low-dimensional smooth structures in parameter space. This is a concrete, quantitative form of implicit bias that goes beyond what existing analyses can see. To be stressed out, it is not an artifact of the limit: the absence of a Lavrentiev gap ensures that the continuum picture faithfully reflects the discrete one. Concretely, the minimizer is found via a single ridge regression, with generalization error bounded by \(C/\alpha\), making \(\alpha\) a directly interpretable robustness parameter rather than a black-box regularizer.

The shallow setting was the necessary first step. Resolving convexity, regularity, and convergence cleanly here opens a concrete path toward multilayer architectures, sharper implicit bias descriptions, and new regularization principles grounded in the underlying variational framework.

Acknowledgments↩︎

MB was supported by internal grant for specific research No. FSI-S-26-8958. IC and BM were supported by National Science Centre Grant 2024/55/B/ST6/016.

5 Technical appendices and supplementary material↩︎

In the case it is clear from the context, we simplify the notation setting \({\mathscr{F}}_{\alpha,\beta}={\mathscr{F}}^{(f)}_{\alpha,\beta}\).

5.1 Proof of Theorem 1↩︎

Theorem 1 is a consequence of two approximation results. In the first one (Proposition 7), we regularize any \(m\in {\mathscr{M}}_1(\Omega)\) by relying on standard convolution and truncation arguments. We then prove that the functional \({\mathscr{R}}(f,\cdot)\) converges along the regularized sequence to \({\mathscr{R}}(f,m)\) by exploiting the growth assumptions satisfied by \(h\). In the second approximation result (Proposition 8), we approximate \(m\) by a sequence of purely atomic measures \(\mu_i\). In order to establish the convergence of \({\mathscr{R}}(f,\mu_i)\) to \({\mathscr{R}}(f,m)\), we strongly rely on the quadratic structure of \({\mathscr{R}}(f,\cdot)\).

Proposition 7. For every \(m\in {\mathscr{M}}_1(\Omega)\), there exists a sequence \((u_j)_{j\geq 1} \subset C^{\infty}_c(\Omega)\) such that \[\lim_{j\to +\infty}{\mathscr{R}}(f,u_j{\,{\rm d}x}) = {\mathscr{R}}(f,m)\,.\]

Proof. We first consider the case when \(m\) belongs to the set \({\mathscr{M}}_c(\Omega)\) of those finite Borel measures with compact support in \(\Omega\). We extend \(m\) as a measure on \({\mathbb{R}}^{d+1}\) simply by assigning to any Borel set \(B\subset {\mathbb{R}}^{d+1}\) the value \(m(B\cap \Omega)\).

Fix \(0<\epsilon_0<\textrm{dist}(\textrm{supp }m, \partial \Omega)\). We then introduce a smooth regularization kernel \((\rho_{\epsilon})_{0<\epsilon<\epsilon_0}\) with \(\rho_{\epsilon}\in C^{\infty}_c(B_\epsilon)\), \(\rho_{\epsilon}\geq 0\) and \(\int_{{\mathbb{R}}^{d+1}}\rho_{\epsilon}=1\) for every \(0<\epsilon<\epsilon_0\). Here, we have denoted by \(B_\epsilon\) the ball of center \(0\) and radius \(\epsilon\) in \({\mathbb{R}}^{d+1}\). We then define \[m_{\epsilon}:=m*\rho_{\epsilon}:x\in \Omega \mapsto \int_{{\mathbb{R}}^{d+1}} \rho_{\epsilon}(x-y){\,{\rm d}}m(y)\,.\] Then, \(m*\rho_{\epsilon}\in C^{\infty}_c(\Omega)\), with \(\textrm{supp }m*\rho_{\epsilon}\subset \textrm{supp }m + B_{\epsilon}\). We claim that \[\label{eq372} \lim_{\epsilon\to 0} {\mathscr{R}}(f, m*\rho_{\epsilon}{\,{\rm d}x}) = {\mathscr{R}}(f,m)\,.\tag{16}\] There exists a compact set \(K\Subset \Omega\) such that \(\textrm{supp }m + B_{\epsilon}\subset K\) for every \(0<\epsilon<\epsilon_0\). Let \(\eta\in C^{\infty}_c(\Omega)\) be a cut-off function; that is, \(0\leq \eta\leq 1\) and \(\eta\equiv 1\) on \(K\). Then, for every \(\kappa\in C(\Omega)\), we have \[\int_{\Omega}\kappa m_{\epsilon}{\,{\rm d}\theta}= \int_{{\mathbb{R}}^{d+1}}(\eta\kappa)m*\rho_{\epsilon}{\,{\rm d}\theta}\,.\] By the Fubini theorem, \[\int_{\Omega}\kappa m_{\epsilon}{\,{\rm d}\theta}= \int_{{\mathbb{R}}^{d+1}} (\eta\kappa)*\check{\rho}_{\epsilon}{\,{\rm d}}m\,,\] where \(\check{\rho}_\epsilon(\cdot)=\rho_\epsilon(-\cdot)\). Since \(((\eta\kappa)*\check{\rho}_{\epsilon})_{\epsilon}\) uniformly converges to \(\eta\kappa\) on \(\Omega\) when \(\epsilon\to 0\), we deduce that \[\lim_{\epsilon\to 0}\int_{\Omega}\kappa m_{\epsilon}{\,{\rm d}\theta}= \lim_{\epsilon\to 0}\int_{{\mathbb{R}}^{d+1}} (\eta\kappa)*\check{\rho}_{\epsilon}{\,{\rm d}}m=\int_{{\mathbb{R}}^{d+1}} \eta\kappa{\,{\rm d}}m\,.\] Since \(\eta\kappa=\kappa\) on \(\textrm{supp } m\), this gives \[\label{eq565} \lim_{\epsilon\to 0}\int_{\Omega}\kappa m_{\epsilon}{\,{\rm d}\theta}= \int_{\Omega}\kappa{\,{\rm d}}m\,.\tag{17}\] We can repeat the same argument with the measure \(\nu= m \otimes m\) on \({\mathbb{R}}^{d+1}\times {\mathbb{R}}^{d+1}\). More specifically, by the Fubini theorem, for every \(\overline{\kappa}\in C(\Omega\times\Omega)\), one has \[\int_{\Omega\times \Omega} \overline{\kappa}{\,{\rm d}}m_{\epsilon}\otimes {\,{\rm d}}m_{\epsilon} =\int_{\Omega\times \Omega} (\widetilde{\kappa}*\widetilde{\rho}_{\epsilon}){\,{\rm d}}m\otimes {\,{\rm d}}m\,,\] where \[\widetilde{\kappa}(x,y):=\eta(x)\eta(y)\overline{\kappa}(x,y)\,, \qquad \widetilde{\rho}_\epsilon(x,y):=\rho_\epsilon(-x)\rho_\epsilon(-y)\,.\] Since \(\widetilde{\kappa}\in C_c(\Omega\times\Omega)\), the integrand in the left-hand side uniformly converges to \(\widetilde{\kappa}\) when \(\epsilon\to 0\), which implies that \[\label{eq585} \lim_{\epsilon\to 0}\int_{\Omega\times \Omega} \overline{\kappa} {\,{\rm d}}m_{\epsilon}\otimes {\,{\rm d}}m_{\epsilon} =\int_{\Omega\times \Omega} \overline{\kappa}{\,{\rm d}}m\otimes {\,{\rm d}}m\,.\tag{18}\]

Recall the decomposition of \({\mathscr{R}}\) from 9 involving \(K\) and \(Q\) given by 10 . Since \(h\) satisfies 3 , for every \(L>0\) and for every \(\theta, {\vartheta}\in B_L\cap \Omega\) it holds \[|h(\theta,x)h({\vartheta},x)|\leq C_{h}^2 (1+L)^2(1+|x|)^2\quad \text{and} \quad |f(x)h(\theta,x)|\leq C_{h}(1+L)(1+|x|)|f(x)|\,.\] For every \(x\in D\), the function \(\theta\mapsto h(\theta,x)\) is continuous and, by assumption, the right-hand sides of both inequalities are summable on the bounded Borel set \(D\). By continuity under the integral sign, we deduce that \(K\) is continuous on \(\Omega\times \Omega\) and that \(Q\) is continuous on \(\Omega\). We can thus apply 18 to \(\overline{\kappa}:=K\) and 17 to \(\kappa:=Q\). This gives \[\begin{align} \lim_{\epsilon\to 0}&\left( \int_{\Omega\times \Omega}K(\theta, {\vartheta})m_{\epsilon}(\theta)m_{\epsilon}( {\vartheta}){\,{\rm d}\theta}{\,{\rm d}}{\vartheta}+ \int_{\Omega}Q(\theta)m_\epsilon(\theta){\,{\rm d}\theta}\right)\\& = \int_{\Omega\times \Omega}K(\theta, {\vartheta}){\,{\rm d}}m(\theta){\,{\rm d}}m ( {\vartheta}) + \int_{\Omega}Q(\theta){\,{\rm d}}m(\theta)\,. \end{align}\] Equivalently, \(\lim_{\epsilon\to 0}{\mathscr{R}}(f,m_\epsilon {\,{\rm d}x})={\mathscr{R}}(f,m)\) which completes the proof of 16 . We have thus proved the conclusion when \(m\in {\mathscr{M}}_c(\Omega)\).

In the general case when \(m\in {\mathscr{M}}_1(\Omega)\), we introduce a sequence \((K_j)_{j\geq 1}\) of compact sets such that \(K_j\subset K_{j+1}\) and \(\cup_{j\geq 1}K_j = \Omega\). We then set \(m_j:=m\llcorner K_j\) for every \(j\geq 1\). Hence, \[\int_{\Omega}Q(\theta){\,{\rm d}}m_j(\theta)=\int_{\Omega} \mathbb{1}_{K_j}(\theta)Q(\theta) {\,{\rm d}}m (\theta)\,,\] where \(\mathbb{1}_{K_j}\) is the indicator function of \(K_j\) and similarly, \[\int_{\Omega\times }K(\theta, {\vartheta}){\,{\rm d}}m_j(\theta)\otimes {\,{\rm d}}m_j({\vartheta})=\int_{\Omega\times \Omega} \mathbb{1}_{K_j\times K_j}(\theta,{\vartheta})K(\theta,{\vartheta}) {\,{\rm d}}m (\theta)\otimes {\,{\rm d}}m({\vartheta})\,.\] Since \(m\in {\mathscr{M}}_1(\Omega)\), it follows from 11 that \(Q\in L^{1}(\Omega, {\,{\rm d}}|m|)\) and \(K\in L^{1}(\Omega\times \Omega, {\,{\rm d}}|m|\otimes |m|)\). Hence, the dominated convergence implies that \[\lim_{j\to +\infty}\int_{\Omega} Q(\theta){\,{\rm d}}m_j(\theta) = \int_{\Omega} Q(\theta){\,{\rm d}}m(\theta)\,,\] and \[\lim_{j\to +\infty}\int_{\Omega\times \Omega}K(\theta, {\vartheta}){\,{\rm d}}m_j(\theta)\otimes {\,{\rm d}}m_j({\vartheta})=\int_{\Omega\times \Omega} (\theta,{\vartheta})K(\theta,{\vartheta}) {\,{\rm d}}m (\theta)\otimes {\,{\rm d}}m({\vartheta})\,.\] Hence, one can deduce that \[\lim_{j\to +\infty}{\mathscr{R}}(f,m_j) = {\mathscr{R}}(f,m)\,.\] Together with the first part of the proof and a diagonal argument, this completes the proof. ◻

Proposition 8. For every \(m\in {\mathscr{M}}_1(\Omega)\), there exists a sequence \((\mu_i)_{i\geq 1}\subset {{\mathscr{M}}^{\rm at}}(\Omega)\) such that \[\lim_{i\to +\infty}{\mathscr{R}}(f,\mu_i) = {\mathscr{R}}(f,m)\,.\]

Proof. In view of Proposition 7, one can assume without loss of generality that \(m({\rm d}x)= u(x){\,{\rm d}x}\) for some \(u\in C^{\infty}_c(\Omega)\). Then, for every \(i\geq 1\), we divide \({\mathbb{R}}^{d+1}\) into a family of closed cubes \((\Sigma^{i}_j)_{j\geq 1}\) of side \(1/i\) and with pairwise disjoint interiors. Observe that the diameter of each such cube is \(\sqrt{d}/i\). We denote by \(a^{i}_{j}\) the center of the cube \(\Sigma^{i}_j\). Let \(J_i\) the subset of those indices \(j\geq 1\) for which \(\Sigma^{i}_{j}\) intersects \({\rm supp\,}u\) and let \(N_i\) be the cardinality of \(J_i\). Finally, we let \[\mu_i:=\frac{1}{i^d}\sum_{j\in J_i} u(a^{i}_j)\delta_{a^{i}_j}\,.\] For every \(i\) larger than \(i_0:=2\sqrt{d}/\textrm{dist }({\rm supp\,}u, \partial \Omega)\), one has \[\textrm{dist }(\cup_{j\in J_i}\Sigma^{i}_{j}, \partial \Omega)\geq \textrm{dist }({\rm supp\,}u, \partial \Omega) - \tfrac{\sqrt{d}}{i} \geq \textrm{dist }({\rm supp\,}u, \partial \Omega)/2\,.\] In particular, there exists a compact \(K\Subset \Omega\) that contains \(\cup_{j\in J_i}\Sigma_{j}^i\) for every \(i\geq i_0\). Still for \(i\geq i_0\), one has \(\Sigma_{j}^{i}\subset \Omega\) and thus, \(|\Omega\cap \Sigma_{j}^i|=|\Sigma_{j}^i|=i^{-d}\). Let \(\eta\in C(\Omega)\). Then, \[\label{eq936} \int_{\Omega} \eta {\,{\rm d}}\mu_i = \frac{1}{i^d} \sum_{j\in J_i}(\eta u )(a^{i}_j)=\sum_{j\in J_i}(\eta u)(a^{i}_j)|\Omega\cap \Sigma_{j}^i|\,.\tag{19}\]

The function \(\eta u\) being uniformly continuous on \(\Omega\), its modulus of continuity \(\kappa_{\eta u}(r):=\sup_{|x-y|\leq r}|(\eta u)(x)-(\eta u(y))|\) converges to \(0\) when \(r\to 0\). Moreover, \[\begin{align} \left|\int_\Omega \eta u{\,{\rm d}x}-\sum_{j\in J_i} |\Omega\cap \Sigma^{i}_j|(\eta u)(a_{j}^i) \right|&\leq \sum_{j\in J_i} \int_{\Omega\cap \Sigma_{j}^{i}}|\eta u-(\eta u)(a^{i}_j)|\\ &\leq \kappa_{\eta u}(\sqrt{d}/i)|\cup_{j\in J_i}\Sigma^{i}_j|\leq \kappa_{\eta u}(\sqrt{d}/i)|K| \,. \end{align}\] Hence, \[\label{eq946} \lim_{i\to +\infty}\left|\int_\Omega \eta u{\,{\rm d}x}-\sum_{j\in J_i} |\Omega\cap \Sigma^{i}_j|(\eta u)(a_{j}^i) \right|=0\,.\tag{20}\] Using 19 and 20 , we deduce that \[\label{eq736} \lim_{i\to +\infty}\int_{\Omega}\eta{\,{\rm d}\mu}_i = \int_{\Omega}\eta u{\,{\rm d}x}\,.\tag{21}\] The above convergence result will be applied to \(\eta=Q\). To obtain a similar result for \(K\), we observe that for every \(\eta_1, \eta_2 \in C(\Omega)\), by the Fubini theorem, \[\begin{gather} \lim_{i\to +\infty}\int_{\Omega\times \Omega} \eta_1(x)\eta_2(y){\,{\rm d}\mu}_i(x){\,{\rm d}\mu}_i(y) = \lim_{i\to +\infty}\left(\int_{\Omega}\eta_1(x){\,{\rm d}\mu}_i(x)\right)\left(\int_{\Omega}\eta_2(y){\,{\rm d}\mu}_i(y)\right)\\ = \left(\int_{\Omega}\eta_1(x)u(x){\,{\rm d}x}\right)\left(\int_{\Omega}\eta_2(y)u(y){\,{\rm d}y}\right) =\int_{\Omega\times \Omega}\eta_1(x)\eta_2(y)u(x)u(y){\,{\rm d}x}\otimes{\,{\rm d}y}\,. \end{gather}\] By linearity, for every \(\overline{\eta}\) of the form \((x,y)\mapsto\sum_{i\in I} \eta_{i1}(x)\eta_{i2}(y)\) where \(I\) is a finite set of indices and \(\eta_{ij}\in C(\Omega)\) for every \(i\in I\), \(j\in \{1,2\}\), one gets \[\label{eq724} \lim_{i\to +\infty}\int_{\Omega\times \Omega} \overline{\eta}(x,y){\,{\rm d}\mu}_i(x){\,{\rm d}\mu}_i(y) = \int_{\Omega\times \Omega} \overline{\eta}(x,y)u(x)u(y){\,{\rm d}x}{\,{\rm d}y}\,.\tag{22}\] Finally, let \(\overline{\kappa}\in C(\Omega\times \Omega)\) and \(\epsilon >0\). Let \(\overline{\eta}\) as above such that \(\max_{(x,y)\in K\times K}|\overline{\kappa}(x,y)-\overline{\eta}(x,y)|\leq \epsilon\) (here, one relies on the Stone–Weierstrass theorem). Then, \[\left|\int_{\Omega\times \Omega} \overline{\kappa}(x,y){\,{\rm d}\mu}_i(x){\,{\rm d}\mu}_i(y) - \int_{\Omega\times \Omega} \overline{\kappa}(x,y)u(x)u(y){\,{\rm d}x}{\,{\rm d}y} \right| \leq I_1+I_2+I_3\,,\] where \[\begin{align} I_1:&=\int_{\Omega\times \Omega} |\overline{\kappa}(x,y)-\overline{\eta}(x,y)|{\,{\rm d}}|\mu_i|(x){\,{\rm d}}|\mu_i|(y)\,,\\ I_2:&= \left|\int_{\Omega\times \Omega} \overline{\eta}(x,y){\,{\rm d}\mu}_i(x){\,{\rm d}\mu}_i(y) - \int_{\Omega\times \Omega} \overline{\eta}(x,y)u(x)u(y){\,{\rm d}x}{\,{\rm d}y} \right|\,,\\ I_3:&=\int_{\Omega\times \Omega} |\overline{\kappa}(x,y)-\overline{\eta}(x,y)||u(x)u(y)|{\,{\rm d}x}{\,{\rm d}y}\,. \end{align}\] By construction, \[|\mu_i|(\Omega)\leq \frac{1}{i^d}\sum_{j\in J_i}|u(a^{i}_j)|\leq \|u\|_{L^{\infty}(\Omega)}\frac{N_i}{i^d}=\|u\|_{L^{\infty}(\Omega)}|\cup_{j\in J_i}\Sigma_{j}^i|\leq \|u\|_{L^{\infty}(\Omega)}|K|\,.\] Hence, \(|\mu_i|\otimes |\mu_i| (\Omega\times \Omega)\leq (\|u\|_{L^{\infty}(\Omega)}|K|)^2\). Moreover, \[(|u|{\,{\rm d}x}\otimes |u|{\,{\rm d}y}) (\Omega\times \Omega)=\|u\|_{L^{1}(\Omega)}^2\leq (\|u\|_{L^{\infty}(\Omega)}|K|)^2\,.\] It follows that \(I_1, I_3\leq (\|u\|_{L^{\infty}(\Omega)}|K|)^2\epsilon\). Together with 22 , this implies that \[\begin{align} \limsup_{i\to +\infty}& \left|\int_{\Omega\times \Omega} \overline{\kappa}(x,y){\,{\rm d}\mu}_i(x){\,{\rm d}\mu}_i(y) - \int_{\Omega\times \Omega} \overline{\kappa}(x,y)u(x)u(y){\,{\rm d}x}{\,{\rm d}y} \right| \\ &\leq 2\epsilon (\|u\|_{L^{\infty}(\Omega)}|K|)^2 \\ &\qquad+ \limsup_{i\to +\infty}\left|\int_{\Omega\times \Omega} \overline{\eta}(x,y){\,{\rm d}\mu}_i(x){\,{\rm d}\mu}_i(y) - \int_{\Omega\times \Omega} \overline{\eta}(x,y){\,{\rm d}\mu}(x){\,{\rm d}\mu}(y) \right|\\ &=2\epsilon (\|u\|_{L^{\infty}(\Omega)}|K|)^2\,. \end{align}\] Since this is true for every \(\epsilon>0\), we deduce therefrom that \[\lim_{i\to +\infty}\int_{\Omega\times \Omega} \overline{\kappa}(x,y){\,{\rm d}\mu}_i(x){\,{\rm d}\mu}_i(y)= \int_{\Omega\times \Omega} \overline{\kappa}(x,y)u(x)u(y){\,{\rm d}x}{\,{\rm d}y}\,.\] The proof is complete. ◻

We now have all the ingredients to present the proof of Theorem 1.

Proof of Theorem 1. Since \(C_c^\infty(\Omega)\subset {\mathcal{W}}\subset {\mathscr{M}}_1(\Omega)\), we have \[\inf_{m \in {\mathscr{M}}_1 (\Omega)} {\mathscr{R}}(f,m)\leq \inf_{v \in {\mathcal{W}}} {\mathscr{R}}(f,v{\,{\rm d}x}) \leq \inf_{v \in C_c^\infty(\Omega)} {\mathscr{R}}(f,v{\,{\rm d}x})\,.\] By Proposition 7, one also has \[\inf_{v \in C_c^\infty(\Omega)} {\mathscr{R}}(f,v{\,{\rm d}x})\leq \inf_{m \in {\mathscr{M}}_1 (\Omega)} {\mathscr{R}}(f,m)\,.\] We can thus conclude that \[\inf_{m \in {\mathscr{M}}_1 (\Omega)} {\mathscr{R}}(f,m)= \inf_{v \in {\mathcal{W}}} {\mathscr{R}}(f,v{\,{\rm d}x}) = \inf_{v \in C_c^\infty(\Omega)} {\mathscr{R}}(f,v{\,{\rm d}x})\,.\] Similarly, the fact that \({{\mathscr{M}}^{\rm at}}(\Omega)\subset {\mathscr{M}}_1 (\Omega)\) and Proposition 8 imply that \[\inf_{m \in {\mathscr{M}}_1 (\Omega)} {\mathscr{R}}(f,m)= \inf_{\mu\in {{\mathscr{M}}^{\rm at}}(\Omega)} {\mathscr{R}}(f,\mu)\,.\] The proof is complete. ◻

5.2 Proof of Proposition 2↩︎

Inspired by the proof of [1] written for probability measures, we provide related result for the more general setting of finite Borel measures. In all this section, we fix \(f\in L^{2}(D)\). For every \(N\geq 1\), for every \(\theta_i\in \Omega\) and \(w_i\in {\mathbb{R}}\) with \(1\leq i \leq N\), the measure \(\mu_N\) defined in 4 satisfies: \[\label{eq-R0-rhoN} {\mathscr{R}}(f,\mu_N)=\|f\|_{L^2(D)}^2+\frac{1}{N^2}\sum_{i,j=1}^N w_i w_j K(\theta_i, \theta_j)-\frac{2}{N}\sum_{i=1}^N w_i Q(\theta_i)\,.\tag{23}\]

Remark 9. Let \(M\geq 1\) and \(\nu_+, \nu_-\) be two probability measures on \(\Omega\). Let \(\big((\overline{\theta}_{i}^+, \overline{\theta}_{i}^{-})\big)_{1\leq i \leq M}\) be a finite family of independent random variables identically distributed of law \(\nu_+\otimes \nu_-\) on \(\Omega\times \Omega\). Then, the law of each \(\overline{\theta}_{i}^+\) is \(\nu_+\), the law of each \(\overline{\theta}_{i}^{-}\) is \(\nu_-\). Moreover, all those random variables \(\overline{\theta}_{i}^\pm\) are pairwise independent.

We begin with the following observation:

Lemma 1. For every finite Borel measure \(m\in {\mathscr{M}}_1(\Omega)\), \[\label{prop-positive-K} \int_{\Omega\times \Omega}K(\tau,{\vartheta}){\,{\rm d}}m(\tau){\,{\rm d}}m({\vartheta}) \geq 0\,.\tag{24}\]

Proof. By definition of \(K\), one has \[\int_{\Omega\times \Omega}K(\tau,{\vartheta}){\,{\rm d}}m(\tau){\,{\rm d}}m({\vartheta}) =\int_{D}\left(\int_{\Omega}h(\theta,x){\,{\rm d}}m(\theta)\right)^2{\,{\rm d}x}\geq 0\,.\] ◻

Remember that \({\mathscr{M}}_2(\Omega)\) denotes the set of all those finite Borel measures on \(\Omega\) such that \[\int_{\Omega}|\tau|^2{\,{\rm d}}|m|(\tau)<+\infty\,.\]

Lemma 2. Let \(\nu_+, \nu_-\) be two probability measures on \(\Omega\) that belong to \({\mathscr{M}}_2(\Omega)\) and are mutually singular. Let \((\overline{\theta}_{i}^+, \overline{\theta}_{i}^{-})_{1\leq i \leq M}\) be a finite family of independent random variables of law \(\nu_+\otimes \nu_-\) on \(\Omega\times \Omega\). For every \(\alpha_+, \alpha_-\geq 0\), we define: \[\overline{\rho}=\frac{\alpha_+}{N} \sum_{i=1}^{N}\delta_{\overline{\theta}_{i}^+} -\frac{\alpha_-}{N}\sum_{i=1}^{N}\delta_{\overline{\theta}_{i}^-}\,.\] Then, for \(m=\alpha_+\nu_+ - \alpha_-\nu_-\), it holds \[\mathbb{E}\big({\mathscr{R}}(f,\overline{\rho})\big)\leq {\mathscr{R}}(f,m)+\frac{|m|(\Omega)}{N}\int_{\Omega}K(\tau, \tau){\,{\rm d}}|m|(\tau)\,.\]

Proof. Observe that a probability measure that belongs to \({\mathscr{M}}_2(\Omega)\) automatically belongs to \({\mathscr{M}}_1(\Omega)\). As a consequence, \(m\in {\mathscr{M}}_1(\Omega)\) and \({\mathscr{R}}(f,m)\) is well-defined. By 23 , one has \[\begin{align} \mathbb{E}\big({\mathscr{R}}(f,\overline{\rho})\big)=& \|f\|_{L^2(D)}^2+\frac{\alpha_{+}^2}{N^2}\sum_{i,j=1}^N \mathbb{E}\big(K(\overline{\theta}_{i}^+, \overline{\theta}_{j}^+)\big) +\frac{\alpha_{-}^2}{N^2}\sum_{i,j=1}^N \mathbb{E}\big(K(\overline{\theta}_{i}^-, \overline{\theta}_{j}^-)\big)\\ &\,-2\frac{\alpha_+\alpha_-}{N^2}\sum_{i,j=1}^N \mathbb{E}\big(K(\overline{\theta}_{i}^+, \overline{\theta}_{j}^-)\big)-\frac{2\alpha_+}{N}\sum_{i=1}^N \mathbb{E}\big(Q(\overline{\theta}_{i}^+)\big)+\frac{2\alpha_-}{N}\sum_{i=1}^N \mathbb{E}\big(Q(\overline{\theta}_{i}^-)\big)\,. \end{align}\] By Remark 9, the law of \(\theta_{i}^{\pm}\) is \(\nu_{\pm}\). Hence, \[\mathbb{E}(Q(\overline{\theta}_{i}^\pm))=\int_{\Omega}Q(\tau){\,{\rm d}}\nu_\pm(\tau)\,.\] It follows that \[\begin{align} \frac{2\alpha_+}{N}\sum_{i=1}^N \mathbb{E}\big(Q(\overline{\theta}_{i}^+)\big)-\frac{2\alpha_-}{N}\sum_{i=1}^N \mathbb{E}\big(Q(\overline{\theta}_{i}^-)\big) &= 2 \alpha_+\int_{\Omega}Q(\tau){\,{\rm d}}\nu_+(\tau)-2 \alpha_-\int_{\Omega}Q(\tau){\,{\rm d}}\nu_-(\tau)\\ &= 2 \int_{\Omega}Q(\tau){\,{\rm d}}m\,. \end{align}\] By Remark 9 again, the random variables \(\overline{\theta}_{i}^{\pm}\) are pairwise independent. It follows that \[\mathbb{E}\big(K(\overline{\theta}_{i}^+, \overline{\theta}_{j}^-)\big) = \int_{\Omega\times \Omega} K(\tau, {\vartheta}) {\,{\rm d}}\nu_+\otimes {\,{\rm d}}\nu_-(\tau,{\vartheta})\,, \qquad \forall 1\leq i, j\leq N\,,\] and \[\mathbb{E}\big(K(\overline{\theta}_{i}^\pm, \overline{\theta}_{j}^\pm)\big) = \int_{\Omega\times \Omega} K(\tau, {\vartheta}) {\,{\rm d}}\nu_\pm\otimes {\,{\rm d}}\nu_\pm(\tau,{\vartheta})\,, \qquad \forall 1\leq i\not= j\leq N\,.\] Moreover, \[\mathbb{E}\big(K(\overline{\theta}_{i}^\pm, \overline{\theta}_{i}^\pm)\big)=\int_\Omega K(\tau, \tau){\,{\rm d}}\nu_\pm(\tau)\,.\] We thus get \[\begin{align} \mathbb{E}\big({\mathscr{R}}(f,\overline{\rho})\big)&=\|f\|_{L^2(D)}^2+\frac{\alpha_{+}^2(N^2-N)}{N^2}\int_{\Omega\times \Omega}K(\tau,{\vartheta}){\,{\rm d}}\nu_+(\tau){\,{\rm d}}\nu_+({\vartheta})\\ &\quad + \frac{\alpha_{+}^2}{N}\int_{\Omega}K(\tau,\tau){\,{\rm d}}\nu_+(\tau) +\frac{\alpha_{-}^2(N^2-N)}{N^2}\int_{\Omega\times \Omega}K(\tau,{\vartheta}){\,{\rm d}}\nu_-(\tau){\,{\rm d}}\nu_-({\vartheta})\\ &\quad + \frac{\alpha_{-}^2}{N}\int_{\Omega}K(\tau,\tau){\,{\rm d}}\nu_-(\tau) -2\alpha_+\alpha_-\int_{\Omega}K(\tau, {\vartheta}){\,{\rm d}}\nu_+(\tau){\,{\rm d}}\nu_-({\vartheta}) \\ &\quad -2\int_{\Omega}Q(\tau){\,{\rm d}}m(\tau)\,. \end{align}\] Taking into account the definition of \(m=\alpha_+\nu_+-\alpha_-\nu_-\), this gives \[\begin{align} \mathbb{E}\big({\mathscr{R}}(f,\overline{\rho})\big) &={\mathscr{R}}(f,m) -\frac{\alpha_{+}^2}{N}\int_{\Omega\times \Omega}K(\tau,{\vartheta}){\,{\rm d}}\nu_+(\tau){\,{\rm d}}\nu_+({\vartheta})\\ &\quad- \frac{\alpha_{-}^2}{N}\int_{\Omega\times \Omega}K(\tau,{\vartheta}){\,{\rm d}}\nu_-(\tau){\,{\rm d}}\nu_-({\vartheta})\\ &\quad +\frac{\alpha_{+}^2}{N}\int_{\Omega}K(\tau, \tau){\,{\rm d}}\nu_+(\tau)+\frac{\alpha_{-}^2}{N}\int_{\Omega}K({\vartheta}, {\vartheta}){\,{\rm d}}\nu_-({\vartheta})\,. \end{align}\]

Hence, by 24 , \[\mathbb{E}\big({\mathscr{R}}(f,\overline{\rho})\big) \leq {\mathscr{R}}(f,m)+\frac{\alpha_{+}^2}{N}\int_{\Omega}K(\tau, \tau){\,{\rm d}}\nu_+(\tau)+\frac{\alpha_{-}^2}{N}\int_{\Omega}K({\vartheta}, {\vartheta}){\,{\rm d}}\nu_-({\vartheta})\,,\] which implies the desired result since \(\alpha_++\alpha_-=|m|(\Omega)\) and \(|m|=\alpha_+\nu_++\alpha_-\nu_-\). ◻

Proposition 10. For every \(N\geq 1\) and for every \(m\in {\mathscr{M}}_2(\Omega)\), \[\inf_{\rho\in {\mathscr{M}}^{\rm at}_{2N}(\Omega)} {\mathscr{R}}(f,\rho) \leq {\mathscr{R}}(f,m)+ \frac{|m|(\Omega)}{N} \int_{\Omega}K(\tau, \tau){\,{\rm d}}|m|(\tau)\,.\]

Proof. Let \(m=m_+-m_-\) be the Haar decomposition of \(m\). We introduce the nonnegative numbers \(\alpha_+=m_+(\Omega)\) and \(\alpha_-=m_-(\Omega)\). Let \(\nu_+, \nu_-\) be two mutually singular probability measures in \({\mathscr{M}}_2(\Omega)\) such that \(m_+=\alpha_+\nu_+\), \(m_-=\alpha_-\nu_-\). Let \((\overline{\theta}_{i}^+, \overline{\theta}_{i}^{-})_{1\leq i \leq M}\) be a finite family of independent random variables of law \(\nu_+\otimes \nu_-\) on \(\Omega\times \Omega\). We then define \[\overline{\rho}:=\frac{\alpha_+}{N} \sum_{i=1}^{N}\delta_{\overline{\theta}_{i}^+} -\frac{\alpha_-}{N}\sum_{i=1}^{N}\delta_{\overline{\theta}_{i}^-}\,.\] By Lemma 2, \[\mathbb{E}\big({\mathscr{R}}(f,\overline{\rho})\big)\leq {\mathscr{R}}(f,m)+\frac{|m|(\Omega)}{N}\int_{\Omega}K(\tau, \tau){\,{\rm d}}|m|(\tau)\,.\] There exists a choice of \(\theta_{i}^{+}, \theta_{i}^{-}\in \Omega\) with \(1\leq i \leq N\) such that the measure \[\rho=\frac{\alpha_+}{N} \sum_{i=1}^{N}\delta_{\theta_{i}^+} -\frac{\alpha_-}{N}\sum_{i=1}^{N}\delta_{\theta_{i}^-}\] satisfies \[{\mathscr{R}}(f,\rho) \leq \mathbb{E}\big({\mathscr{R}}(f,\overline{\rho})\big)\,.\] Hence, \[{\mathscr{R}}(f,\rho)\leq {\mathscr{R}}(f,m)+\frac{|m|(\Omega)}{N}\int_{\Omega}K(\tau, \tau){\,{\rm d}}|m|(\tau)\,.\] This implies that \[\inf_{\rho\in {\mathscr{M}}_{2N}^{\rm at}(\Omega)}{\mathscr{R}}(f,\rho) \leq {\mathscr{R}}(f,m)+\frac{|m|(\Omega)}{N}\int_{\Omega}K(\tau, \tau){\,{\rm d}}|m|(\tau)\,.\] ◻

We conclude this section with the proof of Proposition 2.

Proof of Proposition 2. Let \(v\in L^{2}_{\omega}(\Omega)\). Then, by the Schwarz inequality and 7 , \[\int_{\Omega} (1+|\theta|)^2 |v(\theta)|{\,{\rm d}}\theta \leq \sqrt{c_\omega} \|v\|_{L^{2}_{\omega}(\Omega)}\,.\] This proves that the measure \(m:=v{\,{\rm d}\theta}\) belongs to \({\mathscr{M}}_2(\Omega)\). Similarly, we have \[|m|(\Omega)=\int_{\Omega}|v(\theta)|{\,{\rm d}}\theta \leq \sqrt{c_\omega} \|v\|_{L^{2}_{\omega}(\Omega)}\,.\] Applying Proposition 10 to this measure \(m\) and using 11 , one gets for every \(N\geq 2\), \[\begin{align} \inf_{\rho\in {\mathscr{M}}^{\rm at}_{N}(\Omega)} {\mathscr{R}}(f,\rho)&\leq \inf_{\rho\in {\mathscr{M}}^{\rm at}_{2\lfloor N/2\rfloor}(\Omega)} {\mathscr{R}}(f,\rho)\\ &\leq {\mathscr{R}}(f,m)+ \frac{4|m|(\Omega)}{N} \int_{\Omega}K(\tau, \tau){\,{\rm d}}|m|(\tau)\\ &\leq {\mathscr{R}}(f,v{\,{\rm d}}\theta)+ \frac{4\sqrt{c_\omega}\|v\|_{L^{2}_{\omega}(\Omega)}}{N} \int_{\Omega}(C_{h}')^2(1+|\tau|)^2|v|(\tau){\,{\rm d}}\tau\\ &\leq {\mathscr{R}}(f,v{\,{\rm d}}\theta)+ \frac{4(C_{h}')^2c_\omega\|v\|^{2}_{L^{2}_{\omega}(\Omega)}}{N}\,. \end{align}\] This proves the first assertion of Proposition 2 with \(C:=4(C_{h}')^2c_\omega\). If one additionnally assumes that there exists a bounded minimizing sequence \((v_j)_{j\geq 1}\) for \({\mathscr{R}}(f,\cdot)\) in \(L^{2}_{\omega}(\Omega)\), then applying the above estimate for every \(j\geq 1\), one gets \[\inf_{\rho\in {\mathscr{M}}^{\rm at}_{N}(\Omega)} {\mathscr{R}}(f,\rho) \leq {\mathscr{R}}(f,v_j{\,{\rm d}}\theta)+ \frac{C}{N}\sup_{k\geq 1}\|v_k\|^{2}_{L^{2}_{\omega}(\Omega)}\,.\] Passing to the limit \(j\to +\infty\), one deduces that \[\inf_{\rho\in {\mathscr{M}}^{\rm at}_{N}(\Omega)} {\mathscr{R}}(f,\rho) \leq \inf_{v\in L^{2}_\omega(\Omega)}{\mathscr{R}}(f,v{\,{\rm d}}\theta)+ \frac{C}{N}\sup_{k\geq 1}\|v_k\|^{2}_{L^{2}_{\omega}(\Omega)}\,,\] which implies the desired conclusion. ◻

5.3 Proof of Theorem 3↩︎

In order to establish Theorem 3, and more specifically the existence and uniqueness of the minimizer of \({\mathscr{F}}^{(f)}_{\alpha, \beta}\), we rely on the uniform convexity of that functional. In turn, the latter is a simple consequence of the observation already formulated in 24 .

Lemma 3. For every \(f\in L^{2}(D)\), the functional \(m\in {\mathscr{M}}_1(\Omega)\to {\mathscr{R}}(f,m)\) is convex.

Proof. By 9 , the functional \({\mathscr{R}}(f,m)\) is the sum of an affine term \(m\mapsto \|f\|_{L^{2}(D)}^2-2\int_{\Omega}Q(\theta){\,{\rm d}}m (\theta)\) and a functional \[m\mapsto \int_{\Omega\times \Omega}K(\theta, {\vartheta}){\,{\rm d}}m (\theta) {\,{\rm d}}m ({\vartheta})\,.\] which is quadratic (here, we use that \(K(\theta, {\vartheta})=K({\vartheta}, \theta)\)) and nonnegative (as already observed in 24 ). It follows that \({\mathscr{R}}(f,\cdot)\) is convex, as desired. ◻

For later use, we introduce for every \(g\in L^{2}_{\omega}(\Omega)\) and every \(\alpha, \beta \geq 0\), the following variant of the functional \({\mathscr{F}}_{\alpha, \beta}\): \[\begin{gather} {\mathscr{J}}_{\alpha, \beta, g}(u):=\alpha \|u\|_{L^{2}_\omega(\Omega)}^2+\beta \|\nabla u\|_{L^{2}(\Omega)}^2 + \int_{\Omega\times \Omega}K(\theta, {\vartheta})u(\theta)u({\vartheta}){\,{\rm d}\theta}{\,{\rm d}}{\vartheta}- \int_{\Omega}g(\theta)u(\theta)\omega(\theta){\,{\rm d}\theta}\,. \end{gather}\]

Lemma 4. For every \(\alpha, \beta>0\) and every \(g\in L^{2}_\omega(\Omega)\), the functional \({\mathscr{J}}_{\alpha, \beta, g}:{\mathcal{W}}\to {\mathbb{R}}\) is \(2\min(\alpha, \beta)\)-convex on \({\mathcal{W}}\), and thus admits a unique minimum \(\bar{u}\) on \({\mathcal{W}}\). Moreover, \(\bar{u}\in W^{2,2}_{loc}(\Omega)\) and \(p(\bar{u}):=\frac{1}{\omega} \Delta \bar{u}\) satisfies that \[p(\bar{u}) = \frac{1}{\beta}\left(\alpha \bar{u} + \frac{1}{\omega}\int_{\Omega}K(\cdot, {\vartheta})\bar{u}({\vartheta}){\,{\rm d}}{\vartheta}-\frac{1}{2}g\right)\in L^{2}_\omega(\Omega)\,.\] Moreover, for every \(v\in {\mathcal{W}}\), \(\int_{\Omega} \nabla \bar{u}\cdot \nabla v = -\int_{\Omega}p(\bar{u})v\omega\).

Proof. The map \[{\mathscr{J}}_{0,0,g}:u\in {\mathcal{W}}\mapsto \int_{\Omega\times \Omega}K(\theta, {\vartheta})u(\theta)u({\vartheta}){\,{\rm d}\theta}{\,{\rm d}}{\vartheta}- \int_{\Omega}g(\theta)u(\theta)\omega(\theta){\,{\rm d}\theta}\] is convex as the sum of a convex quadratic function and an affine function. For every \(u\in {\mathcal{W}}\), \[{\mathscr{J}}_{\alpha, \beta, g}(u)-\min(\alpha, \beta)\|u\|^{2}_{{\mathcal{W}}} =(\alpha-\min(\alpha, \beta))\|u\|^{2}_{L^{2}_\omega(\Omega)}+(\beta-\min(\alpha, \beta))\|\nabla u\|_{L^{2}(\Omega)}^2+{\mathscr{J}}_{0,0,g}(u)\,.\] Since the right-hand side is convex, we deduce that \({\mathscr{J}}_{\alpha, \beta, g}\) is \(2\min(\alpha, \beta)\)-convex on \({\mathcal{W}}\), and thus strictly convex and coercive. Hence, there exists a unique minimum \(\bar{u}\). Moreover, the restriction of \({\mathscr{J}}_{\alpha, \beta, g}\) is smooth on \({\mathcal{W}}\), so that the Euler equation \(D{\mathscr{J}}_{\alpha, \beta, g}(\bar{u})=0\) holds, namely, for all \(v\in {\mathcal{W}}\) \[\label{eq1258} \beta \int_{\Omega}\nabla \bar{u}\cdot \nabla v{\,{\rm d}}\theta + \alpha\int_{\Omega}\bar{u} v \omega{\,{\rm d}}\theta + \int_{\Omega\times \Omega}K(\theta, {\vartheta})\bar{u}(\theta)v({\vartheta}){\,{\rm d}}\theta {\,{\rm d}}{\vartheta}-\frac{1}{2}\int_{\Omega}vg\omega{\,{\rm d}}\theta=0\,,\tag{25}\] and thus, in the distributional sense, \[\label{eq1269} \beta\Delta\bar{u} = \alpha \omega \bar{u}+\int_{\Omega}K(\cdot,{\vartheta})\bar{u}({\vartheta}){\,{\rm d}}{\vartheta}-\frac{1}{2}g\omega\,.\tag{26}\] Since \(\omega\in C^{\infty}(\Omega)\subset L^{\infty}_{loc}(\Omega)\) and \(\bar{u}, g\in L^{2}_{\omega}(\Omega)\), we deduce that \(\alpha \omega \bar{u}-\frac{1}{2}g\omega\in L^{2}_{loc}(\Omega)\). Moreover, by 11 and the Schwarz inequality, \[\left|\int_{\Omega}K(\theta, {\vartheta})\bar{u}({\vartheta}){\,{\rm d}}{\vartheta}\right|\leq (C_{h}')^2\sqrt{c_\omega}(1+|\theta|)\|\bar{u}\|_{L^{2}_\omega(\Omega)}\,.\] Hence, \[\int_{\Omega}\frac{1}{\omega(\theta)}\left( \int_{\Omega}K(\theta, {\vartheta})\bar{u}({\vartheta}){\,{\rm d}}{\vartheta}\right)^2{\,{\rm d}\theta} \leq (C_{h}')^4c_\omega^2\|\bar{u}\|_{L^{2}_\omega(\Omega)}^2\,.\] This proves that the map \[\theta\mapsto \frac{1}{\omega(\theta)}\int_{\Omega}K(\theta,{\vartheta})\bar{u}({\vartheta}){\,{\rm d}}{\vartheta}\] belongs to \(L^{2}_{\omega}(\Omega)\) and thus also to \(L^{2}_{loc}(\Omega)\) (here, we use that \(\omega\geq 1\)). Hence, the right-hand side of 26 is in \(L^{2}_{loc}(\Omega)\), so that \(\bar{u}\in W^{2,2}_{loc}(\Omega)\), see [34]. Setting \(p(\bar{u}):=\frac{\Delta\bar{u}}{\omega}\), we thus have \[\label{eq1275} p(\bar{u})=\frac{1}{\beta}\left(\alpha \bar{u}+\frac{1}{\omega}\int_{\Omega}K(\cdot,{\vartheta})\bar{u}({\vartheta}){\,{\rm d}}{\vartheta}-\frac{1}{2}g \right)\in L^{2}_{\omega}(\Omega)\,.\tag{27}\]

Finally, for every \(v\in {\mathcal{W}}\), we deduce from 25 that \[\int_{\Omega}\nabla \bar{u}\cdot \nabla v{\,{\rm d}}\theta =\frac{-1}{\beta}\left(\alpha\int_{\Omega}\bar{u} v \omega{\,{\rm d}}\theta + \int_{\Omega\times \Omega}K(\theta, {\vartheta})\bar{u}(\theta)v({\vartheta}){\,{\rm d}}\theta {\,{\rm d}}{\vartheta}-\frac{1}{2}\int_{\Omega}vg\omega{\,{\rm d}}\theta\right)\,.\] Hence, by 27 , one gets \[\int_{\Omega}\nabla \bar{u}\cdot \nabla v{\,{\rm d}}\theta =-\int_{\Omega}p(\bar{u})v\omega{\,{\rm d}}\theta\,.\] The proof is complete. ◻

The regularity of the minimizer \(u^*\) stated in Theorem 3 relies on standard elliptic estimates satisfied by the Euler equation associated to the functional \({\mathscr{F}}^{(f)}_{\alpha, \beta}\). In order to fully exploit this elliptic structure, we need to establish some regularity properties of the right-hand side of the Euler equation.

Lemma 5. Assume that \(h\) satisfies 3 and 13 . Given \(u^*\in L^{2}_{\omega}(\Omega)\), we consider the function \[\ell:\theta \in \Omega \mapsto -Q(\theta)+\int_{\Omega}K(\theta,{\vartheta})u^*({\vartheta})\,d{\vartheta}.\] Then, \(\ell\) is locally Lipschitz on \(\Omega\).

Proof. Let \(R>0\) and \(\theta, {\vartheta}\in \Omega \cap B_R\). Then, by definition of \(Q\) and 13 , \[|Q(\theta)-Q({\vartheta})|\leq \int_{D}|f(x)||h(\theta,x)-h({\vartheta},x)|\,dx \leq c_{h,R} |\theta-{\vartheta}| \int_{D}|f(x)|(1+|x|)\,dx.\] By definition of \(K\) and 13 again, for every \(\tau\in \Omega\), \[|K(\theta,\tau)-K({\vartheta},\tau)|\leq c_{h,R}|\theta-{\vartheta}|\int_{D}(1+|x|)|h(\tau,x)|\,dx.\] Hence, integrating over \(\tau\in \Omega\) and using 3 , \[\begin{align} \left|\int_{\Omega}(K(\theta,\tau)-K({\vartheta},\tau))u^*(\tau){\,{\rm d}}\tau\right| &\leq c_{h,R}|\theta-{\vartheta}|\int_{D\times \Omega}(1+|x|)|h(\tau,x)||u^*(\tau)|{\,{\rm d}x}{\,{\rm d}}\tau\\ &\leq C_{h}c_{h,R}|\theta-{\vartheta}|\int_{D}(1+|x|)^2\,dx\int_{\Omega}(1+|\tau|)|u^*(\tau)|\,d\tau\\ &\leq C_{h}c_{h,R}|\theta-{\vartheta}|\int_{D}(1+|x|)^2\,dx\|u^*\|_{L^{2}_\omega(\Omega)}\sqrt{c_\omega}\,, \end{align}\] where the last inequality relies on 3 and the Schwarz inequality. This implies that \(\ell\in W^{1,\infty}_{loc}(\Omega)\) and completes the proof. ◻

The existence and regularity of the minimizer in Theorem 3 can now be obtained through classical tools from convexity theory and bootstrap arguments often used in elliptic regularity.

Proof of Theorem 3. We first consider the restriction of \({\mathscr{F}}_{\alpha, \beta}={\mathscr{F}}^{(f)}_{\alpha,\beta}\) to its domain \({\mathcal{W}}\). From Lemma 4 with \(g=2Q/\omega\), one deduces that \({\mathscr{F}}_{\alpha, \beta}\) is \(2\min(\alpha, \beta)\)-convex on the Hilbert space \({\mathcal{W}}\), and thus attains a unique minimum \(u^*=u^*_f\), which is a weak solution of the linear elliptic equation: \[\label{eq-Euler-proof-weak} -\beta\Delta u^* +\alpha \omega u^* +\int_{\Omega}K(\cdot, {\vartheta})u^*({\vartheta}){\,{\rm d}}{\vartheta}-Q=0\,.\tag{28}\] By Lemma 5, the function \(\ell:=-Q+\int_{\Omega}K(\cdot, {\vartheta})u^*({\vartheta}){\,{\rm d}}{\vartheta}\) is locally Lipschitz on \(\Omega\).

Using that \(u_*\in W^{1,2}(\Omega)\) and \(\omega\in C^{\infty}(\Omega)\), we get that \(\omega u^*\in W^{1,2}_{loc}(\Omega)\). Since \(\ell\in W^{1,2}_{loc}(\Omega)\), classical elliptic estimates yield \(u^*\in W^{3,2}_{loc}(\Omega)\), see [34]. By the Sobolev embeddings, one deduces that \(u^*\in W^{1,2^{**}}_{loc}(\Omega)\), where \[\frac{1}{2^{**}}=\frac{1}{2^*}-\frac{1}{d+1}=\frac{1}{2}-\frac{2}{d+1}\] if \(d\geq 4\), while \(2^{**}\) is any number \(>1\) otherwise. Using that \(\ell\in W^{1,2^{**}}_{loc}(\Omega)\), it follows from 28 again that \(u^*\in W^{3,2^{**}}_{loc}(\Omega)\), see [34]. By a standard bootstrap strategy, we can conclude that \(u^*\in W^{3,p}_{loc}(\Omega)\) for every \(p>1\). By the Morrey embeddings, this implies that \(u^*\in C^{2,s}(\Omega)\) for every \(s\in (0,1)\).

Finally, since the functional \[v \mapsto {\mathscr{F}}_{\alpha, \beta}(v) -\alpha \int_{\Omega} v^2(\theta)\omega(\theta){\,{\rm d}\theta}\] is convex on \({\mathcal{W}}\) (being a convex subset of \(L^{2}_\omega(\Omega)\)), the functional \({\mathscr{F}}_{\alpha, \beta}\) is \(2\alpha\)-convex on \(L^{2}_{\omega}(\Omega)\). ◻

We proceed with the stability results, first in \({\mathcal{W}}\) (Proposition 4) and then in \(C^{2}_{loc}(\Omega)\) (Proposition 5). We also justify the continuity of \(\inf_{L^{2}_{\omega}(\Omega)}{\mathscr{F}}^{(f)}_{\alpha, \beta}\) with respect to \(f\) (Corollary 1).

Proof of Proposition 4. Let \(\alpha, \beta>0\) and \(f\in L^{2}(D)\). Since the corresponding functional \({\mathscr{F}}_{\alpha, \beta}\) has a unique minimizer \(u_f\), the latter is the unique solution of the linear equation: \[\alpha \int_{\Omega}u_fv\omega + \beta \int_{\Omega} \nabla u_f\cdot \nabla v + \int_{\Omega\times \Omega}K(\theta, {\vartheta})u_f(\theta)v({\vartheta}){\,{\rm d}\theta}{\,{\rm d}}{\vartheta}=\int_{\Omega}Qv{\,{\rm d}\theta}\quad \forall v\in {\mathcal{W}}\,.\] We deduce that \(u_f\) depends linearly on \(Q\). Since \(Q\) depends linearly on \(f\), we can conclude that \(u_f\) depends linearly on \(f\). Inserting \(v=u_f\) in the above identity and using 24 , one gets \[\alpha \|u_f\|_{L^{2}_{\omega}(\Omega)}^2 +\beta \|\nabla u_f\|_{L^{2}(\Omega)}^2\leq \int_{\Omega}Qu_f{\,{\rm d}\theta}\,.\] Using the Schwarz and then the Young inequality in the right-hand side, one gets \[\label{eq1267} \frac{\alpha}{2} \|u_f\|_{L^{2}_{\omega}(\Omega)}^2 +\beta \|\nabla u_f\|_{L^{2}(\Omega)}^2\leq \frac{1}{2\alpha}\int_{\Omega}\frac{Q(\theta)^2}{\omega(\theta)}{\,{\rm d}\theta}\,.\tag{29}\] From 11 , one gets \[|Q(\theta)| \leq C_h(1+|\theta|)\int_{D}(1+|x|)|f(x)|{\,{\rm d}x}\leq C' (1+|\theta|)\|f\|_{L^{2}(D)}\,,\] where \(C'=C'(D,h)>0\). Hence, \[\label{eq1284} \int_{\Omega}\frac{Q(\theta)^2}{\omega(\theta)} \leq C'^2 \|f\|_{L^{2}(D)}^2 \int_{\Omega}\frac{(1+|\theta|)^2}{\omega(\theta)}{\,{\rm d}\theta}\leq C'^2 c_\omega\|f\|_{L^{2}(D)}^2\,,\tag{30}\] where \(c_\omega\) is defined in 7 . Inserting the above estimate into 29 yields the desired result. ◻

Proof of Corollary 1. By minimality of \(u_{f_1}^{*}\), \[{\mathscr{F}}_{\alpha, \beta}^{(f_1)}(u_{f_1}^{*}) \leq {\mathscr{F}}_{\alpha, \beta}^{(f_1)}(u_{f_2}^{*})={\mathscr{F}}_{\alpha, \beta}^{(f_2)}(u_{f_2}^{*})+2\int_{\Omega}(Q_{f_2}-Q_{f_1})u_{f_2}^{*}{\,{\rm d}\theta}\,,\] where \(Q_{f_i}=\int_{D}f_i(x) h(\theta,x){\,{\rm d}x}\). Hence, by the fact that \(Q_{f_2}-Q_{f_1}=Q_{f_2-f_1}\) and the Schwarz inequality, one gets \[{\mathscr{F}}_{\alpha, \beta}^{(f_1)}(u_{f_1}^{*})-{\mathscr{F}}_{\alpha, \beta}^{(f_2)}(u_{f_2}^{*})\leq 2\|u_{f_2}^{*}\|_{L^{2}_\omega(\Omega)}\left(\int_{\Omega}\frac{|Q_{f_2-f_1}|^2}{\omega}{\,{\rm d}\theta}\right)^{1/2}\,.\] Relying on 30 and Proposition 4, we obtain \[{\mathscr{F}}_{\alpha, \beta}^{(f_1)}(u_{f_1}^{*})-{\mathscr{F}}_{\alpha, \beta}^{(f_2)}(u_{f_2}^{*})\leq \frac{C}{\alpha} \|f_2\|_{L^{2}(D)}\|f_2-f_1\|_{L^{2}(D)}\,,\] where \(C=C(\Omega, D, \omega,h)>0\). Symetrically, one also has \[{\mathscr{F}}_{\alpha, \beta}^{(f_2)}(u_{f_2}^{*})-{\mathscr{F}}_{\alpha, \beta}^{(f_1)}(u_{f_1}^{*})\leq \frac{C}{\alpha} \|f_1\|_{L^{2}(D)}\|f_2-f_1\|_{L^{2}(D)}\,.\] The two inequalities above yield the desired conclusion. ◻

Proof of Proposition 5. Fix \(\Omega'\Subset \Omega\). Then, by Theorem 3, the restriction \(u_{f}|_{\overline{\Omega'}}\) belongs to \(C^{2}(\overline{\Omega'})\). We only need to establish the continuity of the linear map \[\label{eq1288} f\in L^{2}(D)\mapsto u_{f}|_{\overline{\Omega'}}\in C^{2}(\overline{\Omega'})\,.\tag{31}\] Let \((f_k)_{k\geq 1}\subset L^{2}(D)\) converge to \(f\in L^{2}(D)\) and assume that \((u_{f_k}|_{\overline{\Omega'}})_{k\geq 1}\) converges to some \(v\in C^{2}(\overline{\Omega'})\). By Proposition 4, we know that \((u_{f_k})_{k\geq 1}\) converges to \(u_f\) in \({\mathcal{W}}\). We deduce that \(v=u_{f}|_{\overline{\Omega'}}\). Since \(L^{2}(D)\) and \(C^{2}(\overline{\Omega'})\) are Banach spaces, one is entitled to apply the closed graph theorem and deduce that the map in 31 is continuous, as desired. ◻

5.4 On the gradient flow in \(L^{2}_\omega(\Omega)\)↩︎

5.4.1 The Hille–Yosida approach↩︎

To study the gradient flow associated to the minimization of \({\mathscr{F}}_{\alpha, \beta}\), we first rely on the Hille–Yosida approach, see [35].

Remember that \(A:D(A)\subset L^{2}_\omega(\Omega)\to L^{2}_\omega(\Omega)\) is the unbounded linear operator defined by \[D(A)=\left\lbrace u\in {\mathcal{W}}: \exists p=p(u)\in L^{2}_{\omega}(\Omega) \textrm{ such that } \forall v\in {\mathcal{W}}, \int_{\Omega}\nabla u \cdot \nabla v = -\int_{\Omega}pv\omega\right\rbrace\,,\] \[Au=2\alpha u -2\beta p(u)+\frac{2}{\omega}\int_{\Omega}K(\theta, \cdot)u(\theta){\,{\rm d}\theta}\,.\] Observe that for every \(u\in D(A)\), the function \(p(u)\) is \(\Delta u/\omega\). Hence, \(\Delta u\in L^{2}_{loc}(\Omega)\), so that by [34], the function \(u\) is in \(W^{2,2}_{loc}(\Omega)\).

Lemma 6. If \(\Omega\) is \(C^2\) and bounded and \(\omega\in C^{\infty}(\overline{\Omega})\), then \[D(A)=\left\lbrace u\in W^{2,2}(\Omega) : \tfrac{\partial u}{\partial \nu}|_{\partial \Omega}=0 \right\rbrace \,.\]

Proof. Since \(\Omega\) is bounded and \(\omega\) is bounded from below and from above by positive constants, one has \(L^{2}_{\omega}(\Omega)=L^{2}(\Omega)\) and thus \({\mathcal{W}}=W^{1,2}(\Omega)\). Moreover, for every \(p\in L^{2}(\Omega)\), the function \(p\omega\) belongs to \(L^{2}(\Omega)\). Hence, the conclusion follows from standard elliptic estimates, see e.g. [35]. ◻

In the next lemma, we check that \(A\) satisfies all the required properties to apply the Hille–Yosida theorem.

Lemma 7. The unbounded linear operator \(A:D(A)\subset L^{2}_{\omega}(\Omega)\to L^{2}_\omega(\Omega)\) is an autoadjoint maximal monotone operator.

Proof. For every \(u\in D(A)\), \[\begin{align} \langle A(u),u\rangle_{L^{2}_\omega(\Omega)} &= 2\int_{\Omega}\left( \alpha u - \beta p(u)+\frac{1}{\omega}\int_{\Omega}K(\theta, \cdot)u\right) u\omega\\ &=2\alpha \int_{\Omega}u^2 \omega +2\beta \int_{\Omega}|\nabla u|^2 + 2 \int_{\Omega\times \Omega}K(\theta, {\vartheta})u(\theta)u({\vartheta}){\,{\rm d}\theta}{\,{\rm d}}{\vartheta}\geq 0\,. \end{align}\] In the last inequality, we have used 24 . This proves that \(A\) is monotone.

In order to prove that \(A\) is maximal, let \(g\in L^{2}_\omega(\Omega)\) and consider the functional \({\mathscr{J}}_{\alpha+1/2, \beta, g}\). Then, by Lemma 4, this functional admits a unique minimizer \(\bar{u}\) on \({\mathcal{W}}\) that belongs to \(D(A)\) and satisfies \[\label{eq1271} \beta p(\bar{u})=\beta \frac{\Delta \bar{u}}{\omega} = \left(\frac{1}{2}+\alpha\right)\bar{u} + \frac{1}{\omega}\int_{\Omega}K(\theta, \cdot)\bar{u}(\theta){\,{\rm d}\theta}-\frac{1}{2}g \,.\tag{32}\] We thus get \(\bar{u}+A(\bar{u})=g\). We can conclude that \(A\) is maximal.

In order to prove that \(A\) is self-adjoint, we only need to establish that \(A\) is symmetric, see [35]. Let \(u, v\in D(A)\). Then, \[\int_{\Omega}p(u)v\omega=-\int_{\Omega}\nabla u \cdot \nabla v = \int_{\Omega}p(v)u\omega\,.\] We deduce therefrom that \(\langle Au, v \rangle_{L^{2}_\omega(\Omega)}= \langle u, Av \rangle_{L^{2}_\omega(\Omega)}\). The proof is complete. ◻

In the proof of Proposition 6, we exploit two important results related to gradient flows in Hilbert spaces: the Hille–Yosida theorem on the one hand, and the Brézis–Komura theorem on the other hand. The latter is well-adapted to lower semicontinuous and \(\lambda\)-convex functionals. We have already checked that \({\mathscr{F}}_{\alpha, \beta}\) is \(2\alpha\)-convex on \(L^{2}_\omega(\Omega)\). We now verify that it is lower semicontinuous.

Lemma 8. The functional \({\mathscr{F}}_{\alpha,\beta}\) is lower semicontinuous on \(L^{2}_\omega(\Omega)\).

Proof. Let \((u_j)_{j\geq 1}\) be a sequence on \(L^{2}_\omega(\Omega)\) that converges to some \(u\in L^{2}_{\omega}(\Omega)\). We claim that \[\label{eq1374} \liminf_{j\to +\infty}{\mathscr{F}}_{\alpha, \beta}(u_j)\geq {\mathscr{F}}_{\alpha, \beta}(u).\tag{33}\] We can assume without loss of generality that \(({\mathscr{F}}_{\alpha, \beta}(u_j))_{j\geq 1}\) converges in \({\mathbb{R}}\) and that each \(u_j\) belongs to \({\mathcal{W}}\). By coercivity of \({\mathscr{F}}_{\alpha,\beta}\), this implies that \((u_j)_{j\geq 1}\) is bounded in \({\mathcal{W}}\). Hence, one can extract a subsequence (we do not relabel) that converges weakly in the Hilbert space \({\mathcal{W}}\) to some \(\tilde{u}\). By uniqueness of the limit in \(L^{2}_\omega(\Omega)\), one has \(u=\tilde{u}\). Since \({\mathscr{F}}_{\alpha, \beta}\) is convex and lower-semicontinuous on \({\mathcal{W}}\), it is also weakly sequentially lower semicontinuous on \({\mathcal{W}}\) and 33 follows. ◻

Now that all their assumptions have been verified, it remains to apply the Hille–Yosida theorem and the Brézis–Komura theorem to obtain Proposition 6.

Proof of Proposition 6. Let \(u_0\in L^{2}_{\omega}(\Omega)\). Let \(v\) be the unique minimizer of \({\mathscr{J}}_{\alpha, \beta, 2Q/\omega}\) given by Lemma 4. Then, \(v\in D(A)\) and \(Av=2Q/\omega\). From [35] and Lemma 7, we deduce that there exists a unique \(\bar{u}\in C^{0}([0,\infty[;L^{2}_\omega(\Omega))\cap C^1((0,\infty);L^{2}_\omega(\Omega))\cap C^0((0,\infty);D(A))\) such that \[\begin{cases} \frac{d \bar{u}}{dt}+A\bar{u}=0\,, \\ \bar{u}(0)=u_0-v\,. \end{cases}\] Moreover, \(\|\bar{u}(t)\|_{L^{2}_\omega(\Omega)}\leq \|u_0-v\|_{L^{2}_\omega(\Omega)}\), \(\|\frac{d\bar{u}}{dt}(t)\|_{L^{2}_\omega(\Omega)}=\|A\bar{u}(t)\|_{L^{2}_\omega(\Omega)}\leq \frac{1}{t}\|u_0-v\|_{L^{2}_\omega(\Omega)}\).

We then set \(u:=\bar{u}+v\). Then, \(u\) belongs to \(C^{0}([0,\infty[;L^{2}_\omega(\Omega))\cap C^1((0,\infty);L^{2}_\omega(\Omega))\cap C^0((0,\infty);D(A))\) and satisfies

\[\label{eq1307} \begin{cases} \frac{du}{dt}= \frac{d\bar{u}}{dt} = -A\bar{u} =-A(u-v)=-Au+2\frac{Q}{\omega}\,,\\ u(0)=u_0\,. \end{cases}\tag{34}\]

Using the fact that the unbounded linear operator \(A\) is derived from the functional \({\mathscr{F}}_{\alpha, \beta}\), which is lower semicontinuous (see Lemma 8) and \(2\alpha\)–convex on \(L^{2}_\omega(\Omega)\) (by Theorem 3), we can be more specific on the convergence of the gradient flow to the unique minimum \(u^*\) of \({\mathscr{F}}_{\alpha, \beta}\). More specifically, by smoothness of \({\mathscr{F}}_{\alpha, \beta}\) when restricted to \({\mathcal{W}}\), the domain of the convex subdifferential \(\partial {\mathscr{F}}_{\alpha, \beta}\) coincides with \(D(A)\) and for every \(u\in {\mathcal{W}}\), \[\nabla {\mathscr{F}}_{\alpha, \beta}(u)=Au-2\tfrac{Q}{\omega}\,.\] The Brézis–Komura theorem, see e.g. [25], then states that the gradient flow given in 34 has the following additional properties: there exists a continuous semigroup of contractions \((S_t)_{t\geq 0}\) such that \[u(t)=S_t u(0),\] \[\label{eq1502} \forall v_1, v_2\in L^{2}_\omega(\Omega), \qquad \|S_t v_1 -S_t v_2\|_{L^{2}_\omega(\Omega)} \leq e^{-2\alpha t}\|v_1-v_2\|_{L^{2}_\omega(\Omega)}\,.\tag{35}\] Moreover, \(t\mapsto e^{2\alpha t}\|u'(t)\|_{L^{2}_\omega(\Omega)}\) is nonincreasing on \((0,\infty)\) and \[{\mathscr{F}}_{\alpha, \beta}(u(t))\leq \inf_{v\in {\mathcal{W}}}\left({\mathscr{F}}_{\alpha, \beta}(v)+\frac{\alpha}{e^{2\alpha t}-1} \|u(0)-v\|^{2}_{L^{2}_\omega(\Omega)}\right)\,.\] ◻

Proof of Corollary 3. Let \(u_*\) be the minimizer of \({\mathscr{F}}_{\alpha, \beta}\). Then, Lemma 4 with \(g=2Q/\omega\) implies that \(u_*\in D(A)\) and \(A(u_*)=2Q/\omega\). Hence, the gradient flow associated to the initial condition \(u_*\) is the constant map \(t\mapsto u_*\). For every \(v\in L^{2}_\omega(\Omega)\), the estimate 35 applied to \(v_1=v\) and \(v_2=u_*\) implies that \[\|S_tv-u_*\|_{L^{2}_\omega(\Omega)}=\|S_t v - S_t u_*\|_{L^{2}_\omega(\Omega)}\leq e^{-2\alpha t}\|v-u_*\|_{L^{2}_\omega(\Omega)}\,.\] ◻

We conclude this section by an interesting estimate on the approximation of the solution of a gradient flow by the solution of the corresponding implicit Euler scheme.

Theorem 11 (Theorem 12.5 in [25]). Let \(H\) be a Hilbert space, let \(\widetilde{\mathscr{F}}:H\to [0,\infty]\) be convex and lower semicontinuous, and let \(\varphi \in \mathrm{Dom}(\widetilde{\mathscr{F}})\). Given \(\tau>0\), we define \[\label{eq:jko-appendix} \tilde{\varrho}^0 = \varphi, \qquad \tilde{\varrho}^{k\tau} = \arg\min_{\psi \in H} \left\{ \widetilde{\mathscr{F}}(\psi) + \tfrac{1}{2\tau} \|\psi - \tilde{\varrho}^{(k-1)\tau}\|_{H}^2 \right\}\,.\tag{36}\] We also consider the piecewise constant left-continuous interpolation of \((\tilde{\varrho}^{k\tau})_{k\geq 0}\): \[\label{eq-piecewie-interpolation} \tilde{\varrho}_{\tau}(t):= \begin{cases} \varphi & \textrm{ if } t=0\,,\\ \tilde{\varrho}^{k\tau} & \textrm{ if } k\geq 1 \textrm{ and }t\in ((k-1)\tau, k\tau]\,. \end{cases}\tag{37}\] Then:

  • for every \(t\geq 0\), the family \((\tilde{\varrho}_\tau(t))_{\tau>0}\) is Cauchy as \(\tau \to 0\), so that its limit \(\varrho(t)\) exists ,

  • the curve \(t\mapsto \varrho(t)\) is the gradient flow starting from \(\varphi\) ,

  • \(\|\tilde{\varrho}_\tau(t) - \varrho(t)\|_{H} \le 2(\sqrt{2}+1) \sqrt{\tau \, \widetilde{\mathscr{F}}(\varphi)}\) for every \(\tau>0\) and every \(t\geq 0\) .

5.5 Numerical simulations↩︎

In this section, we provide more detailed explanations and derivations accompanying Section 3.

Derivation of quadratic form↩︎

In this subsection we derive the approximating functional \(\widehat{\mathscr{F}}_{\alpha, \beta}\). We denote \(u \approx \widehat{u} = \sum_{i=1}^{M} a_i \widehat{u}_i\). We consider a given dataset \(\{(x_i, f(x_i))\}_{i=1}^{N_D}\). We start with the approximating risk function: \[\begin{align} \mathscr{R} (\vec{f}, \vec{a}) &= C_D\sum_{j=1}^{N_D} \Big(f(x_j) - \sum_{i=1}^{M} a_i \int_\Omega h(\theta, x_j) \,\widehat{u}_i(\theta) {\,{\rm d}}\theta\Big)^2 = C_D\lvert \vec{f} - U \vec{a} \rvert^2\\ &= C_D\left(\vert \vec{f}\,\rvert^2 - 2\vec{f} \,^{\top} U \vec{a} + \lvert U \vec{a}\rvert^2\right) \, , \end{align}\] where \(C_D=\mathcal{L}^d(D)/N_D\) (here, \(\mathcal{L}^d(D)\) is the Lebesgue measure of \(D\)) and \[U_{ki} = \int_\Omega h(\theta, x_k) \,\widehat{u}_i(\theta) {\,{\rm d}}\theta \, .\] Now we derive the weighted Lebesgue norm \(\lVert\widehat u\rVert^2_{L^2_\omega(\Omega)}\): \[\begin{align} \lVert\widehat u&\rVert^2_{L_\omega(\Omega)} = \Big\lVert \sum_{i=1}^{M} a_u \widehat{u}_i\Big\rVert^2_{L^2_\omega(\Omega)} = \sum_{i,j}^{M} a_i a_j \langle \widehat{u}_i, \widehat{u}_j \rangle_{L^2_\omega (\Omega)} = \vec{a}^{\top} V \, \vec{a} \, ,\quad V_{ik} = \langle \widehat{u}_i, \widehat{u}_j \rangle _{L^2_\omega(\Omega)}\, . \end{align}\] And finally the gradient norm \(\|\nabla \widehat{u}\|_{L^2(\Omega)}^2\): \[\begin{align} \|\nabla \widehat{u}\|_{L^2(\Omega)}^2 &= \sum_{\ell=1}^{d+1} \| \partial_\ell u \|_{L^2(\Omega)}^2 = \sum_{\ell=1}^{d+1} \Big\lVert \sum_{i=1}^{M} a_i \partial_\ell \widehat{ u}_i \Big\rVert_{L^2(\Omega)}^2 = \sum_{d=1}^{\ell+1} \sum_{i,j = 1}^{M} a_i a_j \langle \partial_\ell\widehat{u}_i, \partial_\ell\widehat{u}_j \rangle_{L^2 (\Omega)} \\ & = \vec{a} \cdot W \vec{a} \, , \quad W_{ij} = \langle \nabla \widehat{u}_i, \nabla \widehat{u}_j \rangle_{L^{2}(\Omega)} \, . \end{align}\]

Basis functions↩︎

We have used the following type of basis functions \(\{\widehat {u}_i \}_{i=1}^{M}\):

  1. Polynomials on \(\Omega = B_{R}^d \times (-L,L)\) (where for \(d \geq 1\) the notation \(B_{R}^d\) refers to the ball of radius \(R\) and center \(0\) in \({\mathbb{R}}^d\)): \[\label{eq:polynomial} \widehat{u}_i(\theta) = \prod_{j=1}^{d+1} \theta_j^{p^{i}_j}\tag{38}\]

  2. Trigonometric orthonormal basis on \(\Omega = (-R,R)\times (-L,L)\) (for \(d=1\)): \[\label{eq:harmonic} \widehat{u}_i (\theta) = c_1 c_2 \cos \Big(\tfrac{p_1^i \pi (\theta_0 + L)}{2L} \Big)\cos \Big(\tfrac{p_2^i \pi (\theta_1 + R)}{2R} \Big)\tag{39}\] with \[c_1 = \textstyle\begin{cases} 1/\sqrt{2L} & p_1^i >0 \,,\\ 1/\sqrt{L} & p_1^i =0\,, \end{cases} \qquad \qquad c_2 = \textstyle\begin{cases} 1/\sqrt{2R} & p_2^i >0\,, \\ 1/\sqrt{R} & p_2^i =0\,. \end{cases}\] These functions are the solutions of the eigenvalue problem for the Laplacian: \(\Delta \widehat{u}_i = \lambda\widehat{u}_i\) with zero Neumann boundary condition.

Number of basis functions \(M\)↩︎

Given \(s\in \mathbb{N}_0\), we consider the set of polynomials of degree not larger than \(s\) on the space \(\Omega \subseteq {\mathbb{R}}^{d+1}\). The exponents of the monomials involve all possible tuples \(p^i=(p^i_j)_{1\leq j\leq d+1}\) such that \(\sum_{j=1}^{d+1} p^i_j \leq s\). The number \(M\) of such tuples can be computed as follows.

We define an additional element of each tuple variable \(p^i_{d+2} := s - \sum_{j=1}^{d+1} p^i_j\) and thus the condition \(\sum_{j=1}^{d+1} p^i_j \leq s\) is equivalent to \(p^i_{d+2} \ge 0\). We obtain \(\sum_{j=1}^{d+2} p^i_j = s\) with all variables satisfying \(p^i_j \in \mathbb{N}_{0}\). So the number of all such possible tuples is given by the stars and bars formula: \[\textstyle \binom{s + (d+2) - 1}{(d+2) - 1} = \binom{s + d + 1}{d + 1} \,.\]

Additional examples↩︎

Example 2’: Discontinuous target — sign function (\(d=1\))

We consider the same dataset as in Example 2. We approximate by polynomial functions \(\{u_i\}_{i=1}^{M}\). The minimum of \({\mathscr{F}}^{(f)}_{\alpha, \beta}\) is approximated by a polynomial 38 of order \(s=60\), thus \(M = 1{,}891\). We have chosen \(\Omega=(-R,R)\times (-L,L)\) with \(R=L=1.5\). For the regularized functional, we have taken \(\alpha = 3.2\times 10^{-6}\) and \(\beta = 2 \times 10^{-5}\). See the results on Figure 2. It is worth to emphasize that the smooth minimizer of the regularized functional does not reflect the outlier.

Figure 2: Example 2’

Example 5: Sinus of norm value for (\(d=2\))

We generate \(N_D = 100\) observations on the domain \((-1,1)^2\) of the function \(f=\sin(3\lvert x\vert)\). We approximate the minimum by polynomials 38 of order \(s=5\), so that \(M=112\). We have chosen \(R=L=5\) and \(\alpha = 4\times 10^{-9}\) and \(\beta=4\times 10^{-8}\). We note that the sparsity of the matrix \(U\) is \(\approx10\,\%\). See the result of approximation by the regularized functional \(\widehat{\mathscr{F}}_{\alpha,\beta}^{\;(f)}\) on Figure 3.

Figure 3: Example 5

References↩︎

[1]
Song Mei, Andrea Montanari, and Phan-Minh Nguyen. A mean field view of the landscape of two-layer neural networks. Proceedings of the National Academy of Sciences, 115 (33): E7665–E7671, 2018.
[2]
Lénaïc Chizat and Francis Bach. On the global convergence of gradient descent for over-parameterized models using optimal transport. In Advances in Neural Information Processing Systems, 2018.
[3]
Xavier Fernández-Real and Alessio Figalli. The continuous formulation of shallow neural networks as Wasserstein-type gradient flows. In Analysis at large—dedicated to the life and work of Jean Bourgain, pages 29–57. Springer, Cham, 2022. ISBN 978-3-031-05330-6; 978-3-031-05331-3. . URL https://doi.org/10.1007/978-3-031-05331-3_3.
[4]
Justin Sirignano and Konstantinos Spiliopoulos. Mean field analysis of neural networks: A law of large numbers. SIAM Journal on Applied Mathematics, 80 (2): 725–752, 2020.
[5]
Grant M. Rotskoff and Eric Vanden-Eijnden. Trainability and accuracy of neural networks: An interacting particle system approach. Communications on Pure and Applied Mathematics, 75 (9): 1889–1935, 2022.
[6]
Weinan E, Jiequn Han, and Qianxiao Li. A mean-field optimal control formulation of deep learning. Research in the Mathematical Sciences, 6 (1), December 2018. ISSN 2197-9847. . URL http://dx.doi.org/10.1007/s40687-018-0172-y.
[7]
Weinan E, Chao Ma, and Lei Wu. Machine learning from a continuous viewpoint, I. Science China Mathematics, 63 (11): 2233–2266, Nov 2020. ISSN 1869-1862. . URL https://doi.org/10.1007/s11425-020-1773-8.
[8]
Arthur Jacot, Franck Gabriel, and Clément Hongler. Neural tangent kernel: Convergence and generalization in neural networks. In Advances in Neural Information Processing Systems, 2018.
[9]
Simon S. Du, Jason D. Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai. Gradient descent finds global minima of deep neural networks. In Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 1675–1685, 2019.
[10]
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song. A convergence theory for deep learning via over-parameterization. In Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 242–252, 2019.
[11]
Yihang Chen, Fanghui Liu, Yiping Lu, Grigorios G. Chrysos, and Volkan Cevher. Generalization of scaled deep resnets in the mean-field regime. arXiv preprint arXiv:2403.09889, 2024.
[12]
Atsushi Nitanda. Improved particle approximation error for mean field neural networks. In Advances in Neural Information Processing Systems, 2024.
[13]
Alireza Mousavi-Hosseini, Denny Wu, and Murat A. Erdogdu. Learning multi-index models with neural networks via mean-field Langevin dynamics. In International Conference on Learning Representations, 2025.
[14]
Steffen Dereich, Arnulf Jentzen, and Sebastian Kassing. On the existence of minimizers in shallow residual ReLU neural network optimization landscapes. SIAM Journal on Numerical Analysis, 62 (6): 2640–2666, 2024. .
[15]
Tong Mao, Jonathan W. Siegel, and Jinchao Xu. Approximation by shallow neural networks with ReLU\(^k\) activation: Sobolev spaces and optimal rates via the Radon transform. SIAM Journal on Mathematical Analysis, 58 (2): 1171–1186, 2026. . URL https://doi.org/10.1137/24M1686693.
[16]
Jong Kwon Oh, Hanbaek Lyu, and Hwijae Son. Sobolev acceleration for neural networks. arXiv preprint arXiv:2509.19773, 2025.
[17]
Richard Jordan, David Kinderlehrer, and Felix Otto. The variational formulation of the Fokker-Planck equation. SIAM J. Math. Anal., 29 (1): 1–17, 1998. ISSN 0036-1410,1095-7154. . URL https://doi.org/10.1137/S0036141096303359.
[18]
Luigi Ambrosio, Nicola Gigli, and Giuseppe Savaré. Gradient Flows: In Metric Spaces and in the Space of Probability Measures. Birkhäuser, 2008.
[19]
Filippo Santambrogio. {Euclidean, metric, and Wasserstein} gradient flows: an overview. Bull. Math. Sci., 7 (1): 87–154, 2017. ISSN 1664-3607,1664-3615. . URL https://doi.org/10.1007/s13373-017-0101-1.
[20]
Anna Kh. Balci, Lars Diening, and Mikhail Surnachev. New examples on Lavrentiev gap using fractals. Calc. Var. Partial Differential Equations, 59 (5): Paper No. 180, 34, 2020. ISSN 0944-2669,1432-0835. . URL https://doi.org/10.1007/s00526-020-01818-1.
[21]
Irene Fonseca, Jan Malý, and Giuseppe Mingione. Scalar minimizers with fractal singular sets. Arch. Ration. Mech. Anal., 172 (2): 295–307, 2004. ISSN 0003-9527,1432-0673. . URL https://doi.org/10.1007/s00205-003-0301-6.
[22]
Michał Borowski, Iwona Chlebicka, Filomena De Filippis, and Błażej Miasojedow. Absence and presence of Lavrentiev’s phenomenon for double phase functionals upon every choice of exponents. Calc. Var. Partial Differential Equations, 63 (2): Paper No. 35, 23, 2024. ISSN 0944-2669,1432-0835. . URL https://doi.org/10.1007/s00526-023-02640-1.
[23]
Giuseppe Mingione and Vicenţiu Rǎdulescu. Recent developments in problems with nonstandard growth and nonuniform ellipticity. J. Math. Anal. Appl., 501 (1): Paper No. 125197, 41, 2021. ISSN 0022-247X,1096-0813. . URL https://doi.org/10.1016/j.jmaa.2021.125197.
[24]
Lénaïc Chizat, Maria Colombo, and Xavier Fernández-Real. Convergence of drift-diffusion pdes arising as wasserstein gradient flows of convex functions, 2025. URL https://arxiv.org/abs/2507.12385.
[25]
Luigi Ambrosio, Elia Brué, and Daniele Semola. Lectures on optimal transport, volume 130 of Unitext. Springer, Cham, 2021. ISBN 978-3-030-72161-9; 978-3-030-72162-6. . URL https://doi.org/10.1007/978-3-030-72162-6. La Matematica per il 3+2.
[26]
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal. Reconciling modern machine-learning practice and the classical bias–variance trade-off. Proceedings of the National Academy of Sciences, 116 (32): 15849–15854, 2019. .
[27]
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo. Implicit regularization in deep matrix factorization. Advances in Neural Information Processing Systems, 32, 2019.
[28]
Noam Razin and Nadav Cohen. Implicit regularization in deep learning may not be explainable by norms. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 21174–21187. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper_files/paper/2020/file/f21e255f89e0f258accbe4e984eef486-Paper.pdf.
[29]
Francis Bach. Breaking the curse of dimensionality with convex neural networks. Journal of Machine Learning Research, 18 (19): 1–53, 2017.
[30]
Greg Ongie, Rebecca Willett, Daniel Soudry, and Nathan Srebro. A function space view of bounded norm infinite width relu nets: The multivariate case. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. URL https://openreview.net/forum?id=H1lNPxHKDH.
[31]
Diederik Kingma and Jimmy Ba. Adam: A method for stochastic optimization. International Conference on Learning Representations, 12 2014.
[32]
Bradley Efron, Trevor Hastie, Iain Johnstone, and Robert Tibshirani. Least angle regression. Ann. Statist., 32 (2): 407–499, 2004. ISSN 0090-5364,2168-8966. . URL https://doi.org/10.1214/009053604000000067. With discussion, and a rejoinder by the authors.
[33]
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, et al. Scikit-learn: Machine learning in python. Journal of Machine Learning Research, 12: 2825–2830, 2011.
[34]
David Gilbarg and Neil S. Trudinger. Elliptic partial differential equations of second order. Class. Math. Berlin: Springer, reprint of the 1998 ed. edition, 2001. ISBN 3-540-41160-7.
[35]
Haïm Brézis. Functional Analysis, Sobolev Spaces and Partial Differential Equations. Universitext. Springer, New York, 2011. ISBN 978-0-387-70913-0.