July 11, 2026
It is well known that the study of the Kolmogorov widths of a function class, which is the image of the unit ball of the \(L_q\) space of an integral operator \(J_K\) with the kernel \(K\), is closely connected with the study of sparse approximations of the kernel \(K\) with respect to the classical bilinear dictionary. Recently, it was discovered that if instead of the Kolmogorov widths we study the errors of optimal linear sampling recovery of the same classes, then we need to study sparse approximations of the kernel \(K\) with respect to an adaptive dictionary, which is determined by the kernel \(K\). In this paper we study this important problem of nonlinear approximation with respect to an adaptive dictionary. Also, in this paper we continue to develop the following general approach, which is related to the above nonlinear approximation problem. We study asymptotic behavior of the errors of sampling recovery not for an individual smoothness class, how it is usually done, but for the collection of classes, which are defined by integral operators with kernels coming from a given class of functions. Earlier, such approach was realized for the Kolmogorov widths and very recently for the entropy numbers.
This paper is a followup to the recent author’s paper [1]. In this paper we continue to study approximation of the multivariate functions \(K(\mathbf{x},\mathbf{y})\), \(\mathbf{x}=(x_1,\dots,x_d)\), \(\mathbf{y}=(y_1,\dots,y_d)\) by linear combinations of functions of the form \(u(\mathbf{x})v(\mathbf{y})\). In the case, when we can choose arbitrary functions \(u\) and \(v\), it is a classical problem of best bilinear approximation. In the paper [1] it was pointed out that the problem of optimal linear recovery on function classes defined by an integral operator with the kernel \(K(\mathbf{x},\mathbf{y})\) is closely related to the problem of approximation of \(K\) by linear combinations of functions of the form \(u(\mathbf{x})v(\mathbf{y})\) with functions \(v(\mathbf{y})\) defined by the kernel \(K(\mathbf{x},\mathbf{y})\), namely, \(v(\mathbf{y}) = K(\mathbf{z},\mathbf{y})\) with some \(\mathbf{z}\). This means that we approximate \(K(\mathbf{x},\mathbf{y})\) with respect to a dictionary, which is determined by the function \(K(\mathbf{x},\mathbf{y})\) itself. We call such a process – approximation with adaptive dictionaries (see below for more details). In the paper [1] mostly the case \(d=1\), i.e. the case of functions \(K(x,y)\) of two variables, was studied. In this paper we focus on the general case \(d\ge 1\).
We now proceed to the detailed presentation. Let \((\Omega,\mu)\) be a probability space. By the \(L_p\), \(1\le p< \infty\), norm we understand \[\|f\|_p:=\|f\|_{L_p(\Omega,\mu)} := \left(\int_\Omega |f|^p\,d\mu\right)^{1/p}.\] By the \(L_\infty\)-norm we understand the uniform norm of continuous functions \[\|f\|_\infty := \sup_{\omega\in\Omega} |f(\omega)|\] and with some abuse of notation we occasionally write \(L_\infty(\Omega)\) for the space \({\mathcal{C}}(\Omega)\) of continuous functions on \(\Omega\). We define the vector \(L_{\mathbf{p}}\)-norm, \(\mathbf{p}=(p_1,\dots,p_v)\), of functions of \(v\) variables \(\mathbf{x}=(x_1,\dots,x_v)\) as \[\|f(\mathbf{x})\|_{\mathbf{p}} :=\|f(\mathbf{x})\|_{(p_1,\dots,p_v)} :=\|f(\mathbf{x})\|_{p_1,\dots,p_v} := \|\cdots\|f(\cdot,x_2,\dots,x_v)\|_{p_1}\cdots\|_{p_v}.\]
We now introduce some concepts from nonlinear sparse approximation.
The first example of sparse approximation with respect to redundant dictionaries was considered by E. Schmidt in [2], who studied the approximation of functions \(f(x,y)\) of two variables by bilinear forms, \[\sum_{i=1}^mu_i(x)v_i(y),\] in \(L_2([0,1]^2)\). In this case we use the following dictionary (bilinear dictionary) \[\label{Bi1} \Pi := \{u(x)v(y)\,:\, u,v\in L_2([0,1])\},\tag{1}\] where the functions \(u\) and \(v\) are functions of a single variable. This problem is closely connected with properties of the integral operator \[(J_fg)(x) := \int_0^1 f(x,y)g(y) dy\] with the kernel \(f(x,y)\). E. Schmidt ([2]) gave an expansion (known as the Schmidt expansion) \[\label{Bi2} f(x,y) = \sum_{j=1}^\infty s_j(J_f) \phi_j(x)\psi_j(y),\tag{2}\] where \(\{s_j(J_f)\}\) is a nonincreasing sequence of singular numbers of \(J_f\), i.e. \(s_j(J_f) := \lambda_j(J^*_fJ_f)^{1/2}\), where \(\{\lambda_j(A)\}\) is a sequence of eigenvalues of an operator \(A\), and \(J^*_f\) is the adjoint operator to \(J_f\). The two sequences \(\{\phi_j(x)\}\) and \(\{\psi_j(y)\}\) form orthonormal sequences of eigenfunctions of the operators \(J_fJ_f^*\) and \(J_f^*J_f\), respectively. He also proved that \[\left\|f(x,y) -\sum_{j=1}^m s_j(J_f) \phi_j(x)\psi_j(y)\right\|_{L_2}\] \[\label{Bi3} = \inf_{u_j,v_j\in L_2, \quad j=1,\dots,m}\left\|f(x,y) -\sum_{j=1}^m u_j(x)v_j(y)\right\|_{L_2}.\tag{3}\] The reader can find a detailed discussion of this connection in [3], Ch.2.
In a general setting we are working in a Banach space \(X\) with a redundant system of elements \({\mathcal{D}}\) (dictionary \({\mathcal{D}}\)). An element (function, signal) \(h\in X\) is said to be \(m\)-sparse with respect to \({\mathcal{D}}\) if it has a representation \(h=\sum_{i=1}^mc_ig_i\), \(g_i\in {\mathcal{D}}\), \(i=1,\dots,m\), where \(\{c_i\}\) are real or complex numbers. The set of all \(m\)-sparse elements is denoted by \(\Sigma_m({\mathcal{D}})\). For a given element \(f\) we introduce the error of best \(m\)-term approximation \[\sigma_m(f,{\mathcal{D}})_X := \inf_{h\in\Sigma_m({\mathcal{D}})} \|f-h\|_X.\] We now make a comment on terminology. In the greedy approximation literature we define a dictionary \({\mathcal{D}}\) as a system \(\{g\}\) of elements \(g\in X\) with the following two properties \[\|g\|_X\le 1 \quad\text{for all} \quad g\in {\mathcal{D}}\quad \text{and the closure of}\, \operatorname{span}({\mathcal{D}}) =X.\] The normalization condition \(\|g\|_X\le 1\) is imposed for convenience. Clearly, the characteristic \(\sigma_m(f,{\mathcal{D}})_X\) does not depend on normalization. In this paper we mostly use this characteristic. Let us discuss the second condition. Suppose that a system \({\mathcal{S}}\subset X\) does not satisfy this condition. Then, instead of the Banach space \(X\) we consider a subspace \(X_{\mathcal{S}}\) of \(X\), which is the closure (in \(X\)) of \(\operatorname{span}({\mathcal{S}})\). This makes the system \({\mathcal{S}}\) to be a dictionary in the Banach space \(X_{\mathcal{S}}\). For this reason, we sometimes with a little abuse of exactness freely use both terms system and dictionary for a general system. In the greedy approximation theory there are theorems, which guarantee convergence of certain greedy algorithms with respect to any dictionary \({\mathcal{D}}\) for any element \(f\in X\). Clearly, in the case, when we deal with a system, we can only apply those theorems to \(f\in X_{\mathcal{S}}\).
We stress that the bilinear dictionary \(\Pi\) does not depend on a function under approximation. In this sense it is not adaptive – we use it for approximation of all functions. It turns out that in some problems we need to study approximation of a given kernel \(K(\mathbf{x},\mathbf{y})\) with respect to a dictionary, which is determined by \(K\). We now give a more general definition of the bilinear dictionary (system) and define three adaptive systems. Let \(\mathbf{p}=(p_1,p_2)\), \(1\le p_1,p_2 \le \infty\) be given. In the case \(\mathbf{p}=(p,p)\) for brevity we write \(p\) instead of \(p,p\) in the notations. Sometimes, we drop \(\mathbf{p}\) from the notation.
Bilinear dictionary \(\Pi(\mathbf{p})\). Define \[\Pi(\mathbf{p}):={\mathcal{L}}{\mathcal{L}}(\mathbf{p}) :=\{g: g(\mathbf{x},\mathbf{y}) = u(\mathbf{x})v(\mathbf{y}),\, u\in L_{p_1}(\Omega^1),\, v\in L_{p_2}(\Omega^2) \}.\]
\({\mathcal{L}}{\mathcal{K}}(\mathbf{p})\)-system. Assume that \(K \in L_\mathbf{p}(\Omega^1\times \Omega^2)\) satisfies the following property. For any \(\mathbf{z}\in \Omega^1\) we have \(K(\mathbf{z},\cdot) \in L_{p_2}(\Omega^2)\). Define \[{\mathcal{L}}{\mathcal{K}}(\mathbf{p}) :=\{g: g(\mathbf{x},\mathbf{y}) = u(\mathbf{x},\mathbf{z})K(\mathbf{z},\mathbf{y}),\, \forall \mathbf{z}\in \Omega^1\,\,\text{we have}\,\, u(\cdot,\mathbf{z}) \in L_{p_1}(\Omega^1) \}.\]
\({\mathcal{K}}{\mathcal{L}}(\mathbf{p})\)-system. Assume that \(K \in L_\mathbf{p}(\Omega^1\times \Omega^2)\) satisfies the following property. For any \(\mathbf{z}\in \Omega^2\) we have \(K(\cdot,\mathbf{z}) \in L_{p_1}(\Omega^1)\). Define \[{\mathcal{K}}{\mathcal{L}}(\mathbf{p}) :=\{g: g(\mathbf{x},\mathbf{y}) = K(\mathbf{x},\mathbf{z})v(\mathbf{z},\mathbf{y}),\, \forall \mathbf{z}\in \Omega^2\,\,\text{we have}\,\, v(\mathbf{z},\cdot) \in L_{p_1}(\Omega^2) \}.\]
\({\mathcal{K}}{\mathcal{K}}(\mathbf{p})\)-system. Assume that \(K \in L_\mathbf{p}(\Omega^1\times \Omega^2)\) satisfies the following property. For any \(\mathbf{a}\in \Omega^1\) we have \(K(\mathbf{a},\cdot) \in L_{p_2}(\Omega^2)\) and for any \(\mathbf{b}\in \Omega^2\) we have \(K(\cdot,\mathbf{b}) \in L_{p_1}(\Omega^1)\). Define \[{\mathcal{K}}{\mathcal{K}}(\mathbf{p}) :=\{g: g(\mathbf{x},\mathbf{y}) = K(\mathbf{x},\mathbf{b})K(\mathbf{a},\mathbf{y}), \quad (\mathbf{a},\mathbf{b})\in \Omega^1\times \Omega^2 \}.\] Note that in the literature (see [4]) the functions \(K_{ab}(x,y) := K(x,b)K(a,y)\), \((x,y),(a,b) \in [0,1]^2\) are called cross-functions of the function \(K(x,y)\) and the system \({\mathcal{K}}{\mathcal{K}}(\infty)\) is called the system of cross-functions.
In Section 6 we prove the following general inequalities (see Theorem 20): For any \(b\in(1,2]\) there exists a positive constant \(B=B(b)\) such that for any continuous \(K\) we have \[\label{In1} \sigma_m(K,{\mathcal{L}}{\mathcal{K}}(\infty))_\infty \le Bm^{1/2} \sigma_{\theta (m-1)}(K,\Pi(\infty))_\infty\tag{4}\] with \(\theta =1/b\) in the real case and \(\theta =1/(2b)\) in the complex case. Also, we prove there that the extra factor \(m^{1/2}\) in the inequality (4 ) is sharp (see Proposition 3).
A number of results (upper bounds, lower bounds, and sometimes the right orders) are obtained in the paper [1] in the case \(d=1\) for the following setting: Estimate \[\sup_{K\in {\mathbf{F}}} \sigma_m(K,{\mathcal{L}}{\mathcal{K}}(\mathbf{p}))_\mathbf{p}\] for a certain function class \({\mathbf{F}}\).
In this paper we extend some of those results from the case \(d=1\) to the general case \(d\ge 1\). For that we apply here the same general strategy, which was used in [1]. It is a three step strategy. First, we relate \(\sigma_m(K,{\mathcal{L}}{\mathcal{K}}(\mathbf{p}))_\mathbf{p}\) to the optimal linear recovery characteristic \(\varrho_m({\mathbf{W}}^K_q,L_\mathbf{p})\). Second, we relate \(\sigma_m(K,\Pi(\mathbf{p}))_\mathbf{p}\), \(\mathbf{p}= (p,\infty)\), to the Kolmogorov width \(d_m({\mathbf{W}}^K_1,L_p)\). Third, we use known results, which provide an upper bound on \(\varrho_m({\mathbf{W}},L_p)\) in terms of the Kolmogorov width \(d_n({\mathbf{W}},L_\infty)\). Note that the first inequality of that type, namely, the inequality \[\varrho_{bn}({\mathbf{W}},L_2) \le Bd_n({\mathbf{W}},L_\infty)\] was obtained in [5] (see Theorem 2 below). Later, some generalizations of that inequality to the case \(p\in (2,\infty]\) were proved in [6] and [7]. Here we use Theorem 8.
In Section 6 we prove the following bound (see Corollary 5). Assume that we have \(r_1>1/2\), \(r_2>0\). Then (see the definition of classes \({\mathbf{W}}^\mathbf{r}_2\) in Section 2 below) \[\sup_{K\in {\mathbf{W}}^\mathbf{r}_2} \sigma_m(K,{\mathcal{L}}{\mathcal{K}})_{2} \ll m^{-r_1-r_2}(\log m)^{(d-1)(r_1+r_2)+1/2} .\]
As we already pointed out above the study of linear recovery is closely related to the Kolmogorov widths. Namely, to the Kolmogorov widths in the uniform norm \(L_\infty\). In Section 5 we focus on the case of the uniform norm and complement the results known in the case of \(L_p\), \(p\in [2,\infty)\), by the case \(p=\infty\). The following bound (see Theorem 18 below) is a step in that direction. Assume that we have \(r_1>1/2\), \(r_2>0\). Then (see the definition of classes \({\mathbf{W}}^\mathbf{r}_2\) in Section 2 below) \[\sup_{K\in {\mathbf{W}}^\mathbf{r}_2} d_m({\mathbf{W}}^K_2)_\infty \ll m^{-r_1-r_2-1/2}(\log m)^{(d-1)(r_1+r_2)+1/2} .\]
Thus, the new results of the paper are contained in Sections 5 and 6. In Section 2 we present the definitions of function classes that we discuss in the paper. In Section 3 we formulate results on inequalities between the error of optimal linear sampling recovery and the Kolmogorov widths. In Section 4 we collect some of the known results on the Kolmogorov widths and their relation to the sparse approximation with respect to the bilinear dictionary \(\Pi\).
We begin with the definition of classes \({\mathbf{W}}^\mathbf{a}_\mathbf{q}\) (see, for instance, [8], p.31, in the case of scalar \(q\)).
Definition 1. In the univariate case, for \(a>0\), let \[\label{Bi8} F_a(x):= 1+2\sum_{k=1}^\infty k^{-a}\cos (kx-a\pi/2)\qquad{(1)}\] be the Bernoulli kernel and in the multivariate case, for \(\mathbf{a}=(a_1,\dots,a_v) \in {\mathbb{R}}^v_+\), \(\mathbf{x}=(x_1,\dots,x_v)\in {\mathbb{T}}^v\), let \[\label{Bi8m} F_\mathbf{a}(\mathbf{x}) := \prod_{j=1}^v F_{a_j}(x_j).\qquad{(2)}\] Denote for \(\mathbf{1}\le \mathbf{q}\le \infty\) (we understand the vector inequality coordinate wise) \[{\mathbf{W}}^\mathbf{a}_\mathbf{q}:= \{f:f=\varphi\ast F_\mathbf{a},\quad \|\varphi\|_\mathbf{q}\le 1\},\] where \[( F_\mathbf{a}\ast \varphi)(\mathbf{x}):= (2\pi)^{-v}\int_{{\mathbb{T}}^v} F_\mathbf{a}(\mathbf{x}-\mathbf{y}) \varphi(\mathbf{y})d\mathbf{y},\quad {\mathbb{T}}^v := [0,2\pi)^v.\]
The classes \({\mathbf{W}}^\mathbf{a}_\mathbf{q}\) are classical classes of functions with dominating mixed derivative (Sobolev-type classes of functions with mixed smoothness).
We now proceed to the definition of the classes \({\mathbf{H}}^\mathbf{a}_\mathbf{q}:= {\mathbf{H}}^{\mathbf{a},v}_\mathbf{q}\) of periodic functions of \(v\) variables, which is based on the mixed differences (see, for instance, [8], p.31, in the case of scalar \(q\)).
Definition 2. Let \(\mathbf{t}=(t_1,\dots,t_v)\) and \(\Delta_{\mathbf{t}}^l f(\mathbf{x})\) be the mixed \(l\)-th difference with step \(t_j\) in the variable \(x_j\), that is \[\Delta_{\mathbf{t}}^l f(\mathbf{x}) :=\Delta_{t_v,v}^l\cdots\Delta_{t_1,1}^l f(x_1,\dots ,x_v) .\] Let \(e\) be a subset of natural numbers in \([1,v]\). We denote \[\Delta_{\mathbf{t}}^l (e) :=\prod_{j\in e}\Delta_{t_j,j}^l,\qquad \Delta_{\mathbf{t}}^l (\varnothing) := Id \,-\, \text{identity operator}.\] We define the class \({\mathbf{H}}_{\mathbf{q},l}^\mathbf{a}B\), \(l > \|\mathbf{a}\|_\infty\), as the set of \(f\in L_\mathbf{q}({\mathbb{T}}^v)\) such that for any \(e\) \[\label{Bi9} \bigl\|\Delta_{\mathbf{t}}^l(e)f(\mathbf{x})\bigr\|_\mathbf{q}\le B \prod_{j\in e} |t_j |^{a_j} .\qquad{(3)}\] In the case \(B=1\) we omit it. It is known (see Theorem 1 below) that the classes \({\mathbf{H}}^\mathbf{a}_{\mathbf{q},l}\) with different \(l>\|\mathbf{a}\|_\infty\) are equivalent. So, for convenience we omit \(l\) from the notation.
We now formulate a result, which gives an equivalent description of classes \({\mathbf{H}}^\mathbf{a}_{\mathbf{q},l}\). We need some classical trigonometric polynomials. The univariate Fejér kernel of order \(j - 1\): \[\mathcal{K}_{j} (x) := \sum_{|k|\le j} \bigl(1 - |k|/j\bigr) e^{ikx} =\frac{(\sin (jx/2))^2}{j (\sin (x/2))^2}.\] The Fejér kernel is an even nonnegative trigonometric polynomial of order \(j-1\). It satisfies the obvious relations \[\label{FKm} \| \mathcal{K}_{j} \|_1 = 1, \qquad \| \mathcal{K}_{j} \|_{\infty} = j.\tag{5}\] Let \({\mathcal{K}}_\mathbf{j}(\mathbf{x}):= \prod_{i=1}^v {\mathcal{K}}_{j_i}(x_i)\) be the \(v\)-variate Fejér kernels for \(\mathbf{j}= (j_1,\dots,j_d)\) and \(\mathbf{x}=(x_1,\dots,x_v)\).
The univariate de la Vallée Poussin kernels are defined as follows \[{\mathcal{V}}_m := 2{\mathcal{K}}_{2m} - {\mathcal{K}}_m.\] We also need the following special trigonometric polynomials. Let \(s\) be a nonnegative integer. We define \[\mathcal{A}_0 (x) := 1, \quad \mathcal{A}_1 (x) := \mathcal{V}_1 (x) - 1, \quad \mathcal{A}_s (x) := \mathcal{V}_{2^{s-1}} (x) -\mathcal{V}_{2^{s-2}} (x), \quad s\ge 2,\] where \(\mathcal{V}_m\) are the de la Vallée Poussin kernels defined above. For \(\mathbf{s}=(s_1,\dots,s_v)\in {\mathbb{N}}^v_0\) define \[{\mathcal{A}}_\mathbf{s}(\mathbf{x}) := \prod_{j=1}^v {\mathcal{A}}_{s_j}(x_j),\qquad \mathbf{x}=(x_1,\dots,x_v)\] and \[A_\mathbf{s}(f) := {\mathcal{A}}_\mathbf{s}\ast f.\]
The following result is known (see, for instance, [8], p.32, for the scalar \(q\) and [9] for the vector \(\mathbf{q}\)).
Theorem 1. Let \(f\in {\mathbf{H}}^\mathbf{a}_{\mathbf{q},l}\), \(\mathbf{1} \le \mathbf{q}\le \infty\). Then, for \(\mathbf{s}\ge \mathbf{0}\) \[\label{H1} \|A_\mathbf{s}(f)\|_\mathbf{q}\le C(\mathbf{a},v,l)2^{-(\mathbf{a},\mathbf{s})}.\qquad{(4)}\] Conversely, from (?? ) it follows that there exists \(B>0\), which does not depend on \(f\), such that \(f\in {\mathbf{H}}^\mathbf{a}_{\mathbf{q},l}B\).
The reader can find results on approximation properties of these classes in the books [8], [10], and [11].
Notations for the function classes. In this paper we consider the case, when \(v=2d\), \(d\in {\mathbb{N}}\), \(\mathbf{1}\le \mathbf{q}\le \infty\), and \(\mathbf{a}\) has a special form: \(a_j = r_1\), \(a_{j+d}= r_2\) for \(j=1,\dots,d\). In this case we write \({\mathbf{W}}^{\mathbf{r}}_\mathbf{q}= {\mathbf{W}}^{(\mathbf{r}^1,\mathbf{r}^2)}_\mathbf{q}= {\mathbf{W}}^{r_1,r_2}_\mathbf{q}\) and \({\mathbf{H}}^{\mathbf{r},2d}_\mathbf{q}= {\mathbf{H}}^{(\mathbf{r}^1,\mathbf{r}^2),2d}_\mathbf{q}= {\mathbf{H}}^{r_1,r_2,2d}_\mathbf{q}\), where \(\mathbf{r}^i := (r_i,\dots,r_i) \in {\mathbb{R}}^d\), \(i=1,2\). Sometimes for brevity we omit \(2d\) in the notation for the \({\mathbf{H}}\) classes and write, for instance, \({\mathbf{H}}^{r_1,r_2}_\mathbf{q}\) instead of \({\mathbf{H}}^{r_1,r_2,2d}_\mathbf{q}\).
In this paper we study the case, when the asymptotic characteristic is the error of sampling recovery. Recall the setting of the optimal linear recovery introduced in [12]. For a fixed \(m\) and a set of points \(\xi:=\{\xi^j\}_{j=1}^m\subset \Omega\), let \(\Phi\) be a linear operator from \({\mathbb{C}}^m\) into \(L_p(\Omega,\mu)\). Denote for a class \({\mathbf{F}}\) (usually, centrally symmetric and compact subset of \(L_p(\Omega,\mu)\)) \[\varrho_m({\mathbf{F}},L_p) := \inf_{\xi} \inf_{\text{linear}\, \Phi } \sup_{f\in {\mathbf{F}}} \|f-\Phi(f(\xi^1),\dots,f(\xi^m))\|_p.\] The above described recovery procedure is a linear procedure.
Most of the known results on optimal sampling recovery deal with the linear recovery methods. We now give some very brief comments on recent results in this direction and refer the reader to the books [11], [10] and to the survey paper [6] for a discussion of the previous results in this direction. We are interested in results, which relate the errors of sampling recovery with the Kolmogorov widths for general function classes. We begin with a result from [5].
Theorem 2 ([5]). There exist two positive absolute constants \(b\) and \(B\) such that for any compact subset \(\Omega\) of \({\mathbb{R}}^d\), any probability measure \(\mu\) on it, and any compact subset \({\mathbf{F}}\) of \({\mathcal{C}}(\Omega)\) we have \[\label{kr1} \varrho_{bn}({\mathbf{F}},L_2(\Omega,\mu)) \le Bd_n({\mathbf{F}},L_\infty).\qquad{(5)}\]
The following generalization of Theorem 2 to the case \(2<p\le \infty\) was obtained in [7].
Theorem 3 ([7]). Let \(2\le p\le \infty\). There exists a positive absolute constant \(C\) such that for any compact subset \(\Omega\) of \({\mathbb{R}}^d\), any probability measure \(\mu\) on it, and any compact subset \({\mathbf{F}}\) of \({\mathcal{C}}(\Omega)\) we have \[\label{kr2} \varrho_{4n}({\mathbf{F}},L_p(\Omega,\mu)) \le Cn^{1/2-1/p}d_n({\mathbf{F}},L_\infty).\qquad{(6)}\]
In our applications the following analog of the inequality (?? ), which is contained in Theorem 3, \[\label{kr3} \varrho_{bn}({\mathbf{F}},L_\infty) \le Bn^{1/2}d_n({\mathbf{F}},L_\infty)\tag{6}\] plays a fundamental role. For this reason, we now present a detailed discussion of this inequality.
Let as above \(\Omega\) be a compact subset of \({\mathbb{R}}^d\) and \(X_N\) be an \(N\)-dimensional subspace of the space of continuous functions \({\mathcal{C}}(\Omega)\). Given a fixed \(m\) and a set of points \(\xi^1,\ldots,\xi^m\in\Omega\), we associate with a function \(f\in {\mathcal{C}}(\Omega)\) a vector (sample vector) \[S(f,\xi) := (f(\xi^1),\dots,f(\xi^m)) \in {\mathbb{C}}^m.\] We also consider the discrete norms \[\|S(f,\xi)\|_p:= \left(\frac{1}{m}\sum_{j=1}^m |f(\xi^j)|^p\right)^{1/p},\quad 1\le p<\infty,\] and \(\|S(f,\xi)\|_\infty := \max_{j}|f(\xi^j)|\).
For a positive weight \(\mathbf{w}:=(w_1,\dots,w_m)\in {\mathbb{R}}^m\) consider the following seminorm \[\|S(f,\xi)\|_{p,\mathbf{w}}:= \left(\sum_{j=1}^m w_j |f(\xi^j)|^p\right)^{1/p},\quad 1\le p<\infty.\] Define the best approximation of \(f\in L_p(\Omega,\mu)\), \(1\le p\le \infty\), by elements of \(X_N\) as follows \[d(f,X_N)_p := \inf_{u\in X_N} \|f-u\|_p.\]
Theorem 4 below was proved in [5] under the following assumptions.
A1. Discretization. Let \(1\le p\le \infty\). Suppose that \(\xi:=\{\xi^j\}_{j=1}^m\subset \Omega\) provides the following discretization property: For any \(u\in X_N\) in the case \(p<\infty\) we have \[\|u\|_p \le D \|S(u,\xi)\|_{p,\mathbf{w}}\] and in the case \(p=\infty\) we have \[\|u\|_\infty \le D \|S(u,\xi)\|_{\infty}\] with some positive constant \(D\).
A2. Weights. Suppose that there is a positive constant \(W\) such that \(\sum_{j=1}^m w_j \le W\).
Consider the following well known recovery operator (algorithm) \[\ell p\mathbf{w}(\xi)(f) := \ell p\mathbf{w}(\xi,X_N)(f):=\text{arg}\min_{u\in X_N} \|S(f-u,\xi)\|_{p,\mathbf{w}},\quad 1\le p<\infty,\] \[\ell \infty(\xi)(f) := \ell \infty(\xi,X_N)(f):=\text{arg}\min_{u\in X_N} \|S(f-u,\xi)\|_{\infty}.\] Note that the above algorithm \(\ell p\mathbf{w}(\xi)\) only uses the function values \(f(\xi^j)\), \(j=1,\dots,m\). In the case \(p=2\) it is a linear algorithm – orthogonal projection with respect to the seminorm \(\|\cdot\|_{2,\mathbf{w}}\). Therefore, in the case \(p=2\) the approximation error in the \(L_q\) norm by the algorithm \(\ell 2\mathbf{w}(\xi)\) gives an upper bound for the recovery characteristic \(\varrho_m(\cdot, L_q)\).
Theorem 4 ([5]). Under assumptions A1 and A2 for any \(f\in {\mathcal{C}}(\Omega)\) we have for \(1\le p<\infty\) \[\|f-\ell p\mathbf{w}(\xi)(f)\|_p \le (2DW^{1/p} +1)d(f, X_N)_\infty.\] Under assumption A1 for any \(f\in {\mathcal{C}}(\Omega)\) we have \[\|f-\ell \infty(\xi)(f)\|_\infty \le (2D+1)d(f, X_N)_\infty.\]
The following version of Theorem 4 for the error of \(\|f-\ell p\mathbf{w}(\xi)(f)\|_{\infty}\) under an extra condition on the Nikol’skii inequality for the \(X_N\) was proved in [13]. For completeness we present that proof here. For the reader’s convenience we recall the classical definition of the Nikol’skii inequality.
Nikol’skii-type inequalities. Let \(1\le p\le q\le\infty\) and \(X_N\subset L_q(\Omega, \mu)\). The inequality \[\label{I4} \|f\|_q \leq M\|f\|_p,\; \;\forall f\in X_N\tag{7}\] is called the Nikol’skii inequality for the pair \((p,q)\) with the constant \(M\). We will also use the brief form of this fact: \(X_N \in NI(p,q,M)\). Typically, \(M\) depends on \(N\), for instance, \(M\) can be of order \(N^{\frac{1}{p}-\frac{1}{q}}\).
Theorem 5 ([13]). Let \(1\le p<\infty\). Under assumptions A1, A2, and an extra assumption \(X_N\in NI(p,\infty,M)\) for any \(f\in {\mathcal{C}}(\Omega)\) we have \[\|f-\ell p\mathbf{w}(\xi)(f)\|_{\infty} \le (2MD W^{1/p} +1)d(f, X_N)_\infty.\]
Proof. The proof is simple and goes along the lines of the proof of Theorem 4. Let \(u:= \ell p\mathbf{w}(\xi)(f)\). For an arbitrary \(g \in X_N\) we have the following chain of inequalities. \[\|f-u\|_\infty \le \|f-g\|_\infty + \|g-u\|_\infty \le \|f-g\|_\infty +M\|g-u\|_p\] \[\le \|f-g\|_\infty + MD\|S(g-u,\xi)\|_{p,\mathbf{w}}\] \[\le \|f-g\|_\infty + MD(\|S(f-g,\xi)\|_{p,\mathbf{w}}+ \|S(f-u,\xi)\|_{p,\mathbf{w}})\] \[\le \|f-g\|_\infty + 2MD\|S(f-g,\xi)\|_{p,\mathbf{w}}\] \[\le \|f-g\|_\infty + 2MDW^{1/p}\|S(f-g,\xi)\|_{\infty} \le (1+ 2MDW^{1/p})\|f-g\|_\infty.\] Minimizing over \(g\in X_N\), we complete the proof. ◻
We now explain how to derive inequality (6 ) from Theorem 5. For a given function class \({\mathbf{F}}\subset {\mathcal{C}}(\Omega)\) and any \(\delta>0\) find a subspace \(X_N := X_N^\delta\) such that for any \(f\in {\mathbf{F}}\) we have \[\label{kr4} d(f, X_N)_\infty \le d_N({\mathbf{F}},L_\infty) +\delta.\tag{8}\] We want to apply Theorem 5 to the subspace \(X_N\). We will do that for \(p=2\). For that we need to check that the conditions of that theorem are satisfied. Namely, assumptions A1, A2, and the assumption \(X_N\in NI(p,\infty,M)\). To satisfy those conditions we can choose the measure \(\mu\), points \(\xi^1, \dots, \xi^m\), and weights \(\mathbf{w}\). We begin with the measure \(\mu\). We use the following fundamental result of J. Kiefer and J. Wolfowitz [14], which guarantees that for any finite dimensional subspace \(X_N\) of \({\mathcal{C}}(\Omega)\) there exists a probability measure \(\mu\) on \(\Omega\) such that for all \(f\in X_N\) we have \[\label{KW} \|f\|_\infty \le N^{1/2}\|f\|_{L_2(\Omega,\mu)}.\tag{9}\] In other words, for any subspace \(X_N\) of \({\mathcal{C}}(\Omega)\) we have \(X_N \in NI(2,\infty, N^{1/2})\) with some probability measure \(\mu\). We take this measure \(\mu\) and solve the discretization problem for the \(L_2(\Omega,\mu)\) norm on the subspace \(X_N\).
We use a result on discretization from [15] (see Theorem 3.3 there), which is a generalization to the complex case of an earlier result from [16] established for the real case.
Theorem 6 ([15]). If \(X_N\) is an \(N\)-dimensional subspace of the complex \(L_2(\Omega,\mu)\), then there exist three absolute positive constants \(C_1\), \(c_0\), \(C_0\), a set of \(m\leq C_1N\) points \(\xi^1,\ldots, \xi^m\in\Omega\), and a set of nonnegative weights \(\lambda_j\), \(j=1,\ldots, m\), such that \[\label{kr5} c_0\|f\|_2^2\leq \sum_{j=1}^m \lambda_j |f(\xi^j)|^2 \leq C_0\|f\|_2^2,\; \;\forall f\in X_N.\qquad{(7)}\]
For our application we need to satisfy the assumption A2 on weights. We use the following remark from [5].
Remark 1 ([5]). Considering a new subspace \(X_N' := \{f\,:\, f= g+c, \, g\in X_N,\, c\in {\mathbb{C}}\}\) and applying Theorem 6 to the \(X_N'\) with \(f=1\) (\(g=0\), \(c=1\)) we conclude that a version of Theorem 6 holds with the inequality \(m\le C_1N\) replaced by \(m\le C_1(N+1)\) and with weights satisfying \[\sum_{j=1}^m \lambda_j \le C_0.\]
We apply Theorem 5 with \(p=2\), \(M=N^{1/2}\), \(D= c_0^{-1/2}\), \(\mathbf{w}=(\lambda_1,\dots,\lambda_m)\), \(W=C_0^{1/2}\) and obtain \[\label{kr6} \varrho_{C_1(N+1)}({\mathbf{F}},L_\infty) \le (2N^{1/2}(C_0/c_0)^{1/2}+1)(d_N({\mathbf{F}},L_\infty) +\delta),\tag{10}\] which implies (6 ).
In the inequality (6 ) we only say that the parameter \(b\) can be chosen as an absolute constant. There are results on the inequality (?? ) with \(b\) being arbitrarily close to \(1\). The first result in that direction was proved in [15].
Theorem 7 ([15]). For any \(b\in(1,2]\) there exists a positive constant \(B=B(b)\) such that for any compact subset \(\Omega\) of \({\mathbb{R}}^d\), any probability measure \(\mu\) on it, and any compact subset \({\mathbf{F}}\) of \({\mathcal{C}}(\Omega)\) we have in the real case \[\varrho_{\lceil b(n+1) \rceil}({\mathbf{F}},L_2(\Omega,\mu)) \le Bd_n({\mathbf{F}},L_\infty)\] and in the complex case \[\varrho_{\lceil b(2n+1) \rceil}({\mathbf{F}},L_2(\Omega,\mu)) \le Bd_n({\mathbf{F}},L_\infty).\]
In the same way as we obtained above an analog (6 ) of the original inequality (?? ) we can obtain the following analog of Theorem 7.
Theorem 8. For any \(b\in(1,2]\) there exists a positive constant \(B=B(b)\) such that for any compact subset \(\Omega\) of \({\mathbb{R}}^d\) and any compact subset \({\mathbf{F}}\) of \({\mathcal{C}}(\Omega)\) we have in the real case \[\varrho_{\lceil b(n+1) \rceil}({\mathbf{F}},L_\infty) \le Bn^{1/2}d_n({\mathbf{F}},L_\infty)\] and in the complex case \[\varrho_{\lceil b(2n+1) \rceil}({\mathbf{F}},L_\infty)) \le Bn^{1/2}d_n({\mathbf{F}},L_\infty).\]
We complete this section with a brief historical comment.
Historical comments on weighted discretization. In the case of weighted discretization, namely, when instead of \(\frac{1}{m}\sum_{j=1}^m |f(\xi^j)|^2\) we use the weighted sum \(\sum_{j=1}^m\lambda_j |f(\xi^j)|^2\), the problem of discretization is solved in the sense of order in the case of real subspaces \(X_N\). It is pointed out in [17] that the paper by J. Batson, D.A. Spielman, and N. Srivastava [18] basically solves the discretization problem with weights. We present an explicit formulation of this important result in our notation.
Theorem 9 ([18]). Let \(\Omega_M=\{x^j\}_{j=1}^M\) be a discrete set with the probability measure \(\mu_M(x^j)=1/M\), \(j=1,\dots,M\), and let \(X_N\) be an \(N\)-dimensional subspace of real functions defined on \(\Omega_M\). Then for any number \(b>1\) there exists a set of weights \(\lambda_j\ge 0\) such that \(|\{j: \lambda_j\neq 0\}| \le \lceil bN \rceil\) so that for any \(f\in X_N\) we have \[\|f\|_2^2 \le \sum_{j=1}^M \lambda_jf(x^j)^2 \le \frac{b+1+2\sqrt{b}}{b+1-2\sqrt{b}}\|f\|_2^2.\]
As observed in [19] Theorem 2.13, this last theorem with a general probability space \((\Omega, \mu)\) in place of the discrete space \((\Omega_M, \mu_M)\) remains true (with other constant in the right hand side) if \(X_N\subset L_4(\Omega,\mu)\). It was proved in [16] that the additional assumption \(X_N\subset L_4(\Omega,\mu)\) can be dropped as well.
Theorem 10 ([16]). If \(X_N\) is an \(N\)-dimensional subspace of the real \(L_2(\Omega,\mu)\), then for any \(b\in (1,2]\), there exist a set of \(m\leq \lceil bN \rceil\) points \(\xi^1,\ldots, \xi^m\in\Omega\) and a set of nonnegative weights \(\lambda_j\), \(j=1,\ldots, m\), such that \[\|f\|_2^2\leq \sum_{j=1}^m \lambda_j f(\xi^j)^2 \leq \frac{C}{(b-1)^2} \|f\|_2^2,\; \;\forall f\in X_N,\] where \(C>1\) is an absolute constant.
In this section we discuss the best \(m\)-term bilinear approximations in \(L_\mathbf{p}({\mathbb{T}}^{2d})\) of functions from different classes. Our standard notation for the best \(m\)-term bilinear approximations is the following (see Section 1) \[\sigma_m({\mathbf{F}},\Pi)_\mathbf{p}:= \sup_{f\in {\mathbf{F}}}\sigma_m(f,\Pi)_\mathbf{p}.\] Note, that in a number of papers on this topic the following notation is used as well \[\tau_m({\mathbf{F}})_\mathbf{p}:= \sigma_m({\mathbf{F}},\Pi)_\mathbf{p},\qquad \tau_m(K)_\mathbf{p}:= \sigma_m(K,\Pi)_\mathbf{p}.\] In the formulation of the known results we use the \(\tau\) notation, which is used in the corresponding papers.
We begin with a simple lemma, which was proved (in a particular case) in [9]. For completeness we present a proof here.
Lemma 1 ([9]). We have for \(1\le p \le \infty\) \[\label{KB1} d_n({\mathbf{W}}^K_1,L_p) = \tau_n(K)_{p,\infty}\qquad{(8)}\] and for \(1\le q \le \infty\) \[\label{KB1a} d_n({\mathbf{W}}^K_q,L_p) \le \tau_n(K)_{p,q'}.\qquad{(9)}\]
Proof. It is clear that it is sufficient to prove (?? ) for continuous functions \(K\). For a fixed \(\mathbf{y}\in \Omega^2\) the function \(K(\mathbf{x},\mathbf{y})\) as a function on \(\mathbf{x}\in \Omega^1\) belongs to the closure of the class \({\mathbf{W}}^K_1\). Therefore, \[\label{KB2} d_n({\mathbf{W}}^K_1,L_p) \ge \inf_{u_i,v_i}\left \|K(\mathbf{x},\mathbf{y}) -\sum_{i=1}^n u_i(\mathbf{x})v_i(\mathbf{y})\right\|_{p,\infty} = \tau_n(K)_{p,\infty}.\tag{11}\] We now prove (?? ). Let for \(\varepsilon>0\) the systems of functions \(\{u_i\}_{i=1}^n \subset L_p(\Omega^1)\) and \(\{v_i\}_{i=1}^n \subset L_\infty(\Omega^2)\) be such that \[\label{KB3} \left \|K(\mathbf{x},\mathbf{y}) -\sum_{i=1}^n u_i(\mathbf{x})v_i(\mathbf{y})\right\|_{p,q'} \le \tau_n(K)_{p,q'} +\varepsilon.\tag{12}\] Then for any \(\varphi\in L_q(\Omega^2)\), \(\|\varphi\|_q \le1\), we have \[\label{KB4} \int_{\Omega^2} \left(K(\mathbf{x},\mathbf{y}) -\sum_{i=1}^n u_i(\mathbf{x})v_i(\mathbf{y})\right)\varphi(\mathbf{y})d\mu_2 = f(\mathbf{x}) - \sum_{i=1}^n a_iu_i(\mathbf{x})\tag{13}\] and \[\left\|f(\mathbf{x}) - \sum_{i=1}^n a_iu_i(\mathbf{x})\right\|_p \le \int_{\Omega^2} \left\|K(\cdot,\mathbf{y}) -\sum_{i=1}^n u_i(\cdot)v_i(\mathbf{y})\right\|_p|\varphi(\mathbf{y}|d\mu_2\] \[\le \left\|K(\mathbf{x},\mathbf{y}) -\sum_{i=1}^n u_i(\mathbf{x})v_i(\mathbf{y})\right\|_{p,q'} \le \tau_n(K)_{p,q'} +\varepsilon.\] This implies that \[\label{KB5} d_n({\mathbf{W}}^K_q,L_p) \le \tau_n(K)_{p,q'},\tag{14}\] which proves (?? ). Inequalities (11 ) and (14 ) with \(q=1\) complete the proof of (?? ). ◻
Some useful tricks. For bounded linear operators \(P\,:\, X\to Y\) and \(Q\,:\, Y\to Z\) acting in Banach spaces \(X\), \(Y\), \(Z\) we have the following simple inequality \[\label{KB6} d_{2n}(QP(B_X),Z) \le d_n(P(B_X),Y)d_n(Q(B_Y),Z).\tag{15}\]
Let \(H\) be a Hilbert space and \(J\,:\, H\to H\) be a compact linear operator. Then \[\label{KB7} d_n(J(B_H),H) = s_{n+1}(J).\tag{16}\]
Let \(K\in L_2(\Omega^1\times\Omega^2)\). Then the following Schmidt’s formula holds \[\label{KB8} \tau_n(K)_2 = \left(\sum_{i=n+1}^\infty s_i(J_K)^2\right)^{1/2},\tag{17}\] which implies that \[\label{KB9} s_{2n}(J_K) \le n^{-1/2}\tau_n(K)_2.\tag{18}\]
Some known results on bilinear approximation and singular numbers.
The case \(d\ge 1\). Here is the result from [20].
Theorem 11 ([20], Theorem 2.1). Let \(\mathbf{r}= (r_1,\dots,r_1,r_2,\dots,r_2)\in {\mathbb{R}}_+^{2d}\) have the first \(d\) coordinates equal \(r_1\) and the rest equal \(r_2\). Assume that \(r_i>1/2\), \(i=1,2\) and \(2\le \mathbf{q}\le \infty\), \(2\le \mathbf{p}<\infty\). Then for the class \({\mathbf{W}}^\mathbf{r}_{\mathbf{q}}\) we have \[\tau_m({\mathbf{W}}^\mathbf{r}_{\mathbf{q}})_\mathbf{p}\asymp m^{-r_1-r_2} (\log m)^{(r_1+r_2)(d-1)}.\]
Corollary 1. Under conditions of Theorem 11 we have \[s_m({\mathbf{W}}^\mathbf{r}_{\mathbf{q}}) := \sup_{K\in {\mathbf{W}}^\mathbf{r}_{\mathbf{q}}} s_m(J_K) \ll m^{-r_1-r_2-1/2} (\log m)^{(r_1+r_2)(d-1)}.\]
Here are the corresponding results for the \({\mathbf{H}}\) classes form [20].
Theorem 12 ([20], Theorem 2.2). Let \(\mathbf{r}= (r_1,\dots,r_1,r_2,\dots,r_2)\in {\mathbb{R}}_+^{2d}\) have the first \(d\) coordinates equal \(r_1\) and the rest equal \(r_2\). Assume that \(r_i>1/2\), \(i=1,2\) and \(2\le \mathbf{q}\le \infty\), \(2\le \mathbf{p}<\infty\). Then for the class \({\mathbf{H}}^\mathbf{r}_{\mathbf{q}}\) we have \[\tau_m({\mathbf{H}}^\mathbf{r}_{\mathbf{q}})_\mathbf{p}\asymp m^{-r_1-r_2} (\log m)^{(r_1+r_2+1)(d-1)}.\]
Corollary 2 ([20], Theorem 3.1). Let \(\mathbf{r}= (r_1,\dots,r_1,r_2,\dots,r_2)\in {\mathbb{R}}_+^{2d}\) have the first \(d\) coordinates equal \(r_1\) and the rest equal \(r_2\). Assume that \(r_i>0\), \(i=1,2\) and \(2\le \mathbf{q}\le \infty\). Then for the class \({\mathbf{H}}^\mathbf{r}_{\mathbf{q}}\) we have \[s_m({\mathbf{H}}^\mathbf{r}_{\mathbf{q}}) \asymp m^{-r_1-r_2-1/2} (\log m)^{(r_1+r_2+1)(d-1)}.\]
Remark 2. In the above Theorems 11 and 12 we impose the restriction \(r_i>1/2\), \(i=1,2\). We need this restriction for proving the upper bounds in the case of scalar \(q=2\) and arbitrarily large \(p\). In the case \(p=q=2\) it is sufficient to assume that \(r_i>0\), \(i=1,2\). In the case \(r_1=r_2\) Theorems 11 and 12 were proved in [21].
The case \(d=1\). The case \(d=1\) is better studied than the general case. We now formulate the corresponding results. The following results are from [9]. We use the following notation for \(1\le q,p\le\infty\) \[\label{ksi} \xi(q,p):= \left(\frac{1}{q} - \max\left(\frac{1}{2},\frac{1}{p}\right)\right)_+, \quad (a)_+ :=\max(a,0).\tag{19}\]
Theorem 13 ([9], Theorem 2). Let \(d=1\) and \({\mathbf{F}}^\mathbf{r}_\mathbf{q}\) denote one of the classes \({\mathbf{W}}^\mathbf{r}_\mathbf{q}\) or \({\mathbf{H}}^\mathbf{r}_\mathbf{q}\). Then for \(\mathbf{r}> \mathbf{1}\) and \(1\le q_1\le p_1 \le \infty\), \(1\le q_2,p_2 \le \infty\) we have \[\tau_m({\mathbf{F}}^\mathbf{r}_\mathbf{q})_\mathbf{p}\asymp m^{-r_1-r_2 + \xi(q_1,p_1)}.\]
Remark 3. Note that Theorem 13 is proved in [9] under weaker conditions on \(\mathbf{r}\) than above. That restriction on \(\mathbf{r}\) is needed for the proof of the upper bounds. For the lower bounds it is sufficient to assume that \(\mathbf{r}> (1/q_1-1/p_1, (1/q_2-1/p_2)_+)\).
Corollary 3 ([9], Theorem 3.2). Under conditions of Theorem 13 we have \[\sup_{K\in {\mathbf{F}}^\mathbf{r}_\mathbf{q}} s_m(J_K) \asymp m^{-r_1-r_2 + \max\left(\frac{1}{2},\frac{1}{q_1}\right)-1}.\]
Some known results on the Kolmogorov widths of classes \({\mathbf{W}}^K_q\).
We begin with the case of univariate functions (\(d=1\)), in which case the kernel \(K\) is a function on two variables. The following results are proved in [9].
Theorem 14 ([9], Theorem 4.2). Let \(d=1\) and \({\mathbf{F}}^\mathbf{r}_1\) denote one of the classes \({\mathbf{W}}^\mathbf{r}_1\) or \({\mathbf{H}}^\mathbf{r}_1\). Then for \(1\le q,p \le \infty\) and \(\mathbf{r}> (1,1+\max(1/2,1/q))\) we have \[\sup_{K\in {\mathbf{F}}^\mathbf{r}_1} d_m({\mathbf{W}}^K_q)_p \asymp m^{-r_1-r_2 + \xi(q,p)}\] with \(\xi(q,p)\) defined in (19 ).
For \(\mathbf{q}=(q_1,q_2)\), \(\mathbf{p}=(p_1,p_2)\), \(1\le q_1 \le p_1 \le \infty\), \(1\le q_2,p_2 \le \infty\) denote \[\mathbf{r}(\mathbf{q},\mathbf{p}) := \begin{cases} (1/q_1-1/p_1,(1/q_2-1/p_2)_+), & 1\le q_1\le p_1 \le 2, \\ (1/q_1,1/q_2), & 2\le q_1\le p_1\le \infty, p_1>2,\\ (1/q_1,\max(1/2,1/q_2)), & 1\le q_1<2< p_1 \le \infty. \end{cases}\]
Theorem 15 ([9], Theorem 4.1). Let \(d=1\) and \({\mathbf{F}}^\mathbf{r}_\mathbf{q}\) denote one of the classes \({\mathbf{W}}^\mathbf{r}_\mathbf{q}\) or \({\mathbf{H}}^\mathbf{r}_\mathbf{q}\). Then for \(\mathbf{p}=(p,\infty)\), \(1\le q_1 \le p \le \infty\), \(1\le q_2\le \infty\) and \(\mathbf{r}> \mathbf{r}(\mathbf{q},\mathbf{p})\) we have \[\sup_{K\in {\mathbf{F}}^\mathbf{r}_\mathbf{q}} d_m({\mathbf{W}}^K_1)_p \asymp m^{-r_1-r_2 + \xi(q_1,p)}\] with \(\xi(q,p)\) defined in (19 ).
Note that in the case \(\mathbf{r}> \mathbf{1}\) Theorem 15 follows from Lemma 1 and Theorem 13 (see also Remark 3).
In the above Theorem 14 we consider the case of classes \({\mathbf{F}}^\mathbf{r}_1\). Some results on the classes \({\mathbf{F}}^\mathbf{r}_\mathbf{q}\) are obtained in [20] (see Theorem 3.1’ there). We formulate that result as Theorem 16 and refer the reader to the paper [9] for further results and historical comments on bilinear approximation of functions on two variables with mixed smoothness. Denote \[\mathbf{r}(\mathbf{q}) := ((1/q_1-1/2)_+,(1/q_2-1/2)_+).\]
Theorem 16 ([20], Theorem 3.1’). Let \(d=1\) and \({\mathbf{F}}^\mathbf{r}_\mathbf{q}\) denote one of the classes \({\mathbf{W}}^\mathbf{r}_\mathbf{q}\) or \({\mathbf{H}}^\mathbf{r}_\mathbf{q}\), \(\mathbf{1} \le \mathbf{q}\le \infty\). Then for \(2\le a\le \infty\), \(1\le b\le \infty\) under assumption that \(\mathbf{r}> \mathbf{r}(\mathbf{q})\) for \(1\le b\le 2\) and \(\mathbf{r}> \mathbf{r}(\mathbf{q})+ (1/2,0)\) for \(b>2\) we have \[\sup_{K\in {\mathbf{F}}^\mathbf{r}_\mathbf{q}} d_m({\mathbf{W}}^K_a)_b \asymp m^{-r_1-r_2 + \max(1/q_1,1/2) -1} .\]
Here is an analog of Theorem 14, which holds for \(d=1\), in the case \(d>1\).
Theorem 17 ([20], Theorem 3.2). Let \(d\in{\mathbb{N}}\) and \(\mathbf{2} \le \mathbf{q}\le \infty\), \(2\le a <\infty\), \(1<b<\infty\). Assume that in the case \(b\in (1,2]\) we have \(r_i>0\), \(i=1,2\), and in the case \(b\in (2,\infty)\) we have \(r_1>1/2\), \(r_2>0\). Then \[\sup_{K\in {\mathbf{W}}^\mathbf{r}_\mathbf{q}} d_m({\mathbf{W}}^K_a)_b \asymp (m^{-1}(\log m)^{d-1})^{r_1+r_2}m^{-1/2} .\]
Theorems 2 and 3 show that in the study of linear recovery the Kolmogorov widths in the uniform norm \(L_\infty\) play an important role. In this section we focus on the case of the uniform norm and complement the results known in the case of \(L_p\), \(p\in [2,\infty)\), by the case \(p=\infty\). The following Theorem 18 is a step in that direction from the above Theorem 17.
Theorem 18. Let \(d\in{\mathbb{N}}\). Assume that we have \(r_1>1/2\), \(r_2>0\). Then \[\sup_{K\in {\mathbf{W}}^\mathbf{r}_2} d_m({\mathbf{W}}^K_2)_\infty \ll m^{-r_1-r_2-1/2}(\log m)^{(d-1)(r_1+r_2)+1/2} .\]
Proof. We remind some known results that we use in the proof. E. Belinsky (see [22]) proved the following bounds \[\label{KolWb} d_m({\mathbf{W}}^r_{2},L_\infty) \ll m^{-r}(\log m)^{(d-1)r+1/2}, \qquad r>1/2.\tag{20}\] The following bound was obtained in [9] (see Theorem 3.1 there): For \(g\in {\mathbf{W}}^{a_1,a_2}_2\), \(a_1>0\), \(a_2>0\) we have \[\label{KolWg} d_m({\mathbf{W}}^g_{2},L_2) = s_{m+1}(J_g) \ll m^{-a_1-a_2-1/2}(\log m)^{(d-1)(a_1+a_2)} .\tag{21}\] We also need the following operators of fractional integration and differentiation. We begin with the univariate case. In this subsection we discuss a slightly more general Bernoulli kernels and integral operators related to them (see [10], Section 1.4). In the univariate case, for \(a>0\), and \(\alpha\in {\mathbb{R}}\) let \[F_{a,\alpha}(x):= 1+2\sum_{k=1}^\infty k^{-a}\cos (kx-\alpha\pi/2)\] \[\label{Lb1} = 1+\sum_{k=1}^\infty k^{-a}(e^{i\alpha\pi/2}e^{-ikx}+e^{-i\alpha\pi/2}e^{ikx})\tag{22}\] be the generalised Bernoulli kernel. Clearly, we have \(F_{a}(x) = F_{a,a}(x)\), where \(F_a(x)\) is defined in (?? ). Define the integral operator, acting on trigonometric polynomials \(\phi(x)\), as \[\label{Lb2} (I^{(a,\alpha)}\phi)(x) := (I^{(a,\alpha)}_x\phi)(x) := ( F_{a,\alpha} \ast \phi)(x):= \frac{1}{2\pi}\int_{{\mathbb{T}}} F_{a,\alpha}(x-z) \phi(z)dz.\tag{23}\] The operator \(I^{(a,\alpha)}\) is the multiplier operator: \[\label{Lb3} (I^{(a,\alpha)}\phi)(x) = \hat{\phi}(0)+ \sum_{k<0} |k|^{-a}e^{i\alpha\pi/2} \hat{\phi}(k)e^{ikx}+ \sum_{k>0}k^{-a}e^{-i\alpha\pi/2} \hat{\phi}(k)e^{ikx}.\tag{24}\] Identity (24 ) implies that \[\label{Lb339} I^{(a,\alpha)}I^{(b,\beta)} = I^{(b,\beta)} I^{(a,\alpha)} = I^{(a+b,\alpha+\beta)} .\tag{25}\]
We now define the inverse operator to the operator \(I^{(a,\alpha)}\), acting on the trigonometric polynomials from \({\mathcal{T}}(2n)\) (we take \(2n\) for convenience in the future use). Define \[{\mathcal{D}}^{(a,\alpha)}_{2n}(x) := 1+2\sum_{k=1}^{2n} k^{a}\cos (kx+\alpha\pi/2)\] and the operator (for \(h \in {\mathcal{T}}(2n)\)) \[\label{Lb4} (D^{(a,\alpha)}h)(x) := (D^{(a,\alpha)}_xh)(x) := ({\mathcal{D}}^{a,\alpha}_{2n} \ast h)(x):= \frac{1}{2\pi}\int_{{\mathbb{T}}} {\mathcal{D}}^{a,\alpha}_{2n}(x-z) h(z)dz.\tag{26}\] The operator \(D^{(a,\alpha)}\) is the multiplier operator: \[\label{Lb3a} (D^{(a,\alpha)}h)(x) = \hat{h}(0)+ \sum_{k<0} |k|^{a}e^{-i\alpha\pi/2} \hat{h}(k)e^{ikx}+ \sum_{k>0}k^{a}e^{i\alpha\pi/2} \hat{h}(k)e^{ikx}.\tag{27}\] It is easy to see that for \(h \in {\mathcal{T}}(2n)\) we have \[\label{Lb5} I^{(a,\alpha)}D^{(a,\alpha)}h = h.\tag{28}\] Clearly, the operator \(D^{(a,\alpha)}\) can be defined for smooth enough functions \(h\) instead of the trigonometric polynomials. Then relation (28 ) means that \(I^{(a,\alpha)}D^{(a,\alpha)} =Id\), where \(Id\) is the identity operator.
For convenience, we write \[I^a_x := I^{(a,a)}_x,\qquad D^a_x := D^{(a,a)}_x.\] For vectors \(\mathbf{r}= (r_1,\dots,r_d)\) and \(\mathbf{x}= (x_1,\dots,x_d)\) define \[I^\mathbf{r}_\mathbf{x}:= \prod_{j=1}^d I^{r_j}_{x_j},\qquad D^\mathbf{r}_\mathbf{x}:= \prod_{j=1}^d D^{r_j}_{x_j}.\]
Assume that \(K\in {\mathbf{W}}^\mathbf{r}_2 = {\mathbf{W}}^{(\mathbf{r}^1,\mathbf{r}^2)}_2\) (see below) with \(r_1>1/2\), \(r_2>0\). Let \(u\) be a number satisfying \(1/2 <u< r_1\). This means that there exists \(\phi \in L_{2}({\mathbb{T}}^{2d})\), \(\|\phi\|_2 \le 1\), such that \[K = I^{\mathbf{r}^1}_\mathbf{x}I^{\mathbf{r}^2}_\mathbf{y}\phi,\qquad \mathbf{r}^1 =(r_1,\dots,r_1) \in {\mathbb{R}}^d,\quad \mathbf{r}^2 =(r_2,\dots,r_2) \in {\mathbb{R}}^d.\] Represent \[I^{\mathbf{r}^1}_\mathbf{x}= I^{\mathbf{u}}_\mathbf{x}I^{\mathbf{a}}_\mathbf{x},\quad \mathbf{u}:= (u,\dots,u) \in {\mathbb{R}}^d,\quad \mathbf{a}:= (r_1-u,\dots,r_1-u) \in {\mathbb{R}}^d.\] Denote \(g:= I^{\mathbf{a}}_\mathbf{x}I^{\mathbf{r}^2}_\mathbf{y}\phi\). Then \(g\in {\mathbf{W}}^{(\mathbf{a},\mathbf{r}^2)}_2\) and \(J_K = J_{F_\mathbf{u}}J_g\). Therefore, \[\label{Lb6} d_{2m}({\mathbf{W}}^K_2,L_\infty) \le d_m({\mathbf{W}}^\mathbf{u}_2,L_\infty) d_{m}({\mathbf{W}}^g_2,L_2).\tag{29}\] We now use relations (20 ), (21 ) and complete the proof. ◻
For the future use we formulate the inequality (29 ) proved above as a separate statement.
Lemma 2. Let \(d\in {\mathbb{N}}\) and \(\mathbf{u}\in {\mathbb{R}}^d_+\), \(\mathbf{u}:= (u,\dots,u)\), \(u>1/2\). Assume that \(K\) is such that \(g := D^\mathbf{u}_\mathbf{x}K \in L_2(\Omega^1 \times \Omega^2)\). Then \[\label{Lb7} d_{2m}({\mathbf{W}}^K_2,L_\infty) \le d_m({\mathbf{W}}^\mathbf{u}_2,L_\infty) d_{m}({\mathbf{W}}^g_2,L_2).\qquad{(10)}\]
We now prove an analog of Theorem 18 for the \({\mathbf{H}}\) classes.
Theorem 19. Let \(d\in{\mathbb{N}}\). Assume that we have \(r_1>1/2\), \(r_2>0\). Then \[\sup_{K\in {\mathbf{H}}^{r_1,r_2}_2} d_m({\mathbf{W}}^K_2)_\infty \ll m^{-r_1-r_2-1/2}(\log m)^{(d-1)(r_1+r_2)+d-1/2} .\]
Proof. We use notations from the above proof of Theorem 18. By Lemma 2 we have \[\label{Lb8} d_{2m}({\mathbf{W}}^K_2,L_\infty) \le d_m({\mathbf{W}}^\mathbf{u}_2,L_\infty) d_{m}({\mathbf{W}}^g_2,L_2), \quad g:= D^\mathbf{u}_\mathbf{x}K.\tag{30}\]
We now need one more lemma.
Lemma 3. Assume that we have \(r_1>1/2\), \(r_2>0\). Then for \(K \in {\mathbf{H}}^{r_1,r_2}_2\) and \(\mathbf{u}:= (u,\dots,u)\), \(1/2<u< r_1\) we have \(g := D^\mathbf{u}_\mathbf{x}K \in {\mathbf{H}}^{r_1-u,r_2}_2\).
Proof. By Theorem 1 we get for all \(\mathbf{s}\in {\mathbb{N}}_0^{2d}\) \[\|A_\mathbf{s}(K)\|_2 \ll 2^{-r_1\|\mathbf{s}^1\|_1-r_2\|\mathbf{s}^2\|_1}, \quad \mathbf{s}^1:= (s_1,\dots,s_d), \quad \mathbf{s}^2:= (s_{d+1},\dots,s_{2d}).\] From here we easily obtain \[\|A_\mathbf{s}(g)\|_2 = \|A_\mathbf{s}(D^\mathbf{u}_\mathbf{x}K)\|_2 \ll 2^{u\|\mathbf{s}^1\|_1} \|A_\mathbf{s}(K)\|_2 \ll 2^{-(r_1-u)\|\mathbf{s}^1\|_1-r_2\|\mathbf{s}^2\|_1}.\] By Theorem 1 we conclude \(g := D^\mathbf{u}_\mathbf{x}K \in {\mathbf{H}}^{r_1-u,r_2}_2\), which proves Lemma 3. ◻
We continue proof of Theorem 19. By (20 ) we get \[\label{KB9a} d_m({\mathbf{W}}^\mathbf{u}_2,L_\infty) \ll m^{-u} (\log m)^{(d-1)u +1/2}.\tag{31}\] By Corollary 1 and Remark 2 we have for \(g \in {\mathbf{H}}^{r_1-u,r_2}_2\) \[\label{KB10a} d_{m}({\mathbf{W}}^g_2,L_2) \ll m^{-(r_1-u)-r_2-1/2} (\log m)^{(d-1)(r_1-u+r_2) +(d-1)}.\tag{32}\] Combining (30 ) – (32 ), we complete the proof of Theorem 19. ◻
Sampling recovery. We begin with a simple inequality for the linear recovery \(\varrho_m({\mathbf{W}}^K_q,L_p)\).
Proposition 1. Let \(1\le q,p, \le \infty\). Assume that for every \(\mathbf{z}\in \Omega^1\) we have \(K(\mathbf{z},\cdot) \in L_{q'}(\Omega^2)\), \(q':=q/(q-1)\). Then we have \[\label{RN1} \varrho_m({\mathbf{W}}^K_q,L_p) \le \sigma_m(K,{\mathcal{L}}{\mathcal{K}})_{p,q'}.\qquad{(11)}\]
Proof. Consider an operator \(\Psi_m\) of linear recovery \[\Psi_m(f,\xi,\mathbf{x}) := \sum_{j=1}^m f(\xi^j)\psi_j(\mathbf{x}).\] Then we have for \(f\in {\mathbf{W}}^K_q\) \[\|f-\Psi_m(f,\xi,\mathbf{x})\|_p \le \int_{\Omega^2} \left\|K(\cdot,\mathbf{y})-\sum_{j=1}^m K(\xi^j,\mathbf{y})\psi_j(\cdot)\right\|_p |\varphi(\mathbf{y})|d\mu_2.\] This implies that \[\sup_{f\in{\mathbf{W}}^K_q} \|f-\Psi_m(f,\xi,\mathbf{x})\|_p \le \left\|K(\mathbf{x},\mathbf{y})-\sum_{j=1}^m K(\xi^j,\mathbf{y})\psi_j(\mathbf{x})\right\|_{p,q'} .\] We now take infimum over sets of points \(\{\xi^j\}_{j=1}^m\) and sets of functions \(\{\psi_j\}_{j=1}^m\) and complete the proof. ◻
For the next simple relation we need a new notation. Define for \(p_1,p_2\) \[\|f(\mathbf{x},\mathbf{y})\|_{L^*_{p_1,p_2}}:=\|f(\mathbf{x},\mathbf{y})\|^*_{p_1,p_2} := \|\|f(\mathbf{x},\cdot)\|_{p_2}\|_{p_1},\] which means that first we take the norm with respect to \(\mathbf{y}\) and after that the norm with respect to \(\mathbf{x}\).
Proposition 2. Let \(1\le q \le \infty\). Assume that for every \(\mathbf{z}\in \Omega^1\) we have \(K(\mathbf{z},\cdot) \in L_{q'}(\Omega^2)\), \(q':=q/(q-1)\). Then we have \[\label{RN2} \varrho_m({\mathbf{W}}^K_q,L_\infty) = \sigma_m(K,{\mathcal{L}}{\mathcal{K}})_{L^*_{\infty,q'}}.\qquad{(12)}\]
Proof. In the same way as in the above proof of Proposition 1 we obtain for \(f\in {\mathbf{W}}^K_q\) \[\|f-\Psi_m(f,\xi,\mathbf{x})\|_\infty\] \[=\sup_{\mathbf{x}\in\Omega^1}\left| \int_{\Omega^2} \left(K(\mathbf{x},\mathbf{y})-\sum_{j=1}^m K(\xi^j,\mathbf{y})\psi_j(\mathbf{x})\right) \varphi(\mathbf{y})d\mu_2 \right| =: \sup_{\mathbf{x}\in\Omega^1}E(\mathbf{x},\varphi).\] Therefore, \[\sup_{f\in{\mathbf{W}}^K_q}\|f-\Psi_m(f,\xi,\mathbf{x})\|_\infty= \sup_{f\in{\mathbf{W}}^K_q}\sup_{\mathbf{x}\in\Omega^1}E(\mathbf{x},\varphi) = \sup_{\mathbf{x}\in\Omega^1} \sup_{f\in{\mathbf{W}}^K_q}E(\mathbf{x},\varphi)\] \[=\left\|K(\mathbf{x},\mathbf{y})-\sum_{j=1}^m K(\xi^j,\mathbf{y})\psi_j(\mathbf{x})\right\|^*_{\infty,q'}.\] We now take infimum over sets of points \(\{\xi^j\}_{j=1}^m\) and sets of functions \(\{\psi_j\}_{j=1}^m\) and complete the proof. ◻
In the Section 1 we defined the systems \(\Pi={\mathcal{L}}{\mathcal{L}}\), \({\mathcal{L}}{\mathcal{K}}\), \({\mathcal{K}}{\mathcal{L}}\), and \({\mathcal{K}}{\mathcal{K}}\). In this section we only discuss the best \(m\)-term approximations with respect to some of these systems and therefore normalization of elements of these systems does not play any role. Obviously, we have the following inclusions for any \(\mathbf{p}\) \[{\mathcal{K}}{\mathcal{K}}(\mathbf{p}) \subset {\mathcal{L}}{\mathcal{K}}(\mathbf{p}) \subset \Pi(\mathbf{p}).\] These inclusions immediately imply the following trivial inequalities \[\sigma_m(K,\Pi(\mathbf{p}))_\mathbf{p}\le \sigma_m(K,{\mathcal{L}}{\mathcal{K}}(\mathbf{p}))_\mathbf{p}\le \sigma_m(K,{\mathcal{K}}{\mathcal{K}}(\mathbf{p}))_\mathbf{p}.\]
In this section we discuss the following fundamental problems.
Problem \({\mathcal{L}}{\mathcal{K}}-\Pi\). Find an upper bound for \(\sigma_m(K,{\mathcal{L}}{\mathcal{K}}(\mathbf{p}))_\mathbf{p}\) in terms of \(\sigma_n(K,\Pi(\mathbf{p}))_\mathbf{p}\) with \(n\) close to \(m\).
General Problem \({\mathcal{L}}{\mathcal{K}}-\Pi\). Find an upper bound for \(\sigma_m(K,{\mathcal{L}}{\mathcal{K}}(\mathbf{p}))_\mathbf{p}\) in terms of \(\sigma_n(K^*,\Pi(\mathbf{p}))_\mathbf{p}\) with \(n\) close to \(m\), where \(K^*\) is a new function build from \(K\). Certainly, we would like the operator mapping \(K\) to \(K^*\) to be as simple as possible. For instance, it might be a differentiation operator, which is popular in approximation theory.
We begin with a result on the Problem \({\mathcal{L}}{\mathcal{K}}-\Pi\).
Theorem 20. For any \(b\in(1,2]\) there exists a positive constant \(B=B(b)\) such that for any continuous on \(\Omega^1\times\Omega^2\) function \(K(\mathbf{x},\mathbf{y})\) we have \[\label{2In1} \sigma_m(K,{\mathcal{L}}{\mathcal{K}}(\infty))_\infty \le Bm^{1/2} \sigma_{\theta (m-1)}(K,\Pi(\infty))_\infty\qquad{(13)}\] with \(\theta =1/b\) in the real case and \(\theta =1/(2b)\) in the complex case.
Proof. By Proposition 2 with \(q=1\) we know that \[\label{2In2} \sigma_m(K,{\mathcal{L}}{\mathcal{K}}(\infty))_\infty = \varrho_m({\mathbf{W}}^K_1,L_\infty).\tag{33}\] By Theorem 8 with \({\mathbf{F}}= {\mathbf{W}}^K_1\) we obtain \[\label{2In3} \varrho_m({\mathbf{W}}^K_1,L_\infty) \le Bm^{1/2} d_{\theta (m-1)}({\mathbf{W}}^K_1, L_\infty)\tag{34}\] with \(\theta =1/b\) in the real case and \(\theta =1/(2b)\) in the complex case. Finally, by Lemma 1 with \(p=\infty\) we get \[\label{2In4} d_{\theta (m-1)}({\mathbf{W}}^K_1, L_\infty) = \sigma_{\theta (m-1)}(K,\Pi(\infty))_\infty.\tag{35}\] Combining relations (33 ) – (35 ), we complete the proof of Theorem 20. ◻
Proposition 3. The extra factor \(m^{1/2}\) in the inequality (?? ) of Theorem 20 is sharp.
Proof. The claim of Proposition 3 follows from known results. The following result on bilinear approximations is known.
Theorem 21 ([9], Theorem 2). Let \(d=1\) and \({\mathbf{F}}^\mathbf{r}_\mathbf{q}\) denote one of the classes \({\mathbf{W}}^\mathbf{r}_\mathbf{q}\) or \({\mathbf{H}}^\mathbf{r}_\mathbf{q}\). Then for \(\mathbf{r}> \mathbf{1}\) and \(1\le q_1\le p_1 \le \infty\), \(1\le q_2,p_2 \le \infty\) we have \[\sup_{K\in {\mathbf{F}}^\mathbf{r}_\mathbf{q}} \sigma_m(K,\Pi(\mathbf{p}))_\mathbf{p}\asymp m^{-r_1-r_2 + \xi(q_1,p_1)}\] where \(\xi(q,p)\) is defined in (19 ).
The following result of approximation with respect to adaptive dictionaries was obtained in the recent paper [1] (see Theorem 1.8 there).
Theorem 22. Let \(d=1\) and \({\mathbf{F}}^\mathbf{r}_\mathbf{q}\) denote one of the classes \({\mathbf{W}}^\mathbf{r}_\mathbf{q}\) or \({\mathbf{H}}^\mathbf{r}_\mathbf{q}\) (see the definition in Section 2 below) of functions of two variables. Then for \(1\le q_1\le 2\), \(1\le q_2 \le \infty\), and \(\mathbf{r}> \mathbf{r}(\mathbf{q})\) we have \[\sup_{K\in {\mathbf{F}}^\mathbf{r}_\mathbf{q}} \sigma_m(K,{\mathcal{L}}{\mathcal{K}}(\infty))_\infty \asymp m^{-r_1-r_2 + 1/q_1} .\]
In the case of scalar \(p = \infty\) and scalar \(q\in [1,2]\) we have \[\xi(q,p) = 1/q -1/2.\] It remains to compare Theorems 21 and 22 in this case. ◻
We now proceed to the General Problem \({\mathcal{L}}{\mathcal{K}}-\Pi\) in the case of scalar \(p=2\).
Theorem 23. Let \(\mathbf{u}= (u,\dots,u)\in {\mathbb{R}}^d\), \(u>1/2\). Assume that for every \(\mathbf{z}\in \Omega^1\) we have \(K(\mathbf{z},\cdot) \in L_{2}(\Omega^2)\) and \(K^{(u)} := D^{\mathbf{u}}_\mathbf{x}K \in L_2(\Omega^1\times\Omega^2)\). Then we have \[\label{RN2a} \sigma_m(K,{\mathcal{L}}{\mathcal{K}})_{L^*_{\infty,2}} \le C(u,d) m^{-u}(\log m)^{u(d-1)+1/2} \sigma_{m/8}(K^{(u)},\Pi)_2.\qquad{(14)}\]
Here is a direct corollary of Theorem 23.
Corollary 4. Under conditions of Theorem 23 we have \[\label{RI2} \sigma_m(K,{\mathcal{L}}{\mathcal{K}})_2 \le C(u,d) m^{-u}(\log m)^{u(d-1)+1/2} \sigma_{m/16}(K^{(u)},\Pi)_2.\qquad{(15)}\]
Proof of Theorem 23. By Proposition 2 with \(q=2\) we find that \[\label{RI3} \sigma_m(K,{\mathcal{L}}{\mathcal{K}})_{L^*_{\infty,2}} = \varrho_m({\mathbf{W}}^K_2,L_\infty).\tag{36}\] By Theorem 3 with \({\mathbf{F}}= {\mathbf{W}}^K_2\) and \(p=\infty\) we obtain \[\label{RI4} \varrho_m({\mathbf{W}}^K_2,L_\infty) \le Cm^{1/2} d_{m/4}({\mathbf{W}}^K_2, L_\infty).\tag{37}\] By Lemma 2 we get \[\label{RI5} d_{m/4}({\mathbf{W}}^K_2,L_\infty) \le d_{m/8}({\mathbf{W}}^\mathbf{u}_2,L_\infty) d_{m/8}({\mathbf{W}}^g_2,L_2), \quad g:= K^{(u)}.\tag{38}\] By (20 ) we find \[\label{RI6} d_{m/8}({\mathbf{W}}^u_{2},L_\infty) \ll m^{-u}(\log m)^{(d-1)u+1/2}, \qquad u>1/2.\tag{39}\] Next by (16 ) and (18 ) we conclude \[\label{RI7} d_{m/8}({\mathbf{W}}^g_2,L_2) \le (m/8)^{-1/2} \sigma_{m/16}(g,\Pi)_2.\tag{40}\] Combining (36 ) – (40 ), we complete the proof of Theorem 23.
We formulate corollaries of Proposition 2, Theorems 18, 19, and Theorem 8.
Theorem 24. Let \(d\in{\mathbb{N}}\). Assume that we have \(r_1>1/2\), \(r_2>0\). Then \[\sup_{K\in {\mathbf{W}}^\mathbf{r}_2} \sigma_m(K,{\mathcal{L}}{\mathcal{K}})_{L^*_{\infty,2}} \ll m^{-r_1-r_2}(\log m)^{(d-1)(r_1+r_2)+1/2} .\]
Corollary 5. Let \(d\in{\mathbb{N}}\). Assume that we have \(r_1>1/2\), \(r_2>0\). Then \[\sup_{K\in {\mathbf{W}}^\mathbf{r}_2} \sigma_m(K,{\mathcal{L}}{\mathcal{K}})_{2} \ll m^{-r_1-r_2}(\log m)^{(d-1)(r_1+r_2)+1/2} .\]
Theorem 25. Let \(d\in{\mathbb{N}}\). Assume that we have \(r_1>1/2\), \(r_2>0\). Then \[\sup_{K\in {\mathbf{H}}^{r_1,r_2}_2} \sigma_m(K,{\mathcal{L}}{\mathcal{K}})_{L^*_{\infty,2}} \ll m^{-r_1-r_2}(\log m)^{(d-1)(r_1+r_2)+d-1/2} .\]
Corollary 6. Let \(d\in{\mathbb{N}}\). Assume that we have \(r_1>1/2\), \(r_2>0\). Then \[\sup_{K\in {\mathbf{H}}^{r_1,r_2}_2} \sigma_m(K,{\mathcal{L}}{\mathcal{K}})_{2} \ll m^{-r_1-r_2}(\log m)^{(d-1)(r_1+r_2)+d-1/2} .\]
V.N. Temlyakov, University of South Carolina, USA,
Steklov Mathematical Institute of Russian Academy of Sciences, Russia;
Lomonosov Moscow State University, Russia;
Moscow Center of Fundamental and Applied Mathematics, Russia.
E-mail: temlyakovv@gmail.com