Jackknife Variance Estimation for Hájek-Dominated Generalized U-Statistics


Abstract

Valid uncertainty quantification for subsampling-based and randomized estimators often depends on variance estimators whose behavior is much less understood than that of the underlying point estimator. We prove ratio-consistency of the jackknife variance estimator, and certain delete-\(d\) variants, for a broad class of generalized U-statistics whose variance is asymptotically dominated by their Hájek projection and whose normalized first-projection squares satisfy a row-wise \(L^r\) weak law, with the classical fixed-order case recovered as a special instance. This projection-dominance plus square-LLN structure unifies and generalizes several criteria from the existing literature, clarifies when the simple nonparametric jackknife is theoretically justified in the generalized setting, and yields consistent variance estimation for the two-scale distributional nearest-neighbor regression estimator under substantially weaker conditions than previously required.

1 Introduction↩︎

Variance estimation is often the bottleneck in turning modern subsampling-based estimators into usable inference. Many estimators of current interest in statistics and econometrics, including random-subsample learners and localized nonparametric estimators, admit generalized U-statistic representations. Their asymptotic behavior is increasingly well understood, but variance estimation for this class remains much less developed outside the classical fixed-order setting. The jackknife is attractive because it is simple, generic, and computationally convenient, but theoretical guarantees in these generalized settings remain scarce.

We establish conditions under which the jackknife yields ratio-consistent variance estimates for generalized U-statistics. The central message is that jackknife consistency is governed by two first-projection requirements rather than by the full combinatorial expansion of the statistic. The statistic’s variance must be asymptotically dominated by its first-order (Hájek) projection, and the normalized squares of that projection must obey a mild row-wise weak law. The dominance condition keeps the higher-order Hoeffding terms asymptotically negligible. The square weak law then ensures that the diagonal averaged-square term appearing in the delete-\(d\) calculation tracks the same variance scale. For the delete-1 jackknife, this second requirement is just the familiar law of large numbers for squared first-order influence terms; it becomes a separate condition here because the generalized kernel itself changes with \(n\). Together, the two requirements give a unified treatment of the ordinary jackknife and its delete-\(d\) variants. They extend classical jackknife results for U-statistics (e.g., [1]) to the generalized setting without modifying the base estimator. They also place the plain jackknife on the same conceptual footing as infinitesimal-jackknife-style procedures in generalized settings, beginning with Jaeckel’s original Bell Laboratories memorandum [2] and later developments such as [3], while using milder conditions than many existing alternatives.

As an application, we revisit the Two-Scale Distributional Nearest-Neighbor (TDNN) regression estimator of [4]. Under mild regularity conditions, this estimator satisfies asymptotic Hájek dominance whenever the kernel order grows more slowly than the sample size, substantially weakening the original assumptions used to justify jackknife-based variance estimates.

Related Literature↩︎

Our starting point is the classical theory of U-statistics initiated by [5]; for general background, see [6]. For variance estimation, fixed-order benchmarks include [1][7], and [8] for jackknife methods, and [9] for the bootstrap. In a closely related fixed-order setting, [10] study small-sample variance estimators for U-statistics and find the ordinary jackknife to compare favorably with direct term-by-term corrections. These papers are the baseline against which our generalized results should be read.

A second strand studies U-statistics whose order grows with the sample size and related subsampling-based statistics with the same combinatorial structure. An early reference is [11], which develops large-sample properties for infinite-order U-statistics. In the modern ensemble literature, [12] show that predictions from subsampled tree ensembles can be analyzed through a U-statistic representation, while [13] develop a complementary V-statistic viewpoint for with-replacement ensembles. Relatedly, [14] use generalized U-statistics to obtain rates of convergence for random forests, and [4] develop pointwise inference for the two-scale distributional nearest-neighbor estimator that motivates our application.

A third strand studies variance estimation when the kernel order is not small relative to the sample size. For general U-statistics, [15] propose an unbiased variance estimator together with a partition-resampling implementation, and [16] extend that line with a pseudo-kernel construction aimed at large kernel sizes and cross-validation problems. For ensemble predictors, [17] study jackknife and bootstrap standard errors for bagging and random forests, [18] develop jackknife and infinitesimal-jackknife variance estimators for bagged learners and random forests, and [19] propose an unbiased variance estimator for subsampling-based ensembles under a U-statistic formulation.

Recent work also emphasizes that leading-term approximations can fail when the kernel order is large. In that spirit, [20] study variance estimation for random forests through what they call peak-region dominance and establish ratio consistency for their unbiased estimator, [21] analyze the bias-consistency of the infinitesimal jackknife and extend the discussion to subsampling-based U-statistics, and [22] argue for looking beyond the leading term in the exact variance expansion through a dominant-region viewpoint.

Our contribution answers a different question. Rather than designing a new estimator for a particular generalized U-statistic, we ask when the plain delete-\(d\) jackknife is ratio-consistent. The answer is a projection-level one: asymptotic Hájek dominance plus a row-wise weak law for squared first-order projections is enough to make the ordinary jackknife and its delete-\(d\) variants work in both complete and incomplete Bernoulli-sampled settings. This isolates the first-projection structure behind jackknife consistency, clarifies when the unmodified jackknife is theoretically justified in generalized settings, and yields the TDNN application under the substantially weaker growth condition \(s_2 = o(n)\). At the application level, this perspective is close to nearest-neighbor and forest views of localization, including the adaptive-nearest-neighbor interpretation of forests in [23] and the localized-weight perspective of generalized random forests in [24], and more broadly to localized conditional moment estimation with DNN weights as the localizer.

Notation↩︎

We use the following notation. Let \([n] = \{1, \dotsc, n\}\). Given a finite index set \(\mathcal{I}\subset \mathbb{N}\), we introduce the following notational conventions. \[L_{s}(\mathcal{I}) = \set*{(l_1, \dotsc, l_s) \in \mathcal{I}^{s} \;\middle|\;l_{1} < l_{2} < \dotsc < l_{s}} \quad \text{and} \quad L_{n,s} = L_s\left([n]\right)\] Write \(D_{[n]}=(Z_1,\dotsc,Z_n)\) for an i.i.d.sample from \(F_Z\) with associated measure \(\mu_Z\), so \(D_{[n]} \sim \bigotimes_{i=1}^n \mu_Z\). A realization is denoted by \(d_{[n]}\). For \(\ell \in L_{n,s}\), \(D_{[n],-\ell}\) is the data set with indices in \(\ell\) removed, and \(D_\ell\) is the sub-sample indexed by \(\ell\). For a single deleted observation we write \(D_{[n],-i}\). We use analogous notation for covariates, responses, and realizations, e.g. \(X_{[c]}=(X_1,\dotsc,X_c)\), \(Y_{[c]}=(Y_1,\dotsc,Y_c)\), \(x_{[n]}\), and \(y_{[n]}\). If two index vectors \(\ell\) and \(\iota\) are disjoint, \(\ell \cup \iota\) denotes their concatenation; for example, if \(\ell=(8,2,5)\) and \(\iota=(1,6)\), then \(\ell \cup \iota=(8,2,5,1,6)\). Finally, \(\rightsquigarrow\) denotes weak convergence and \(\xrightarrow{\mathbb{P}}\) convergence in probability. We write \(A \lesssim B\) to mean \(A \leq C B\) for a universal constant \(C\) that does not depend on \(n\) or \(s\), for all sufficiently large \(n\).

Paper Organization↩︎

Section 2 introduces the generalized U-statistic framework, states the Hájek-dominance, square-LLN, and sampling conditions, and proves ratio-consistency of ordinary and delete-\(d\) jackknife variance estimators in complete and incomplete settings. Section 3 applies the framework to TDNN estimation, showing that jackknife variance estimation remains valid when \(s_2=o(n)\) and that studentized inference follows whenever TDNN asymptotic normality holds on the same variance scale. The appendix, beginning with 6, contains the general delete-\(d\) jackknife proofs, the single-scale DNN inputs, and the TDNN extension.

2 Generalized U-Statistics↩︎

We use the generalized U-statistic framework of [14]. This framework unifies incomplete, randomized, and infinite-order U-statistics, covering random forests and a broad class of ensemble estimators. The appeal of the framework is that many modern subsampling estimators can be studied through a common projection structure rather than through estimator-specific algebra. Developing jackknife theory at this level therefore supports the use of such estimators in fields such as computer science and economics.

Definition 1 (Generalized U-Statistic).
* Suppose \(D_{[n]} = \left(Z_1, \ldots, Z_n\right)\) is a data set consisting of i.i.d.observations from \(F_Z\). Let \(h\) denote a (possibly randomized) real-valued function utilizing \(s\) of these observations that is permutation-symmetric in those \(s\) arguments. A generalized U-statistic with kernel \(h\) of order (rank) \(s\) is any estimator of the form \[\label{eq:genUStat} U_{n, s, N, \omega}\left(D_{[n]}\right) = \frac{1}{\widehat{N}} \sum_{\ell \in L_{n,s}} \rho_{\ell} h\left(D_{\ell} ; \omega\right)\tag{1}\] where \(\omega\) denotes i.i.d.randomness, independent of the original data. \(\left(\rho_{\ell}\right)_{\ell \in L_{n,s}}\) denotes a collection of i.i.d.Bernoulli random variables determining which subsamples are selected and is independent of all other inputs to the U-statistic. Set \(p \mathrel{\vcenter{:}}=\Pb*{\rho_{\ell}=1} = N /\binom{n}{s}\), and let the actual number of selected subsamples be \(\widehat{N}=\sum_{\ell \in L_{n,s}} \rho_{\ell}\), so \(\Ex*{\widehat{N}}=N\). The Bernoulli design therefore has \(0<N\leq\binom{n}{s}\), so \(0<p\leq1\); \(N\) is the expected selected count and need not be an integer. The ratio in 1 is understood on the event \(\widehat N>0\); on the event \(\widehat N=0\), where the numerator is also zero, we set \(U_{n,s,N,\omega}\left(D_{[n]}\right)=0\). This zero-count event has probability \(\Pb*{\widehat N=0}=(1-p)^{\binom{n}{s}}\leq \exp(-N)\), so the convention is asymptotically immaterial whenever \(N \to \infty\), as in the incomplete-statistic regimes considered below. When \(N=\binom{n}{s}\), the estimator in 1 is a complete generalized U-statistic and is denoted as \(U_{n, s, \omega}\). When \(N<\binom{n}{s}\), these estimators are incomplete generalized U-statistics.

Here \(N\) is the target number of subsamples, \(\widehat{N}\) is the realized number selected by the Bernoulli design, and \(\rho_{\ell}\) records whether the kernel is evaluated on subset \(\ell\). The complete case sets every \(\rho_{\ell}=1\), while the incomplete case keeps the same kernel but randomizes which subsamples are used. Separating the kernel from the sampling scheme lets one variance-estimation argument cover both settings.

We later apply the general results to the two-scale distributional nearest-neighbor estimator of [4]. For a kernel \(h_s\), write \(\theta_{s} = \Ex*{h_{s}\left(D_{[s]}\right)}\) for the nominal kernel mean. In the complete case this is exactly the expectation of the statistic. In the incomplete Bernoulli-sampled case, the zero-count convention in 1 changes the literal expectation to \(\theta_s\Pb*{\widehat N>0}\), because the statistic is set to zero on \(\set*{\widehat N=0}\). The discrepancy is confined to an exponentially unlikely event, so \(\theta_s\) remains the natural centering scale. When convenient, the complete-case proof may set \(\theta_s=0\) by exact centering. The incomplete-case proof keeps the same centering interpretation but uses an explicit zero-count centering-transfer argument to account for the convention. For some results, boundedness of \(\theta_s\) in \(s\) is the more important requirement, and this is benign in most applications. For TDNN estimation, we keep the nonparametric regression function explicit because it is central to the application. The auxiliary randomness can be viewed as \(\omega=(W_\ell)_{\ell \in L_{n,s}}\), where each \(W_\ell\) contains only the extra randomness injected into the kernel on subsample \(\ell\).

Much of classical U-statistic theory rests on the celebrated Hoeffding decomposition of [5]. This technique decomposes a classical U-statistic into uncorrelated components of orders one through \(s\), with each term capturing progressively higher-order interactions. Equivalently, the order-\(i\) component is the projection of the kernel onto the space of functions that depend on exactly \(i\) arguments and are orthogonal to all lower-order components. It is a powerful tool for studying variance and limiting distributions because the components are orthogonal and higher-order terms are often negligible. [14] extend this decomposition to generalized U-statistics, with the classical result as a special case.

Definition 2 (Generalized Hoeffding Decomposition).
* Suppose \(D_{[n]} = \left(Z_1, \ldots, Z_n\right)\) is a data set consisting of i.i.d.samples from \(F_Z\) and \(d_{[n]}\) is a fixed realization of the data set. Let \(h_{s}\left(D_{[s]} ; \omega\right)\) be a (possibly randomized) real valued function that is permutation-symmetric in \(D_{[s]}\). Let \[h_{s|i}\left(d_{[i]}\right) = \Ex*{h_{s}\left(d_{[i]}, D_{[s-i]} ; \omega\right)} - \Ex*{h_s(D_{[s]}; \omega)}\] for \(i=1, \ldots, s\) and let \[\begin{align} h_{s}^{(i)}\left(D_{[i]}\right) & = h_{s|i}\left(D_{[i]}\right) - \sum_{j=1}^{i-1} \sum_{\ell \in L_{i,j}} h_{s}^{(j)}\left(D_{\ell}\right) && \text{for } i=1, \ldots, s-1, \\ h_{s}^{(s)}\left(D_{[s]}; \omega\right) & = h_{s}\left(D_{[s]} ; \omega\right)-\sum_{j=1}^{s-1} \sum_{\ell \in L_{s,j}} h_{s}^{(j)}\left(D_{\ell}\right), \\ H_{s}^{i}\left(D_{[n]}\right) & = \binom{n}{i}^{-1} \sum_{\ell \in L_{n, i}} h_{s}^{(i)}\left(D_{\ell}\right), \quad \text{for } i=1, \ldots, s-1 \quad \text{and} \\ H_{s}^{s}\left(D_{[n]}; \omega\right) & = \binom{n}{s}^{-1} \sum_{\ell \in L_{n, s}} h_{s}^{(s)}\left(D_{\ell}; \omega\right). \end{align}\] The \(H\)-decomposition of a generalized complete U-statistic is expressed as \[\begin{align} U_{n, s, \omega}\left(D_{[n]}\right) & = \sum_{i=1}^{s-1}\left[\binom{s}{i}\binom{n}{i}^{-1} \sum_{\ell \in L_{n,i}} h_s^{(i)}\left(D_{\ell}\right)\right] + \binom{n}{s}^{-1} \sum_{\ell \in L_{n,s}} h_s^{(s)}\left(D_{\ell}; \omega\right) \\ & = \sum_{i=1}^{s-1}\binom{s}{i} H_{s}^{i}\left(D_{[n]}\right) + H_{s}^{s}\left(D_{[n]}; \omega\right). \end{align}\]

The lower-order projection kernels \(h_s^{(i)}\) for \(i < s\) carry no \(\omega\) argument because conditional expectations integrate out the external randomization; \(\omega\) only remains in the residual term \(h_s^{(s)}\).

The decomposition in 2 identifies the key variance-estimation objects: the first-projection kernel \(h_s^{(1)}\), the Hájek term \((s/n)\sum_{i=1}^n h_s^{(1)}(Z_i)\), and the higher-order remainder \(\sum_{j=2}^s \binom{s}{j} H_s^j\left(D_{[n]}\right)\). Operationally, \(h_s^{(1)}(Z_i) = \Ex*{h_s(D_{[s]};\omega) \;\middle|\;Z_i} - \theta_s\) measures how observation \(i\) shifts the kernel’s conditional mean, while the higher-order terms collect interaction effects that are harder for resampling methods to track. The argument asks when these interactions are small enough that jackknife perturbations behave as if applied to the linear Hájek term.

Continuing with the notation of [14], define the following variance terms. \[\begin{align} \zeta_{s, \omega}^{c} & = \Covp*{h\left(D_{[c]}, D_{[s-c]}; \omega\right), h\left(D_{[c]}, D_{[s-c]}^{\prime}; \omega^{\prime}\right)} \quad \text{for} \quad c = 1, \dotsc, s-1 \tag{2} \\ \zeta_{s}^{s} & = \Covp*{h\left(D_{[s]}; \omega\right), h\left(D_{[s]}; \omega\right)} = \Varb*{h_{s}\left(D_{[s]}; \omega\right)}\tag{3} \\ V_{s, \omega}^{c} & = \Varb*{h_{s}^{(c)}\left(D_{[c]}\right)} \quad \text{for} \quad c = 1, \dotsc, s-1 \\ V_{s}^{s} & = \Varb*{h_{s}^{(s)}\left(D_{[s]}; \omega\right)} \end{align}\] Here, variables with a prime (such as \(D_{[s-c]}^{\prime}\) or \(\omega^{\prime}\)) denote random variables that follow the same distribution as their non-prime counterparts. They are also independent of all other input variables, including each other and their non-prime counterparts. The subscript \(\omega\) for terms of order \(1 \leq c < s\) indicates that expectations include \(\omega\), so additional randomization appears only in the final residual terms of the expansions. For \(c=1\), the covariance definition coincides with the variance of the first Hoeffding projection. Indeed, the two kernel evaluations in \(\zeta_{s,\omega}^{1}\) are conditionally independent given the shared observation \(Z_1\), so \[\zeta_{s,\omega}^{1} = \Varb*{\Ex*{h_s(D_{[s]};\omega)\;\middle|\;Z_1}} = \Varb*{h_s^{(1)}(Z_1)}.\] Thus, \(\zeta_{s,\omega}^1\) is the first-projection variance scale and will be the key quantity governing the leading term in both the variance decomposition and the jackknife analysis below. By contrast, \(\zeta_s^s = \Varb*{h_s\left(D_{[s]}; \omega\right)}\) is the full-kernel variance scale. Standard U-statistic results, extended to the generalized setting, give the following variance decomposition in terms of Hoeffding projection variances. \[\label{eq:Var95decomp} \begin{align} \Varb*{U_{n, s, \omega}\left(D_{[n]}\right)} & = \sum_{j = 1}^{s-1} \binom{s}{j}^2 \Varb*{H_{s}^{j}\left(D_{[n]}\right)} + \Varb*{H_{s}^{s}\left(D_{[n]}; \omega\right)} \\ & = \sum_{j = 1}^{s - 1} \binom{s}{j}^2 \binom{n}{j}^{-1} \Varb*{h_{s}^{(j)}\left(D_{[j]}\right)} + \binom{n}{s}^{-1} \Varb*{h_{s}^{(s)}\left(D_{[s]}; \omega\right)} \\ & = \sum_{j = 1}^{s-1} \binom{s}{j}^2 \binom{n}{j}^{-1} V_{s, \omega}^{j} + \binom{n}{s}^{-1} V_{s}^{s} \\ \end{align}\tag{4}\] For fixed \(s\) and no additional randomization, 4 recovers the classical U-statistic variance decomposition.

3 Consistent Variance Estimation↩︎

Theorems 1 and 2 of [14] establish normality of generalized U-statistics under suitable Hájek-projection variance dominance conditions. The same condition drives jackknife consistency here, so we state it separately. The delete-\(d\) jackknife results below are stated for triangular-array rows with integer kernel order \(s=s_n\geq2\) eventually; the rank-one case is not treated separately here because the higher-order Hoeffding remainder is absent.

Assumption 1 (Asymptotic Hájek Dominance Condition).
* Consider a potentially incomplete generalized U-statistic \(U_{n, s, N, \omega}\) along rows with \(\zeta_{s,\omega}^{1}>0\) eventually. If \[\frac{s}{n} \left(\frac{\zeta_{s}^{s}}{s \zeta_{s, \omega}^{1}} - 1\right) \longrightarrow 0,\] we say \(U_{n, s, N, \omega}\) satisfies the asymptotic Hájek dominance condition.

This condition places the statistic in a first-order variance regime. As \(n \rightarrow \infty\) and \(s = s(n) \rightarrow \infty\), both the full-kernel variance and the first-projection variance may vanish, but their relative scale must leave the first-order term asymptotically decisive. In that regime, deleting observations perturbs the statistic primarily through its linear Hájek term rather than through higher-order interactions.

Lemma 1 (Dominance of Hájek Projection Variance).
* Let \(\mathrm{U}_{n, s, \omega}\left(D_{[n]}\right)\) be a complete generalized U-statistic. Let the kernel variance terms \(\zeta_{s}^{s}\) and \(\zeta_{s, \omega}^{1}\) be defined as in 2 3 . Then under 1, the Hájek projection term asymptotically dominates the variance of the generalized U-statistic \[\frac{n}{s^2}\frac{\Varb*{\mathrm{U}_{n, s, \omega}\left(D_{[n]}\right)}}{\zeta_{s, \omega}^{1}} \longrightarrow 1.\]

Thus 1 says that individual-observation contributions, not higher-order interactions, explain the leading variance, with scale \((s^2/n)\zeta_{s,\omega}^1\). That variance comparison controls the higher-order Hoeffding remainder, but the delete-\(d\) proof also contains a diagonal averaged-square term built from the normalized first projection. To make that term track the same variance scale, we impose the following row-wise weak-law condition.

Assumption 2 (Row-wise \(L^r\) Square-LLN Condition).
* Consider a potentially incomplete generalized U-statistic \(U_{n, s, N, \omega}\) along rows with \(\zeta_{s,\omega}^{1}>0\) eventually. Write \[W_{n,1} \mathrel{\vcenter{:}}= \frac{h_s^{(1)}(Z_1)^2}{\zeta_{s,\omega}^1}.\] If there exists \(r \in (1,2]\) such that \[\frac{\Ex*{W_{n,1}^r}}{n^{r-1}} \longrightarrow 0,\] we say \(U_{n,s,N,\omega}\) satisfies the row-wise \(L^r\) square-LLN condition.

This is the standard \(L^r\) sufficient condition for a weak law of row-wise i.i.d.triangular arrays; see, for example, [25]. In Appendix A we use the von Bahr–Esseen inequality [26] to record the short proof specialized to the normalized first-projection squares. For the ordinary delete-1 jackknife with fixed kernel order, this condition is essentially automatic: the first projection is an ordinary i.i.d.influence term, so the usual law of large numbers controls the diagonal average. It becomes visible here because \(h_s^{(1)}\) and \(W_{n,1}\) form a triangular array as the kernel order grows with \(n\).

Remark 1. A stronger but sometimes convenient sufficient condition is the standardized first-projection moment bound \[\sup_n \frac{\Ex*{\abs*{h_s^{(1)}(Z_1)}^{2+\eta}}}{\left(\zeta_{s,\omega}^{1}\right)^{1+\eta/2}} < \infty \qquad \text{for some } \eta \in (0,2].\] Then 2 follows with \(r = 1 + \eta/2\). This is the same scale of projection moment condition used by [14] for Berry–Esseen analysis, but it is stronger than what the jackknife consistency proof itself requires.

The assumptions now separate the two sources of difficulty in the complete case. 1 suppresses the higher-order Hoeffding remainder, while 2 controls the diagonal average of squared first-order contributions. Incomplete generalized U-statistics introduce a third layer because the jackknife replicates reuse a random subset of possible kernel evaluations. For that setting we impose the following sampling condition from [21], which ensures that each deleted-sample statistic is still based on asymptotically enough sampled subsamples.

Assumption 3 (Asymptotically-Sufficient Sampling Condition).
* Consider a potentially incomplete generalized U-statistic \(U_{n, s, N, \omega}\) along rows with \(\zeta_{s,\omega}^{1}>0\) eventually. If \[\frac{n}{N s \zeta_{s, \omega}^{1}} \longrightarrow 0,\] we say \(U_{n, s, N, \omega}\) satisfies the asymptotically-sufficient sampling condition.

This condition is specific to the incomplete case. Each jackknife replicate reuses only the subsamples that were actually drawn, so the Bernoulli sampling layer must be rich enough that each observation still appears many times on average. Otherwise, sampling noise from the subsample-selection scheme can dominate the jackknife comparison before the Hájek projection has a chance to govern the variance.

When \(\zeta_{s,\omega}^{1}\asymp1\), this condition is equivalent up to constants to \(n/(Ns)\to0\). In that case the total number of observation appearances across selected subsamples grows faster than the sample size, or equivalently each observation appears a diverging number of times in expectation. For example, this occurs when \(N = n^{1+\alpha}\) and \(s \rightarrow \infty\) with \(s = o(n)\) and \(\alpha > 0\). Taken together, the three assumptions describe three separate requirements. The statistic is asymptotically governed by its linear Hájek projection, the diagonal average of squared linear contributions obeys a law of large numbers, and Bernoulli thinning is dense enough to remain a second-order perturbation. This is precisely the regime in which delete-\(d\) jackknife perturbations should behave as if they were applied to a linear statistic. We now justify jackknife variance estimation under this projection-dominance and square-LLN structure, beginning with the variance target \[\label{eq:genUStat95Var} \sigma^{2}_{n} = \Varb*{U_{n, s, N, \omega}\left(D_{[n]}\right)}.\tag{5}\] The subscript in 5 emphasizes dependence on \(n\), and hence on \(s\) and \(N\). We consider the following variance estimators.

Definition 3 (Jackknife Variance Estimators).
* We consider the jackknife variance estimator \[\label{eq:JK95Var95Est} \widehat{\sigma}_{JK}^2\left(D_{[n]}; \omega\right) \mathrel{\vcenter{:}}=\frac{n-1}{n} \sum_{i = 1}^{n}\left(U_{n, s, N, \omega}\left(D_{[n], -i}\right) - U_{n, s, N, \omega}\left(D_{[n]}\right)\right)^2\tag{6}\] and the delete-\(d\) jackknife variance estimator \[\label{eq:JKD95Var95Est} \widehat{\sigma}_{JKD}^2\left(D_{[n]}; d, \omega\right) \mathrel{\vcenter{:}}=\frac{n-d}{d}\binom{n}{d}^{-1} \sum_{\ell \in L_{n,d}}\left(U_{n, s, N, \omega}\left(D_{[n], -\ell}\right) - U_{n, s, N, \omega}\left(D_{[n]}\right)\right)^2.\tag{7}\] For incomplete generalized U-statistics, the deleted-sample evaluations use the same realizations of \(\rho\) and \(\omega\) as the full-sample statistic. This keeps the jackknife comparison focused on deleting observations rather than mixing that effect with fresh subsampling or auxiliary-randomization noise. Specifically, for a delete set \(\ell \in L_{n,d}\), define \[\mathcal{O}_{s,0}(\ell) \mathrel{\vcenter{:}}= \set*{\iota \in L_{n,s}: \iota \cap \ell = \emptyset}, \qquad \widehat N_{\ell}^{\circ} \mathrel{\vcenter{:}}= \sum_{\iota \in \mathcal{O}_{s,0}(\ell)} \rho_{\iota}.\] Then \(U_{n,s,N,\omega}(D_{[n],-\ell})\) is evaluated by retaining only the originally sampled subsamples that avoid \(\ell\): \[U_{n,s,N,\omega}\left(D_{[n],-\ell}\right) = \frac{1}{\widehat N_{\ell}^{\circ}} \sum_{\iota \in \mathcal{O}_{s,0}(\ell)} \rho_{\iota} h\left(D_{\iota}; \omega\right), \qquad \text{when } \widehat N_{\ell}^{\circ}>0.\] If \(\widehat N_{\ell}^{\circ}=0\), we again set this deleted-sample evaluation to zero. For each fixed \(\ell\), this zero-count event has probability at most \(\exp(-N_{d}^{\circ})\), where \(N_{d}^{\circ}=p\binom{n-d}{s}\); when \(N\to\infty\) and \(sd=o(n)\), \(N_{d}^{\circ}\to\infty\) because \(N_{d}^{\circ}/N\to 1\). No Bernoulli subsamples are redrawn after deletion.

For \(d=1\), 7 reduces to the ordinary jackknife estimator in 6 . We first consider complete generalized U-statistics and point out that analogous results on the consistency of the pseudo-infinitesimal jackknife estimator are derived in [21].

Theorem 1 (Variance Estimation for Complete Generalized U-Statistics).
* Let \(U_{n, s, \omega}\) be a complete generalized U-statistic satisfying 1 2. Suppose \(2\leq s=s_n\leq n\), \(1\leq d=d_n\), \(s+d\leq n\) eventually, \(s=o(n)\), \(sd=o(n)\), \(\zeta_{s,\omega}^{1}>0\) eventually, and \(\sigma_n^2>0\) eventually. Let \(\widehat{\sigma}_{JKD}^2\left(D_{[n]}; d, \omega\right)\) be the associated delete-\(d\) jackknife variance estimator as defined in 7 . Then \[\frac{\widehat{\sigma}_{JKD}^2\left(D_{[n]}; d, \omega\right)}{\sigma_{n}^{2}} \xrightarrow{\mathbb{P}}1.\] In particular, for \(d = 1\), \[\frac{\widehat{\sigma}_{JK}^2\left(D_{[n]}; \omega\right)}{\sigma_{n}^{2}} \xrightarrow{\mathbb{P}}1.\]

6.3 contains the full proof.

Remark 2. Although \(sd = o(n)\) allows \(d\) to grow, the most relevant practical regimes have fixed \(d\). The result then allows one to combine delete-\(d\) estimates for different fixed values of \(d\) to eliminate lower-order bias terms. Although such combinations are beyond this paper, they could improve finite-sample performance. In contrast, for a classical U-statistic with fixed \(s\) this result justifies the use of delete-\(d\) jackknife variance estimators in the regime \(d = o(n)\).

The extension to incomplete generalized U-statistics is practically important. When no closed-form evaluation is available, evaluating the kernel on all subsets of size \(s\) is often prohibitive, so one uses a Bernoulli sampling scheme as in 1. The following result covers that setting.

Theorem 2 (Variance Estimation for Incomplete Generalized U-Statistics).
* Let \(U_{n, s, N, \omega}\) be a potentially incomplete generalized U-statistic satisfying 1 2 3. Suppose \(2\leq s=s_n\leq n\), \(1\leq d=d_n\), \(s+d\leq n\) eventually, \(s=o(n)\), \(sd=o(n)\), \(\zeta_{s,\omega}^{1}>0\) eventually, and \(\sigma_n^2>0\) eventually. Furthermore, let \(\zeta_{s}^{s}\) and \(\theta_s=\Ex*{h_s(D_{[s]};\omega)}\) be bounded in \(s\). Let \(\widehat{\sigma}_{JKD}^2\left(D_{[n]}; d, \omega\right)\) be the associated delete-\(d\) jackknife variance estimator as defined in 7 . Then \[\frac{\widehat{\sigma}_{JKD}^2\left(D_{[n]}; d, \omega\right)}{\sigma_{n}^{2}} \xrightarrow{\mathbb{P}}1.\] In particular, for \(d = 1\), \[\frac{\widehat{\sigma}_{JK}^2\left(D_{[n]}; \omega\right)}{\sigma_{n}^{2}} \xrightarrow{\mathbb{P}}1.\]

6.4 contains the full proof.

4 Application to Two-Scale Distributional Nearest Neighbor Estimator↩︎

We illustrate the inferential payoff of the general theory with the two-scale distributional nearest-neighbor regression estimator of [4]. This estimator targets pointwise nonparametric regression at a fixed covariate value \(x\), so valid uncertainty quantification requires a consistent variance estimator for a localized subsampling-based object. The TDNN estimator is therefore a natural application of the generalized U-statistic results above. We work under the following setup.

Assumption 4 (Nonparametric Regression DGP).
* The observed data consist of an i.i.d.sample taking the following form. \[D_{[n]} = \{Z_{i} = (X_{i}, Y_{i})\}_{i = 1}^{n} \quad \text{from the model} \quad Y = \mu(X) + \varepsilon,\] where \(Y \in \mathcal{Y}\subset \mathbb{R}\) is the response, \(X \in \mathcal{X}\subset \mathbb{R}^k\) is a feature vector of fixed dimension \(k\) distributed according to a density function \(f\) with associated probability measure \(\varphi\) on \(\mathcal{X}\), and \(\mu(x)\) is the unknown mean regression function. \(\varepsilon\) is the unobservable model error on which we impose the following conditions. \[\Ex*{\varepsilon \;\middle|\;X} = 0, \quad \Varb*{\varepsilon \;\middle|\;X = x} = \sigma_{\varepsilon}^2\left(x\right)\] Let the distribution induced by this model be denoted by \(P\) and thus \(Z_{i} = \left(X_{i}, Y_{i}\right) \overset{\text{iid}}{\sim} P\).

The pointwise asymptotic-normality result of [4] assumes homoskedastic errors independent of \(X\) and centers the estimator by an explicit residual bias term. For studentization, we combine that result with the heteroskedastic localization adjustment below and require the residual bias to be negligible on the TDNN standard-error scale.

Remark 3. The heteroskedastic version replaces the uses of \(\varepsilon \perp X\) in [4] with localized conditional moments. The required changes occur only where independence is used to factor an error moment away from a DNN selector weight. For the cubic term below, we also use continuity at \(x\) of \(m_3(u) \coloneq \Ex*{\varepsilon^3 \;\middle|\;X=u}\). Let \(q_s(X_1) = \Ex*{\kappa(x;Z_1,D_{[s]})\;\middle|\;X_1}\) be the DNN selector weight from 16. (i) Second-moment terms. The upper-bound argument uses the second moment of the error weighted by the DNN selector. Under independence this appears as the factorization \(\Ex*{\varepsilon_1^2\,q_s(X_1)} = \sigma_\varepsilon^2\,\Ex*{q_s(X_1)}\). Under 4, the corresponding expression is \(\Ex*{\varepsilon_1^2\,q_s(X_1)} = \Ex*{\sigma_\varepsilon^2(X_1)\,q_s(X_1)}\). Since \(s\,\Ex*{\sigma_\varepsilon^2(X_1)\,q_s(X_1)} \to \sigma_\varepsilon^2(x)\) by DNN localization and the variance-function continuity in 5 [asm:tdnn95design95variance], the upper bound holds with the local value \(\sigma_\varepsilon^2(x)\). (ii) Variance lower bound. The first-projection variance lower bound is also localized. The homoskedastic bound \(\eta_1 \geq \sigma_\varepsilon^2/(2s-1)\) uses \(\varepsilon \perp X\). Under 4, the same role is played by \(\eta_1 \geq \Ex*{\sigma_\varepsilon^2(X_1)\,q_s^2(X_1)} \geq \underline{\sigma}_\varepsilon^2/(2s-1)\), exactly as established in 17. (iii) Fourth-moment expansion. The only fourth-moment term that changes qualitatively is the cubic cross-term \(\Ex*{\mu(X_1)\,m_3(X_1)\,q_s(X_1)}\) in the expansion of \(\Ex*{Y_1^4\,q_s(X_1)}\) (Demirkaya Lemma 3), which no longer factors. Since \(\mu\) is bounded on \(\mathcal{X}\), \(\Ex*{Y^4}<\infty\) implies \(\mu(\cdot)m_3(\cdot)\in L^1\), and continuity of \(m_3\) at \(x\) gives continuity of \(\mu(\cdot)m_3(\cdot)\) at \(x\); therefore 7 yields convergence to \(\mu(x)\,\Ex*{\varepsilon^3\;\middle|\;X=x}\).

With this localization convention in place, we separate the inputs used for jackknife variance consistency from the smoother inputs used only when invoking external TDNN asymptotic-normality results.

Assumption 5 (TDNN Variance-Design Conditions).
* For the jackknife variance results, assume the following design and second-moment regularity conditions.

(i) Compact Support: The feature space \(\mathcal{X}= \operatorname{supp}(X)\) is a bounded, compact subset of \(\mathbb{R}^k\).

(ii) Density Bounds: The density \(f(\cdot)\) is bounded away from 0 and \(\infty\), i.e., \[\forall u \in \mathcal{X}: \quad 0 < \underline{\mathfrak{f}} \leq f(u) \leq \overline{\mathfrak{f}} < \infty.\]

(iii) Regression Regularity: \(\mu(\cdot)\) is continuous on \(\mathcal{X}\) and \(\mu(\cdot) \in L^{2}\left(\mathcal{X}\right)\) with respect to \(\varphi\).

(iv) Variance Regularity: \(\sigma^2_{\varepsilon}: \mathcal{X}\longrightarrow \mathbb{R}_{>0}\) is continuous on \(\mathcal{X}\) and square-integrable with respect to \(\varphi\). Since \(\mathcal{X}\) is compact, we write \[0 < \underline{\sigma}_{\varepsilon}^{2} \leq \sigma^2_\varepsilon(u) \leq \overline{\sigma}_{\varepsilon}^{2} < \infty, \qquad u \in \mathcal{X}.\]

Assumption 6 (TDNN Smoothness Inputs for External CLT Results).
* When applying the external TDNN asymptotic-normality result of [4], assume in addition that \(f(\cdot)\) and \(\mu(\cdot)\) are four times continuously differentiable with bounded second, third, and fourth-order partial derivatives. Specifically, for all \(u \in \mathcal{X}\) and all \((i,j,l,m) \in [k]^4\), \[\begin{align} {14} -\infty & \; < \; & \underline{\mathfrak{f}}^{\prime} & \; \leq \; & & \partial_{i,j} f(u), \; & & \partial_{i,j,m} f(u), \; & & \partial_{i,j,l,m} f(u) & & \; \leq \; & \overline{\mathfrak{f}}^{\prime} & \; < \; & \infty \\ -\infty & \; < \; & \underline{\mathfrak{m}}^{\prime} & \; \leq \; & & \partial_{i,j} \mu(u), \; & & \partial_{i,j,m} \mu(u), \; & & \partial_{i,j,l,m} \mu(u) & & \; \leq \; & \overline{\mathfrak{m}}^{\prime} & \; < \; & \infty. \end{align}\] For the heteroskedastic localization adjustment in the preceding remark, also assume the conditional third-moment function \(m_3(u)\coloneq\Ex*{\varepsilon^3 \;\middle|\;X=u}\) exists in a neighborhood of \(x\) and is continuous at \(x\) whenever the corresponding fourth-moment expansion is used.

Remark 4. The bounded-derivatives condition in 6 is used only by external TDNN asymptotic-normality inputs invoked through 5. 3 4 require only 5, together with the response moment condition below for the row-wise square-LLN. In particular, the global fourth-order differentiability could be replaced by a local condition near \(x\) at the cost of a more involved bias argument, without affecting the jackknife consistency result.

We now define the DNN and TDNN kernels under this setup. The single-scale distributional nearest-neighbor (DNN) estimator, developed in the U-statistic-type setting of [27] and [28], averages over subsamples after keeping the observation closest to the target point \(x\) within each subsample. To describe that selection rule, order the sample by distance to the fixed feature vector of interest \(x\): \[\label{eq:ordering} \norm*{X_{(1)} - x}_2 < \norm*{X_{(2)} - x}_2 < \dotsc < \norm*{X_{(n)} - x}_2\tag{8}\] The ordering in 8 is well-defined because \(X\) is continuously distributed; for notational simplicity, tied observations receive the same rank. Let \(\mathop{\mathrm{rk}}(x; X, D)\) denote the rank relative to a point of interest \(x\) that would be assigned to an observation with covariate vector \(X\) if it was added to a sample \(D\). The data-driven selector \[\kappa(x; Z_{i}, D_{\ell}) = \1*{\mathop{\mathrm{rk}}(x; X_{i}, D_{\ell}) = 1}\] therefore records whether observation \(i\) is the nearest neighbor to \(x\) within the subsample \(D_\ell\). With kernel \(h_{s}(x; D_{\ell}) \mathrel{\vcenter{:}}=\sum_{i = 1}^{n} \1*{i \in \ell} \kappa(x; Z_{i}, D_{\ell}) Y_{i}\), averaging over all subsamples gives the DNN estimator its U-statistic representation: \[\label{eq:U95stat} \tilde{\mu}_{s}(x; D_{[n]}) = \binom{n}{s}^{-1} \sum_{\ell \in L_{n,s}} h_{s}(x; D_{\ell})\tag{9}\] The TDNN estimator combines two such DNN averages at subsampling scales \(1 \leq s_1 < s_2 \leq n\). The two-scale combination eliminates the leading bias term, analogously to higher-order kernels in nonparametric kernel regression. Write \(\mathfrak{S}=(s_1,s_2)\) and use this subscript for quantities associated with the TDNN sequence. The corresponding bias-correction weights are \[w_{1}^{*} = \frac{1}{1-(s_1/s_2)^{-2/k}} \quad\text{and}\quad w_2^{*} = 1 - w_{1}^{*}(s_1, s_2)\] With these weights, [4] define the TDNN estimator as a weighted combination of the two DNN scales. Because the smaller-scale DNN average can be embedded inside each \(s_2\)-subsample, the same estimator can be written as an order-\(s_2\) generalized U-statistic: \[\begin{align} \widehat{\mu}_{\mathfrak{S}}\left(x; D_{[n]}\right) & = w_{1}^{*}\tilde{\mu}_{s_1}\left(x; D_{[n]}\right) + w_2^{*}\tilde{\mu}_{s_2}\left(x; D_{[n]}\right) = \binom{n}{s_2}^{-1} \sum_{\ell \in L_{n,s_2}} h_{\mathfrak{S}}(x; D_{\ell}) \end{align}\] The corresponding order-\(s_2\) TDNN kernel is \[\begin{align} h_{\mathfrak{S}}\left(x; D_{[s_2]}\right) & = w_{1}^{*}\left[\binom{s_2}{s_1}^{-1}\sum_{\ell \in L_{s_2, s_1}} h_{s_1}\left(x; D_{\ell}\right)\right] + w_{2}^{*} h_{s_2}\left(x; D_{[s_2]}\right) \\ & = w_{1}^{*} \tilde{\mu}_{s_1}\left(x; D_{[s_2]}\right) + w_{2}^{*} h_{s_2}\left(x; D_{[s_2]}\right) \end{align}\] A bounded-ratio condition prevents the TDNN estimator from degenerating into an effectively single-scale DNN estimator.

Assumption 7 (Bounded Ratio of Kernel-Orders).
* There is a constant \(\mathfrak{c} \in (0,1/2)\) such that the ratio of kernel orders is bounded in the following way. \[\forall n: \quad 0 < \mathfrak{c} \leq s_1 / s_2 \leq 1 - \mathfrak{c} < 1.\]

If \(s_1/s_2\) is too close to \(1\), the two scales are nearly redundant; if it is too small, the bias-correction weights become overly uneven.

The inferential gain comes from applying the general jackknife theorem to this order-\(s_2\) kernel. In the original paper, consistency of the jackknife variance estimator is established under the strong condition \(s_{2} = o(n^{1/3})\), which narrows the range of kernel orders for which jackknife-based standard errors are theoretically justified. The projection argument developed here relaxes the variance-estimation requirement to \(s_{2} = o(n)\), thereby enlarging the regime in which the TDNN estimator admits consistent jackknife variance estimation. Studentized pointwise inference then follows in any subregime where a TDNN asymptotic-normality result is available on the same variance scale. Here the variance target is localized at \(x\), so we write \[\label{eq:TDNN95Var} \sigma^{2}_{n}(x) = \Varb*{\widehat{\mu}_{\mathfrak{S}}\left(x; D_{[n]}\right)}\tag{10}\] We also include \(x\) in the jackknife variance estimators to make the variance target in 10 explicit.

Theorem 3 (The TDNN Estimator satisfies the Asymptotic Hájek Dominance Condition).
* Consider a data-generating process as outlined in 4 and 5. Let \(0 < s_1 < s_2 = o(n)\) be such that 7 holds. Then the single-scale DNN and two-scale TDNN regression estimators satisfy the Asymptotic Hájek Dominance Condition.

7 first establishes the required single-scale DNN bounds. 8 then lifts those bounds to TDNN by decomposing the two-scale kernel into an embedded \(s_1\)-scale DNN average inside an \(s_2\)-sample plus the ordinary \(s_2\)-scale DNN kernel. The key point is that the TDNN first projection is a controlled linear combination of two single-scale DNN first projections, and the selector coefficients do not cancel its first-order variance.

Thus 3 supplies the variance-dominance half of the general jackknife theory for TDNN estimation. To turn that dominance statement into jackknife ratio consistency, it remains to verify the row-wise square-LLN input. The moment condition needed for that step is weaker and more local than the smoothness assumptions used for asymptotic normality and first-order bias removal in the TDNN literature. Because the nearest-neighbor selector is a function of the covariates, the response envelope is naturally stated conditionally on the covariate value.

Assumption 8 (Response Moment Condition).
* There exist constants \(\eta > 0\) and \(C_Y < \infty\) such that \[\sup_{u \in \mathcal{X}} \Ex*{\abs*{Y}^{2+\eta} \;\middle|\;X = u} \leq C_Y.\]

A global unconditional \((2+\eta)\)-moment bound would suffice, but it is stronger than needed because the selector probes the response distribution only through neighborhoods of \(x\) with nonnegligible weight. Since \(\mu\) is bounded on the compact support under 5 [asm:tdnn95design95compact] [asm:tdnn95design95regression], the same assumption also gives a uniform conditional \((2+\eta)\)-moment bound for \(\varepsilon=Y-\mu(X)\). In the row-wise verifications below we use the fixed exponent \(r_\eta \mathrel{\vcenter{:}}= 1 + \frac{1}{2}\min\{\eta,2\}\), which lies in \((1,2]\) and satisfies \(2r_\eta \leq 2+\eta\). At this exponent, 18 verifies the row-wise \(L^r\) square-LLN condition from 2 for the single-scale DNN first projection, and 20 proves the TDNN analogue through 39 and the resulting moment bound. Together with Hájek dominance, this gives the inputs for the following variance-consistency theorem.

Theorem 4 (Ratio-Consistent Variance Estimation for the TDNN Estimator).
* Let \(0 < s_1 < s_2 = o(n)\), \(1 \leq d=d_n\), \(s_2+d\leq n\) eventually, and \(s_2 d = o(n)\). Under the data-generating process outlined in 4 5 8, and for kernel orders satisfying 7, the associated jackknife variance estimators for the TDNN regression estimator satisfy the following ratio consistency result: \[\frac{\widehat{\sigma}_{JKD}^2\left(x, D_{[n]}; d\right)}{\sigma_{n}^{2}\left(x\right)} \xrightarrow{\mathbb{P}}1.\] In particular, for \(d = 1\), \[\frac{\widehat{\sigma}_{JK}^2\left(x, D_{[n]}\right)}{\sigma_{n}^{2}\left(x\right)} \xrightarrow{\mathbb{P}}1.\]

4 is the general complete-case jackknife theorem specialized to the localized TDNN kernel. Once 3 supplies Hájek dominance at order \(s_2\) and 18 20 verify the row-wise square-LLN input under 8, 1 applies without further structural changes. The same verification also pins down the variance scale: \(\sigma_n^2(x) = O(s_2/n)\), so the local standard error contracts at rate \(\sqrt{s_2/n}\).

We do not treat incomplete variants here because the DNN and TDNN regression estimators admit convenient closed-form representations. The preceding variance result has an immediate inference implication: any TDNN asymptotic-normality statement on the \(\sigma_n(x)\) scale can be combined with the jackknife variance estimator to obtain feasible studentization. We state that implication explicitly.

Theorem 5 (Studentized Inference from TDNN Asymptotic Normality).
* Let \(0 < s_1 < s_2 = o(n)\), \(1 \leq d=d_n\), \(s_2+d\leq n\) eventually, and \(s_2 d = o(n)\). Under the data-generating process outlined in 4 5 8, and for kernel orders satisfying 7, suppose that \[\frac{\widehat{\mu}_{\mathfrak{S}}\left(x; D_{[n]}\right) - \mu\left(x\right)}{\sigma_{n}\left(x\right)} \rightsquigarrow\mathcal{N}\left(0,1\right).\] Then the associated jackknife variance estimators yield the following studentized limit: \[\frac{\widehat{\mu}_{\mathfrak{S}}\left(x; D_{[n]}\right) - \mu\left(x\right)}{\sqrt{\widehat{\sigma}_{JKD}^2\left(x, D_{[n]}; d\right)}} \rightsquigarrow\mathcal{N}\left(0,1\right).\] In particular, for \(d = 1\), \[\frac{\widehat{\mu}_{\mathfrak{S}}\left(x; D_{[n]}\right) - \mu\left(x\right)}{\sqrt{\widehat{\sigma}_{JK}^2\left(x, D_{[n]}\right)}} \rightsquigarrow\mathcal{N}\left(0,1\right).\]

5 is a direct Slutsky corollary of 4. When a TDNN central limit theorem is available on the variance scale \(\sigma_n(x)\), jackknife ratio consistency converts it into feasible pointwise inference. For Theorem 3 of [4], one must also verify the external smoothness inputs in 6 and check that its residual bias term is negligible relative to \(\sigma_n(x)\) after the heteroskedastic localization adjustment above.

Thus, whenever TDNN is asymptotically normal on its own variance scale, jackknife studentization yields valid inference without further modifying the variance estimator.

Remark 5. Several other base learners have already been investigated by [14] for whether they satisfy the asymptotic Hájek dominance condition. These include classical U-statistics such as the mean and sample variance, giving another view of jackknife consistency in well-understood cases. They also include classical k-nearest neighbors (kNN), k-potential nearest neighbors (kPNN), and tree-based learners such as honest and double-sampling trees. Thus the results here apply immediately to other generalized U-statistics once the remaining square-LLN input is verified. More broadly, DNN localization applied to targets other than the conditional mean — such as conditional quantile regression, parameters defined by localized conditional estimating equations, or other localized distributional functionals at a point — falls within the same framework, since 1 is governed by the localizer geometry rather than the specific functional being estimated.

5 Conclusion and Outlook↩︎

This paper establishes a simple criterion for consistent nonparametric jackknife variance estimation for generalized U-statistics: asymptotic dominance of the Hájek projection variance. The same first-order dominance condition already plays a central role in generalized U-statistic asymptotics, and our results show that it also governs jackknife validity. For complete generalized U-statistics, the jackknife and its delete-\(d\) variants are ratio-consistent without modifying the base estimator. For incomplete generalized U-statistics, the same conclusion holds under an additional asymptotically sufficient sampling condition. The TDNN application shows that the results are not merely abstract: they deliver consistent jackknife standard errors throughout the broader regime \(s_{2} = o(n)\), substantially extending the range previously justified in [4], and turn any TDNN asymptotic-normality result on that variance scale into feasible pointwise inference.

The same projection-based viewpoint suggests several extensions. In degenerate problems, the relevant dominance condition should shift from the first projection to the first nonvanishing Hoeffding component. For plug-in and forest-type procedures, the main task is to verify that the localized or subsampling-based first-stage object still induces a controlled projection variance of the kind studied here. A related direction is conditional Z-estimation, where DNN localizer weights enter a score-based rather than regression-based kernel; the Hájek-dominance machinery applies once the score-kernel first-projection variance is controlled, with leave-one-out stability playing a complementary role on the bias side. The present paper does not resolve those broader questions. It instead isolates projection dominance as a tractable criterion for when simple jackknife-based uncertainty quantification remains valid for modern nonparametric and machine-learning estimators, including settings of econometric interest.

6 General Results for Delete-\(d\) Jackknife Consistency↩︎

6.1 Notation↩︎

We collect the overlap bookkeeping used throughout the delete-\(d\) jackknife arguments.

Remark 6 (Delete-difference notation).
* Let \(V\) be any statistic with a corresponding deleted-sample evaluation. Write \[\Delta_{\ell}[V] \mathrel{\vcenter{:}}= V(D_{[n]}) - V(D_{[n],-\ell}), \qquad \ell \in L_{n,d}.\] The sign convention is immaterial inside the jackknife square, but this orientation aligns with the Hoeffding-order expansions below. When the statistic is clear from context, we write simply \(\Delta_{\ell}\); for auxiliary statistics, we add the corresponding adornment, such as \(\overline{\Delta}_{\ell}\).

Remark 7 (Delete-\(d\) overlap bookkeeping).
* Fix a delete set \(\ell \in L_{n,d}\), an order \(1 \leq j \leq s\), and an overlap level \(0 \leq a \leq \min\{j,d\}\). In the delete-\(d\) proofs this bookkeeping is used along rows with \(s+d\leq n\), so the post-deletion coefficients \(\binom{n-d}{j}^{-1}\) are well-defined for \(1\leq j\leq s\). We write \[\mathcal{O}_{j,a}(\ell) \mathrel{\vcenter{:}}= \left\{\iota \in L_{n,j} : \abs*{\iota \cap \ell} = a\right\}.\] This partitions \(L_{n,j}\) by the number of deleted indices captured by the tuple \(\iota\). By direct counting, \[\abs*{\mathcal{O}_{j,a}(\ell)} = \binom{d}{a}\binom{n-d}{j-a}.\] For \(1 \leq j \leq s-1\), define the grouped projection sums \[G_{j,a}^{(\ell)}\left(h_{s,\omega}^{(j)}; D_{[n]}\right) \mathrel{\vcenter{:}}= \sum_{\iota \in \mathcal{O}_{j,a}(\ell)} h_{s,\omega}^{(j)}(D_{\iota}),\] and for the final kernel level set \[G_{s,a}^{(\ell)}\left(h_{s}^{(s)}; D_{[n]}\right) \mathrel{\vcenter{:}}= \sum_{\iota \in \mathcal{O}_{s,a}(\ell)} h_{s}^{(s)}(D_{\iota}).\] Since the ambient kernel and sample will always be clear in the delete-\(d\) variance proofs, we suppress these arguments below and write simply \(G_{j,a}^{(\ell)}\) and \(G_{s,a}^{(\ell)}\). To keep track of the overlap combinatorics and the two coefficient types that appear in the delete-\(d\) difference, define \[p_{j,a}^{(n,d)} \mathrel{\vcenter{:}}= \frac{\binom{d}{a}\binom{n-d}{j-a}}{\binom{n}{j}}, \qquad q_{j}^{(n,d)} \mathrel{\vcenter{:}}= \binom{n}{j}^{-1}, \qquad r_{j}^{(n,d)} \mathrel{\vcenter{:}}= \binom{n}{j}^{-1} - \binom{n-d}{j}^{-1}.\] In particular, \(\sum_{a=0}^{\min\{j,d\}} p_{j,a}^{(n,d)} = 1\). With this notation, the delete-\(d\) difference of the \(j\)th Hoeffding projection can be written as \[\begin{align} \delta_{j,\ell} & \mathrel{\vcenter{:}}= \Delta_{\ell}[H_s^j] \\ & = \binom{n}{j}^{-1}\sum_{\iota \in L_{n,j}} h_{s,\omega}^{(j)}(D_{\iota}) - \binom{n-d}{j}^{-1} \sum_{\iota \in L_{j}\left([n]\backslash \ell\right)} h_{s,\omega}^{(j)}(D_{\iota}) \\ & = q_{j}^{(n,d)} \sum_{a=1}^{\min\{j,d\}} G_{j,a}^{(\ell)} + r_{j}^{(n,d)} G_{j,0}^{(\ell)} \end{align}\] for \(1 \leq j \leq s-1\), and the final-order difference is \[\delta_{s,\ell} \mathrel{\vcenter{:}}= \Delta_{\ell}[H_s^s] = q_{s}^{(n,d)} \sum_{a=1}^{\min\{s,d\}} G_{s,a}^{(\ell)} + r_{s}^{(n,d)} G_{s,0}^{(\ell)}.\]

6.2 Preliminary Lemmas↩︎

To show the consistency of the jackknife variance estimators under consideration, we will in part rely on the following basic result from [21].

Lemma 2 ([21] - Lemma C.1.).
* Let \(I_n\) be a finite index set. Suppose that \[\sum_{\alpha\in I_n} X_{\alpha,n}^2 \xrightarrow{\mathbb{P}}1, \qquad \sum_{\alpha\in I_n} \Ex*{X_{\alpha,n}^2} \longrightarrow 1, \qquad \sum_{\alpha\in I_n} \Ex*{Y_{\alpha,n}^2} \longrightarrow 0.\] Then \[\sum_{\alpha\in I_n}\left[X_{\alpha,n}+Y_{\alpha,n}\right]^2 \xrightarrow{\mathbb{P}}1 \quad \text{and} \quad \Ex*{\sum_{\alpha\in I_n}\left(X_{\alpha,n}+Y_{\alpha,n}\right)^2} \longrightarrow 1.\]

Proof. Expand \(\sum_{\alpha\in I_n}(X_{\alpha,n}+Y_{\alpha,n})^2\) into the three terms \[\sum_{\alpha\in I_n} X_{\alpha,n}^2 + 2\sum_{\alpha\in I_n} X_{\alpha,n}Y_{\alpha,n} + \sum_{\alpha\in I_n} Y_{\alpha,n}^2.\] The first term converges in probability to \(1\) by assumption. For the last term, \(\Ex*{\sum_{\alpha\in I_n} Y_{\alpha,n}^2} \to 0\) by assumption, so \(\sum_{\alpha\in I_n} Y_{\alpha,n}^2 \xrightarrow{\mathbb{P}}0\) by Markov’s inequality. For the cross term, Cauchy–Schwarz gives \[\abs*{2\sum_{\alpha\in I_n} X_{\alpha,n}Y_{\alpha,n}} \leq 2 \left(\sum_{\alpha\in I_n} X_{\alpha,n}^2\right)^{1/2} \left(\sum_{\alpha\in I_n} Y_{\alpha,n}^2\right)^{1/2} \xrightarrow{\mathbb{P}}0.\] For expectations, the same inequality and Cauchy–Schwarz give \[\Ex*{\abs*{2\sum_{\alpha\in I_n} X_{\alpha,n}Y_{\alpha,n}}} \leq 2\left(\sum_{\alpha\in I_n}\Ex*{X_{\alpha,n}^2}\right)^{1/2} \left(\sum_{\alpha\in I_n}\Ex*{Y_{\alpha,n}^2}\right)^{1/2} \longrightarrow 0.\] Together with \(\sum_{\alpha\in I_n}\Ex*{X_{\alpha,n}^2}\to1\) and \(\sum_{\alpha\in I_n}\Ex*{Y_{\alpha,n}^2}\to0\), this proves the expectation claim. ◻

Lemma 3 (Square-LLN for the first projection).
* Under 2, \[\frac{1}{n\zeta_{s,\omega}^{1}} \sum_{i=1}^{n} h_s^{(1)}(Z_i)^2 \xrightarrow{\mathbb{P}}1.\]

This lemma controls the diagonal term that appears when the jackknife differences are squared. In both the complete and incomplete proofs, the leading Hájek contribution produces an averaged square of first-projection values, and this normalized diagonal average converges to its population target on the \(\zeta_{s,\omega}^{1}\) scale. Once the higher-order remainder has been shown negligible, this lemma transfers the linear variance scale to the jackknife.

Proof. This is exactly the row-wise i.i.d.triangular-array weak law highlighted in Section 2; see [25]. In the present setting, the proof is a direct application of the von Bahr–Esseen inequality [26]. Write \[W_{n,i} \mathrel{\vcenter{:}}= \frac{h_s^{(1)}(Z_i)^2}{\zeta_{s,\omega}^{1}}, \qquad X_{n,i} \mathrel{\vcenter{:}}= W_{n,i} - 1, \qquad i = 1,\dotsc,n.\] Because the first projection is centered, \(\Ex{W_{n,1}} = 1\), so the variables \(X_{n,1},\dotsc,X_{n,n}\) are row-wise i.i.d.and mean zero. Let \(r \in (1,2]\) be as in 2. Then the von Bahr–Esseen inequality gives \[\mathbb{E}\left[ \abs*{ \frac{1}{n}\sum_{i=1}^{n} X_{n,i} }^{r} \right] \leq \frac{2}{n^{r}} \sum_{i=1}^{n} \Ex*{\abs*{X_{n,i}}^{r}} = \frac{2\Ex*{\abs*{X_{n,1}}^{r}}}{n^{r-1}}.\] Since \(r \geq 1\), \[\abs*{x-1}^{r} \leq 2^{r-1}\left(\abs*{x}^{r} + 1\right),\] so \[\frac{\Ex*{\abs*{X_{n,1}}^{r}}}{n^{r-1}} \lesssim \frac{\Ex*{W_{n,1}^{r}} + 1}{n^{r-1}} \longrightarrow 0\] by 2. Therefore, \[\frac{1}{n\zeta_{s,\omega}^{1}} \sum_{i=1}^{n} h_s^{(1)}(Z_i)^2 = \frac{1}{n}\sum_{i=1}^{n} W_{n,i} \xrightarrow{\mathbb{P}}1.\] ◻

Proof of 1. \[\begin{align} 1 & \leq \frac{n}{s^2}\frac{\Varb*{\mathrm{U}_{n, s, \omega}\left(D_{[n]}\right)}}{\zeta_{s, \omega}^{1}} = \left(\frac{s^2}{n} \zeta_{s, \omega}^{1}\right)^{-1} \left(\sum_{j=1}^{s-1}\binom{s}{j}^2\binom{n}{j}^{-1} V_{s, \omega}^{j} + \binom{n}{s}^{-1} V_{s}^{s}\right) \\ & \leq 1 + \left(\frac{s^2}{n} \zeta_{s, \omega}^{1}\right)^{-1} \frac{s^2}{n^2} \left(\sum_{j=2}^{s-1}\binom{s}{j} V_{s, \omega}^{j} + V_{s}^{s}\right) \\ & \leq 1 + \frac{s}{n}\left(\frac{\zeta_{s}^{s}}{s\zeta_{s, \omega}^{1}} - 1\right) \longrightarrow 1. \end{align}\] ◻

6.3 Complete Generalized U-Statistics↩︎

Proof of 1.
* For each delete set \(\ell \in L_{n,d}\), write \[\Delta_{\ell} \mathrel{\vcenter{:}}= \Delta_{\ell}[U_{n,s,\omega}].\] Then the delete-\(d\) jackknife variance estimator takes the form \[\widehat{\sigma}_{JKD}^2\left(D_{[n]}; d\right) = \frac{n-d}{d}\binom{n}{d}^{-1}\sum_{\ell \in L_{n,d}} \Delta_{\ell}^{2}.\] We first expand each delete-\(d\) difference by Hoeffding order and isolate its first-projection, or Hájek, component. We then show that the remaining higher-order terms are negligible on the \(\zeta_{s,\omega}^{1}\) scale, identify the averaged square of the leading term with the empirical average of \(h_i^2\) up to negligible off-diagonal corrections, and finally invoke 3 2. The last step is to compare the resulting scale back to the actual variance target via 1. By the generalized Hoeffding decomposition, \[\Delta_{\ell} = \sum_{j=1}^{s-1}\binom{s}{j}\delta_{j,\ell} + \delta_{s,\ell},\] where the projection-level differences \(\delta_{j,\ell}\), \(1 \leq j \leq s\), are defined in 7. To isolate the leading Hájek term, write \[h_{i} \mathrel{\vcenter{:}}= h_{s,\omega}^{(1)}(Z_{i}), \qquad i = 1, \dotsc, n.\] Since \[G_{1,1}^{(\ell)} = \sum_{i \in \ell} h_{i}, \qquad G_{1,0}^{(\ell)} = \sum_{i \in [n]\backslash \ell} h_{i},\] and \[q_{1}^{(n,d)} = \frac{1}{n}, \qquad r_{1}^{(n,d)} = \frac{1}{n} - \frac{1}{n-d} = -\frac{d}{n(n-d)},\] the first projection contributes \[s \delta_{1,\ell} = \frac{s}{n}\sum_{i \in \ell} h_{i} - \frac{sd}{n(n-d)}\sum_{i \in [n]\backslash \ell} h_{i}.\] We therefore write \[\Delta_{\ell} = \frac{s\sqrt{d}}{n} \left[ \frac{1}{\sqrt{d}}\sum_{i \in \ell} h_{i} + T_{\ell} \right],\] where \[T_{\ell} \mathrel{\vcenter{:}}= -\frac{\sqrt{d}}{n-d}\sum_{i \in [n]\backslash \ell} h_{i} + \frac{n}{s\sqrt{d}} \left\{ \sum_{j=2}^{s-1}\binom{s}{j}\delta_{j,\ell} + \delta_{s,\ell} \right\}.\] Consequently, \[\widehat{\sigma}_{JKD}^2\left(D_{[n]}; d\right) = \frac{(n-d)s^2}{n^2}\binom{n}{d}^{-1} \sum_{\ell \in L_{n,d}} \left[ \frac{1}{\sqrt{d}}\sum_{i \in \ell} h_{i} + T_{\ell} \right]^2.\] Write \[A_{\ell} \mathrel{\vcenter{:}}= \frac{1}{\sqrt{d}}\sum_{i \in \ell} h_{i}.\] We first show that \(T_{\ell}\) is negligible relative to the variance scale \(\zeta_{s,\omega}^{1}\). Since distinct Hoeffding orders are orthogonal, we have \[\Ex*{T_{\ell}^{2}} = \frac{d}{n-d}\zeta_{s,\omega}^{1} + \frac{n^2}{s^2 d} \left\{ \sum_{j=2}^{s-1}\binom{s}{j}^{2}\Ex*{\delta_{j,\ell}^{2}} + \Ex*{\delta_{s,\ell}^{2}} \right\}.\] The key simplification is that, once the delete-\(d\) difference has been organized by Hoeffding order and overlap strata, the second-moment calculation is mostly combinatorial. Distinct Hoeffding orders are orthogonal, and within a fixed canonical order the degeneracy of the projection kills the mixed overlap terms. What remains is to count how many tuples land in each overlap stratum and multiply by the variance attached to that order. The useful partition is therefore not really the pair of cases \(d<j\) and \(d \geq j\), but the simpler split between \(j\)-tuples that hit the delete set and \(j\)-tuples that avoid it. The exact overlap level \(a=\abs*{\iota\cap\ell}\) is only a device for counting the first class: all nonzero-overlap strata carry the same coefficient \(q_j^{(n,d)}\), while the zero-overlap stratum carries \(r_j^{(n,d)}\). With the standard convention that impossible binomial coefficients are zero, Vandermonde’s identity gives \[\sum_{a=1}^{\min\{j,d\}}\binom{d}{a}\binom{n-d}{j-a} = \binom{n}{j}-\binom{n-d}{j},\] while \(\abs*{\mathcal{O}_{j,0}(\ell)}=\binom{n-d}{j}\). This is why the two apparent regimes collapse to the same coefficient after squaring and taking expectations. For each \(2 \leq j \leq s-1\), the overlap-stratum representation and degeneracy of the canonical \(j\)th projection imply \[\begin{align} \Ex*{\delta_{j,\ell}^{2}} & = \left[ \left(q_{j}^{(n,d)}\right)^{2} \sum_{a=1}^{\min\{j,d\}} \abs*{\mathcal{O}_{j,a}(\ell)} + \left(r_{j}^{(n,d)}\right)^{2}\abs*{\mathcal{O}_{j,0}(\ell)} \right] V_{s,\omega}^{j} \\ & = \left[ \binom{n-d}{j}^{-1} - \binom{n}{j}^{-1} \right] V_{s,\omega}^{j}, \end{align}\] and likewise \[\Ex*{\delta_{s,\ell}^{2}} = \left[ \binom{n-d}{s}^{-1} - \binom{n}{s}^{-1} \right] V_{s}^{s}.\] Substituting these identities gives \[\begin{align} \Ex*{T_{\ell}^{2}} & = \frac{d}{n-d}\zeta_{s,\omega}^{1} \\ & \quad + \frac{n}{s d}\Bigg\{ \sum_{j=2}^{s-1} \frac{\binom{s-1}{j-1}}{\binom{n-1}{j-1}} \left( \binom{n}{j}\binom{n-d}{j}^{-1} - 1 \right)\binom{s}{j}V_{s,\omega}^{j} \\ & \qquad\qquad + \binom{n-1}{s-1}^{-1} \left( \binom{n}{s}\binom{n-d}{s}^{-1} - 1 \right)V_{s}^{s} \Bigg\}. \end{align}\] For \(2 \leq j \leq s\), write \(B_{j,n,d}\) for the amount by which the \(j\)-tuple normalization changes after deleting \(d\) observations. The theorem’s row-wise delete-domain condition \(s+d\leq n\) ensures that the denominator \(\binom{n-d}{j}\) is nonzero for every \(2\leq j\leq s\) considered here. Since each deleted position affects a \(j\)-tuple only at order \(j/n\), the product below formalizes the heuristic that the cumulative perturbation is \(O(jd/n)\). \[B_{j,n,d} \mathrel{\vcenter{:}}= \binom{n}{j}\binom{n-d}{j}^{-1} = \prod_{m=0}^{d-1}\frac{n-m}{n-j-m} = \prod_{m=0}^{d-1}\left(1 + \frac{j}{n-j-m}\right).\] Since \(sd = o(n)\), we have \(jd/(n-j-d+1) \leq 1/2\) uniformly for \(2 \leq j \leq s\) and \(n\) large enough. A Weierstrass-product bound then yields \[B_{j,n,d} - 1 \lesssim \frac{2jd}{n-j-d+1} \lesssim \frac{3jd}{n}, \qquad 2 \leq j \leq s.\] Using also \[\frac{\binom{s-1}{j-1}}{\binom{n-1}{j-1}} \leq \left(\frac{e(s-1)}{n-1}\right)^{j-1},\] we obtain \[\begin{align} \Ex*{T_{\ell}^{2}} & \lesssim \frac{d}{n-d}\zeta_{s,\omega}^{1} + \frac{3}{s}\Bigg\{ \sum_{j=2}^{s-1} j \left(\frac{e(s-1)}{n-1}\right)^{j-1} \binom{s}{j}V_{s,\omega}^{j} + s\left(\frac{e(s-1)}{n-1}\right)^{s-1}V_{s}^{s} \Bigg\}. \end{align}\] Splitting \(j = 1 + (j-1)\), bounding the first part by \(\frac{e(s-1)}{n-1}\), and using \[\zeta_{s}^{s} = \sum_{j=1}^{s-1}\binom{s}{j}V_{s,\omega}^{j} + V_{s}^{s},\] we further obtain \[\begin{align} \Ex*{T_{\ell}^{2}} & \lesssim \frac{d}{n-d}\zeta_{s,\omega}^{1} + \frac{3e(s-1)}{s(n-1)} \left(\zeta_{s}^{s} - s\zeta_{s,\omega}^{1}\right) \\ & \quad + \frac{3}{s}\zeta_{s}^{s} \sum_{j=1}^{\infty} j\left(\frac{e(s-1)}{n-1}\right)^{j}. \end{align}\] Evaluating the geometric derivative series gives \[\begin{align} \Ex*{T_{\ell}^{2}} & \lesssim \Bigg[ \frac{d}{n-d} + \frac{3e(s-1)(n-1)}{\left(n-1-e(s-1)\right)^{2}} \Bigg]\zeta_{s,\omega}^{1} \\ & \quad + \frac{3e(s-1)}{s} \Bigg[ \frac{1}{n-1} + \frac{n-1}{\left(n-1-e(s-1)\right)^{2}} \Bigg] \left(\zeta_{s}^{s} - s\zeta_{s,\omega}^{1}\right). \end{align}\] Therefore, \[\frac{\Ex*{T_{\ell}^{2}}}{\zeta_{s,\omega}^{1}} \longrightarrow 0\] by \(sd = o(n)\), \(s = o(n)\), and 1. Now define \[X_{\ell} \mathrel{\vcenter{:}}= \sqrt{ \frac{n-d}{n\binom{n}{d}\zeta_{s,\omega}^{1}} }\,A_{\ell}, \qquad Y_{\ell} \mathrel{\vcenter{:}}= \sqrt{ \frac{n-d}{n\binom{n}{d}\zeta_{s,\omega}^{1}} }\,T_{\ell}.\] Then \[\sum_{\ell \in L_{n,d}} \left(X_{\ell} + Y_{\ell}\right)^{2} = \frac{n\widehat{\sigma}_{JKD}^{2}\left(D_{[n]}; d\right)}{s^2\zeta_{s,\omega}^{1}}.\] Moreover, by exchangeability, \[\sum_{\ell \in L_{n,d}} \Ex*{Y_{\ell}^{2}} = \frac{n-d}{n\zeta_{s,\omega}^{1}}\Ex*{T_{\ell}^{2}} \longrightarrow 0,\] and \[\sum_{\ell \in L_{n,d}} \Ex*{X_{\ell}^{2}} = \frac{n-d}{n\zeta_{s,\omega}^{1}}\Ex*{A_{\ell}^{2}} = \frac{n-d}{n} \longrightarrow 1.\] It remains to verify \(\sum_{\ell \in L_{n,d}} X_{\ell}^{2} \xrightarrow{\mathbb{P}}1\). This is the diagonal-square step. Averaging \(A_{\ell}^{2}\) over all delete sets extracts the empirical average of the squared first-projection values, while the cross-products are washed out by the combinatorial averaging over deletion patterns. In that sense, the jackknife behaves here as though it were applied to a linear statistic with influence values \(h_1,\dotsc,h_n\). Here the exact averaged-square identity is \[\begin{align} \binom{n}{d}^{-1}\sum_{\ell \in L_{n,d}} A_{\ell}^{2} & = \frac{1}{d}\binom{n}{d}^{-1} \sum_{\ell \in L_{n,d}}\sum_{i \in \ell}\sum_{j \in \ell} h_{i}h_{j} \\ & = \frac{1}{n}\sum_{i=1}^{n} h_{i}^{2} + \frac{d-1}{n(n-1)} \sum_{i \neq j} h_{i}h_{j}. \end{align}\] Note that despite \(d > 1\), the averaged-square identity reduces the diagonal to \(n^{-1}\sum_i h_i^2\)—the same empirical average of squared first-projection values that appears in the delete-\(1\) case and is controlled by 2. The off-diagonal term is negligible because \[\begin{align} \Ex*{\abs*{ \frac{d-1}{n(n-1)} \sum_{i \neq j} h_{i}h_{j} }} & \leq \frac{d-1}{n(n-1)} \Ex*{\abs*{ \left(\sum_{i=1}^{n} h_{i}\right)^{2} - \sum_{i=1}^{n} h_{i}^{2} }} \\ & \leq \frac{2(d-1)}{n-1}\zeta_{s,\omega}^{1} = o\left(\zeta_{s,\omega}^{1}\right). \end{align}\] Hence \[\sum_{\ell \in L_{n,d}} X_{\ell}^{2} = \frac{n-d}{n\zeta_{s,\omega}^{1}} \left[ \frac{1}{n}\sum_{i=1}^{n} h_{i}^{2} + o_{p}\left(\zeta_{s,\omega}^{1}\right) \right].\] By 3, \[\frac{1}{n\zeta_{s,\omega}^{1}}\sum_{i=1}^{n} h_{i}^{2} \xrightarrow{\mathbb{P}}1.\] Granting this input, \(\sum_{\ell \in L_{n,d}} X_{\ell}^{2} \xrightarrow{\mathbb{P}}1\), so 2 yields \[\frac{n\widehat{\sigma}_{JKD}^{2}\left(D_{[n]}; d\right)}{s^2\zeta_{s,\omega}^{1}} \xrightarrow{\mathbb{P}} 1.\] The desired ratio consistency then follows from 1. ◻

6.4 Incomplete Generalized U-Statistics↩︎

Lemma 4 (Reciprocal binomial good-count bounds).
* Let \(K=1+B\), where \(B\sim\mathrm{Binomial}(m-1,p)\), and let \(\mu=mp\to\infty\). Define \(G=\set*{\mu/2\leq K\leq 2\mu}\). Then there are universal constants \(C,c>0\) such that, for all large \(\mu\), \[\Pb*{G^c} \leq C\exp(-c\mu), \qquad \Ex*{K\1*{G^c}} \leq C\mu\exp(-c\mu), \qquad \Ex*{K^2\1*{G^c}} \leq C\mu^2\exp(-c\mu),\] and \[\Ex*{K^{-3}} \leq C\mu^{-3}, \qquad \Ex*{(K-\mu)^2K^{-5}} \leq C\mu^{-4}.\]

Proof. Chernoff’s inequality gives \(\Pb*{G^c}\leq C\exp(-c\mu)\), after reducing \(c\) to absorb the deterministic shift from \(B\) to \(K=1+B\). The same Chernoff bound and the standard identity \[\Ex*{K\1*{G^c}} = \int_0^\infty \Pb*{K\1*{G^c}>t} \,\;\mathrm{d}t\] give \(\Ex*{K\1*{G^c}}\leq C\mu\exp(-c\mu)\), again with a possibly smaller \(c\). Applying the same integration argument to \(K^2\1*{G^c}\) gives \(\Ex*{K^2\1*{G^c}}\leq C\mu^2\exp(-c\mu)\). On \(G\), \(K^{-3}\leq C\mu^{-3}\) and \(K^{-5}\leq C\mu^{-5}\). On \(G^c\), the trivial bound \(K^{-3}\leq 1\) and the tail bound show that the contribution is exponentially small, hence \(O(\mu^{-3})\). Similarly, \[\Ex*{(K-\mu)^2K^{-5}\1*{G}} \leq C\mu^{-5}\Ex*{(K-\mu)^2} \leq C\mu^{-4},\] while on \(G^c\), \((K-\mu)^2K^{-5}\leq (K+\mu)^2\) and the preceding tail estimates make the contribution exponentially small. ◻

Lemma 5 (Zero-count centering transfer).
* Let \(h_s^0=h_s-\theta_s\), where \(\theta_s=\Ex*{h_s(D_{[s]};\omega)}\). Consider the incomplete statistic and delete-\(d\) jackknife estimator under the zero-count conventions in 1 3, using the same Bernoulli variables and auxiliary randomness for \(h_s\) and \(h_s^0\). Suppose \(N_d^\circ/N\to1\), \(N/n\to\infty\), \(\theta_s\) is bounded, and \[\frac{1}{\zeta_{s,\omega}^{1}} = o\left(\frac{Ns}{n}\right).\] If the centered-kernel statistic satisfies \[\frac{n\widehat\sigma_{JKD,h^0}^{2}}{s^2\zeta_{s,\omega}^{1}} = O_p(1), \qquad \frac{n\Varb*{U_{h^0}}}{s^2\zeta_{s,\omega}^{1}} = O(1),\] then \[\frac{n}{s^2\zeta_{s,\omega}^{1}} \abs*{ \widehat\sigma_{JKD,h}^{2} - \widehat\sigma_{JKD,h^0}^{2} } \xrightarrow{\mathbb{P}} 0\] and \[\frac{n}{s^2\zeta_{s,\omega}^{1}} \abs*{ \Varb*{U_h} - \Varb*{U_{h^0}} } \longrightarrow 0.\]

Proof. Write \(A=\set*{\widehat N>0}\) and \(B_\ell=\set*{\widehat N_\ell^\circ>0}\). Since \(B_\ell\subset A\), the zero-count convention gives \[U_h=U_{h^0}+\theta_s\1*{A}, \qquad U_{h,\ell}=U_{h^0,\ell}+\theta_s\1*{B_\ell},\] and hence \[\Delta_{\ell,h} = \Delta_{\ell,h^0} + C_\ell, \qquad C_\ell \mathrel{\vcenter{:}}= \theta_s\1*{A\cap B_\ell^c}.\] If \(H_1\) is the number of selected subsamples intersecting the delete set, then \[A\cap B_\ell^c = \set*{\widehat N_\ell^\circ=0,\;H_1>0}.\] The avoid and hit Bernoulli layers are independent, so \[\Pb*{A\cap B_\ell^c} = (1-p)^{\binom{n-d}{s}} \left\{ 1-(1-p)^{\binom{n}{s}-\binom{n-d}{s}} \right\} \leq \exp(-N_d^\circ).\] Define \[B_n \mathrel{\vcenter{:}}= \frac{n}{s^2\zeta_{s,\omega}^{1}} \frac{n-d}{d}\binom{n}{d}^{-1} \sum_{\ell\in L_{n,d}} C_\ell^2.\] By exchangeability and the preceding bound, \[\Ex*{B_n} \leq \theta_s^2 \frac{n(n-d)}{d\,s^2\zeta_{s,\omega}^{1}} \exp(-N_d^\circ).\] The assumptions imply this upper bound tends to zero: indeed \(1/\zeta_{s,\omega}^{1}=o(Ns/n)\), \(d\geq1\), \(s\geq1\), and \(N_d^\circ\asymp N\) reduce the polynomial factor to \(o(Nn)\), which is dominated by \(\exp(-N_d^\circ)\) because \(N/n\to\infty\). Hence \(B_n\xrightarrow{\mathbb{P}}0\). With \[A_n \mathrel{\vcenter{:}}= \frac{n\widehat\sigma_{JKD,h^0}^{2}}{s^2\zeta_{s,\omega}^{1}},\] Cauchy–Schwarz gives \[\frac{n}{s^2\zeta_{s,\omega}^{1}} \abs*{\widehat\sigma_{JKD,h}^{2}-\widehat\sigma_{JKD,h^0}^{2}} \leq 2A_n^{1/2}B_n^{1/2}+B_n \xrightarrow{\mathbb{P}} 0.\]

For the variance target, \(U_h=U_{h^0}+\theta_s\1*{A}\) and \(\Varb*{\theta_s\1*{A}}\leq\theta_s^2\Pb*{A^c}\leq\theta_s^2\exp(-N)\). The inequality \[\abs*{\Varb*{X+Y}-\Varb*{X}} \leq 2\sqrt{\Varb*{X}\Varb*{Y}}+\Varb*{Y}\] with \(X=U_{h^0}\) and \(Y=\theta_s\1*{A}\), together with the centered variance tightness and the same exponential domination argument, proves the variance-transfer claim. ◻

Proof of 2.
* We first prove the result for the centered kernel, so \(\theta_{s}=0\) throughout the main normalization argument. The zero-count centering-transfer step at the end removes this temporary restriction. The reciprocal-count expressions below use the zero-count conventions from 1 3. Whenever a selected kernel term appears inside the expectation, the corresponding realized count is automatically positive, so the reciprocal manipulations are on positive-count events. Let \[\hbar_{s}\left(D_{[s]}; \rho, \omega\right) \mathrel{\vcenter{:}}= \frac{\rho}{p} h_{s}\left(D_{[s]}; \omega\right), \qquad p \mathrel{\vcenter{:}}= \frac{N}{\binom{n}{s}},\] and define the Horvitz-Thompson normalized auxiliary complete generalized U-statistic \[\overline{U}_{n,s,\rho,\omega}\left(D_{[n]}\right) \mathrel{\vcenter{:}}= \binom{n}{s}^{-1} \sum_{\iota \in L_{n,s}} \hbar_{s}\left(D_{\iota}; \rho_{\iota}, \omega\right) = \frac{1}{N} \sum_{\iota \in L_{n,s}} \rho_{\iota} h_{s}\left(D_{\iota}; \omega\right).\] This differs from \(U_{n,s,N,\omega}\) only through the denominator: the random count \(\widehat N\) is replaced by its expectation \(N\). Working first with \(\overline{U}_{n,s,\rho,\omega}\) separates the random-normalization effect from the delete-\(d\) Hoeffding expansion. The overlap decomposition can then be carried out for a kernel with the same first projection as the original incomplete statistic, and the comparison between \(N\) and \(\widehat N\) is handled afterward as a separate normalization step.

By construction, for every \(1 \leq j \leq s-1\), \[\hbar_{s}^{(j)}\left(D_{[j]}\right) = h_{s}^{(j)}\left(D_{[j]}\right).\] In particular, \[\hbar_{s}^{(1)}(Z_1) = h_{s}^{(1)}(Z_1),\] so the normalized first-projection square is unchanged: \[\frac{\hbar_{s}^{(1)}(Z_1)^2}{\zeta_{s,\omega}^{1}} = \frac{h_{s}^{(1)}(Z_1)^2}{\zeta_{s,\omega}^{1}}.\] Consequently, the Horvitz-Thompson kernel has the same first-projection scale and satisfies the same row-wise square law. It remains to show that Bernoulli sampling noise is a negligible remainder.

Write \(\overline{V}_{s}^{s}\) and \(\overline{\zeta}_{s}^{s}\) for the final-order variance terms corresponding to \(\hbar_{s}\). Since the lower Hoeffding orders are unchanged and \(\theta_{s}=0\), \[\overline{\zeta}_{s}^{s} = \Varb*{\hbar_{s}\left(D_{[s]}; \rho, \omega\right)} = \Varb*{\frac{\rho}{p} h_{s}\left(D_{[s]}; \omega\right)} = \frac{1}{p}\zeta_{s}^{s},\] and therefore \[\overline{V}_{s}^{s} = V_{s}^{s} + \frac{1-p}{p}\zeta_{s}^{s}.\] Thus Horvitz-Thompson normalization leaves every lower Hoeffding order untouched and only inflates the final, purely \(s\)-fold component.

Let \(\overline{\sigma}_{n}^{2} \mathrel{\vcenter{:}}=\Varb*{\overline{U}_{n,s,\rho,\omega}\left(D_{[n]}\right)}\). The variance decomposition from the proof of 1 yields \[\begin{align} 1 & \leq \frac{n}{s^{2}} \frac{\overline{\sigma}_{n}^{2}}{\zeta_{s,\omega}^{1}} \\ & = \left(\frac{s^{2}}{n}\zeta_{s,\omega}^{1}\right)^{-1} \left( \sum_{j=1}^{s-1} \binom{s}{j}^{2}\binom{n}{j}^{-1}V_{s,\omega}^{j} + \binom{n}{s}^{-1}\overline{V}_{s}^{s} \right) \\ & \leq 1 + \frac{s}{n} \left( \frac{\zeta_{s}^{s}}{s\zeta_{s,\omega}^{1}} - 1 \right) + \frac{n}{s^{2}\zeta_{s,\omega}^{1}} \binom{n}{s}^{-1}\frac{1-p}{p}\zeta_{s}^{s} \\ & \leq 1 + \frac{s}{n} \left( \frac{\zeta_{s}^{s}}{s\zeta_{s,\omega}^{1}} - 1 \right) + \frac{n\zeta_{s}^{s}}{N s^{2}\zeta_{s,\omega}^{1}} \longrightarrow 1, \end{align}\] where the last step uses 1 3 and the boundedness of \(\zeta_{s}^{s}\). Thus, \[\label{eq:var95scale95u95bar} \frac{n}{s^{2}} \frac{\overline{\sigma}_{n}^{2}}{\zeta_{s,\omega}^{1}} \longrightarrow 1.\tag{11}\] So the Horvitz-Thompson normalized statistic already lives on the same variance scale as the complete statistic. The asymptotically-sufficient sampling condition is used here in its most transparent role: it says that the extra top-order variance created by Bernoulli thinning is too small to compete with the first projection.

We now analyze the delete-\(d\) jackknife built from \(\overline{U}_{n,s,\rho,\omega}\). For each delete set \(\ell \in L_{n,d}\), define \[\overline{\Delta}_{\ell} \mathrel{\vcenter{:}}= \Delta_{\ell}[\overline{U}_{n,s,\rho,\omega}] = \overline{U}_{n,s,\rho,\omega}\left(D_{[n]}\right) - \overline{U}_{n,s,\rho,\omega}\left(D_{[n],-\ell}\right).\] By the generalized Hoeffding decomposition, \[\overline{\Delta}_{\ell} = \sum_{j=1}^{s-1}\binom{s}{j}\delta_{j,\ell} + \overline{\delta}_{s,\ell},\] where the lower-order terms \(\delta_{j,\ell}\), \(1 \leq j \leq s-1\), are exactly the complete-case projection-level differences from 7, while the final-order term becomes \[\overline{\delta}_{s,\ell} = q_{s}^{(n,d)} \sum_{a=1}^{\min\{s,d\}} \overline{G}_{s,a}^{(\ell)} + r_{s}^{(n,d)} \overline{G}_{s,0}^{(\ell)},\] with \(\overline{G}_{s,a}^{(\ell)}\) denoting the overlap-stratum sums built from \(\hbar_{s}^{(s)}\). Thus the delete-\(d\) expansion differs from the complete-case expansion only through the final-order term. All lower-order overlap combinatorics and leading-term algebra are unchanged, so it remains to bound this modified top-order contribution.

Write \[h_{i} \mathrel{\vcenter{:}}= h_{s,\omega}^{(1)}(Z_i) = \hbar_{s}^{(1)}(Z_i), \qquad i = 1,\dotsc,n.\] Since the first projection is unchanged, the same algebra as in the proof of 1 gives \[\overline{\Delta}_{\ell} = \frac{s\sqrt d}{n} \left[ A_{\ell} + \overline{T}_{\ell} \right],\] where \[A_{\ell} \mathrel{\vcenter{:}}= \frac{1}{\sqrt d} \sum_{i \in \ell} h_{i}\] and \[\overline{T}_{\ell} \mathrel{\vcenter{:}}= -\frac{\sqrt d}{n-d} \sum_{i \in [n]\backslash \ell} h_{i} + \frac{n}{s\sqrt d} \left\{ \sum_{j=2}^{s-1}\binom{s}{j}\delta_{j,\ell} + \overline{\delta}_{s,\ell} \right\}.\]

Orthogonality of distinct Hoeffding orders yields \[\Ex*{\overline{T}_{\ell}^{2}} = \frac{d}{n-d}\zeta_{s,\omega}^{1} + \frac{n^{2}}{s^{2}d} \left\{ \sum_{j=2}^{s-1}\binom{s}{j}^{2}\Ex*{\delta_{j,\ell}^{2}} + \Ex*{\overline{\delta}_{s,\ell}^{2}} \right\}.\] For \(2 \leq j \leq s-1\), the lower-order terms are unchanged, so \[\Ex*{\delta_{j,\ell}^{2}} = \left[ \binom{n-d}{j}^{-1} - \binom{n}{j}^{-1} \right] V_{s,\omega}^{j},\] while the final-order term becomes \[\Ex*{\overline{\delta}_{s,\ell}^{2}} = \left[ \binom{n-d}{s}^{-1} - \binom{n}{s}^{-1} \right]\overline{V}_{s}^{s}.\] Using the same product and geometric-series bounds as in the complete-case analysis gives \[\begin{align} \Ex*{\overline{T}_{\ell}^{2}} & \lesssim \Bigg[ \frac{d}{n-d} + \frac{3e(s-1)(n-1)}{\left(n-1-e(s-1)\right)^{2}} \Bigg]\zeta_{s,\omega}^{1} \\ & \quad + \frac{3e(s-1)}{s} \Bigg[ \frac{1}{n-1} + \frac{n-1}{\left(n-1-e(s-1)\right)^{2}} \Bigg] \left( \zeta_{s}^{s} - s\zeta_{s,\omega}^{1} \right) \\ & \quad + 3\binom{n-1}{s-1}^{-1}\frac{1-p}{p}\zeta_{s}^{s}. \end{align}\] The last term is the only new Bernoulli-sampling contribution, and \[\binom{n-1}{s-1}^{-1}\frac{1-p}{p} = \binom{n-1}{s-1}^{-1} \left( \frac{\binom{n}{s}}{N} - 1 \right) = \frac{n}{Ns} - \binom{n-1}{s-1}^{-1}.\] Therefore, \[\frac{1}{\zeta_{s,\omega}^{1}} \binom{n-1}{s-1}^{-1}\frac{1-p}{p}\zeta_{s}^{s} \longrightarrow 0\] by 3 and the boundedness of \(\zeta_{s}^{s}\). The remaining terms satisfy the corresponding complete-case bounds, so \[\label{eq:overline95T95negligible} \frac{\Ex*{\overline{T}_{\ell}^{2}}}{\zeta_{s,\omega}^{1}} \longrightarrow 0.\tag{12}\] Thus, after subtracting off the linear Hájek piece, the remainder is negligible on the \(\zeta_{s,\omega}^{1}\) scale. The additional Bernoulli-sampling contribution is absorbed by 3.

Define \[\overline{X}_{\ell} \mathrel{\vcenter{:}}= \sqrt{ \frac{n-d}{n\binom{n}{d}\zeta_{s,\omega}^{1}} }\,A_{\ell}, \qquad \overline{Y}_{\ell} \mathrel{\vcenter{:}}= \sqrt{ \frac{n-d}{n\binom{n}{d}\zeta_{s,\omega}^{1}} }\,\overline{T}_{\ell}.\] Then \[\sum_{\ell \in L_{n,d}} \left( \overline{X}_{\ell} + \overline{Y}_{\ell} \right)^{2} = \frac{n\widehat{\overline{\sigma}}_{JKD}^{2}\left(D_{[n]}; d, \rho, \omega\right)}{s^{2}\zeta_{s,\omega}^{1}},\] where \[\widehat{\overline{\sigma}}_{JKD}^{2}\left(D_{[n]}; d, \rho, \omega\right) \mathrel{\vcenter{:}}= \frac{n-d}{d}\binom{n}{d}^{-1} \sum_{\ell \in L_{n,d}} \overline{\Delta}_{\ell}^{2}.\] Moreover, 12 implies \[\sum_{\ell \in L_{n,d}} \Ex*{\overline{Y}_{\ell}^{2}} = \frac{n-d}{n\zeta_{s,\omega}^{1}} \Ex*{\overline{T}_{\ell}^{2}} \longrightarrow 0.\] The corresponding statements for \(\overline{X}_{\ell}\) are exactly the same as in the complete-case proof because \(A_{\ell}\) is unchanged: \[\sum_{\ell \in L_{n,d}} \Ex*{\overline{X}_{\ell}^{2}} \longrightarrow 1, \qquad \sum_{\ell \in L_{n,d}} \overline{X}_{\ell}^{2} \xrightarrow{\mathbb{P}}1,\] where the second convergence again follows from the exact averaged-square identity and 3. Hence 2 gives \[\label{eq:scaled95ht95jkd} \frac{n\widehat{\overline{\sigma}}_{JKD}^{2}\left(D_{[n]}; d, \rho, \omega\right)}{s^{2}\zeta_{s,\omega}^{1}} \xrightarrow{\mathbb{P}}1.\tag{13}\] The limit in 13 is the Horvitz-Thompson target that remains to transfer to realized-count normalization. At this point the delete-\(d\) jackknife has already been shown to work for the Horvitz-Thompson normalized statistic. What remains is no longer a Hoeffding-decomposition problem. It is purely a question of whether replacing the deterministic counts \(N\) and \(N_{d}^{\circ}\) by their realized versions \(\widehat N\) and \(\widehat N_{\ell}^{\circ}\) can change the jackknife differences at first order.

It remains to replace the Horvitz-Thompson normalization by the empirical counts. For each delete set \(\ell\), define \[\widehat N_{\ell}^{\circ} \mathrel{\vcenter{:}}= \sum_{\iota \in \mathcal{O}_{s,0}(\ell)} \rho_{\iota}, \qquad N_{d}^{\circ} \mathrel{\vcenter{:}}= \Ex*{\widehat N_{\ell}^{\circ}} = p\binom{n-d}{s} = N\binom{n-d}{s}\binom{n}{s}^{-1}.\] Then \[U_{n,s,N,\omega}\left(D_{[n]}\right) - \overline{U}_{n,s,\rho,\omega}\left(D_{[n]}\right) = \left( \widehat N^{-1} - N^{-1} \right) \sum_{\iota \in L_{n,s}} \rho_{\iota} h_{s}\left(D_{\iota}; \omega\right),\] and \[U_{n,s,N,\omega}\left(D_{[n],-\ell}\right) - \overline{U}_{n,s,\rho,\omega}\left(D_{[n],-\ell}\right) = \left( \left(\widehat N_{\ell}^{\circ}\right)^{-1} - \left(N_{d}^{\circ}\right)^{-1} \right) \sum_{\iota \in \mathcal{O}_{s,0}(\ell)} \rho_{\iota} h_{s}\left(D_{\iota}; \omega\right).\] The actual incomplete statistic and its Horvitz-Thompson analogue differ only through the reciprocal-count factors. It remains to show that count fluctuations are negligible after jackknife rescaling. We start with the full-sample normalization error. Write \[S \mathrel{\vcenter{:}}= \sum_{\iota \in L_{n,s}} \rho_{\iota} h_{s}\left(D_{\iota}; \omega\right).\] By Cauchy-Schwarz, \[S^{2} \leq \widehat N \sum_{\iota \in L_{n,s}} \rho_{\iota} h_{s}\left(D_{\iota}; \omega\right)^{2}.\] Therefore, \[\begin{align} \Ex*{ \left( U_{n,s,N,\omega}\left(D_{[n]}\right) - \overline{U}_{n,s,\rho,\omega}\left(D_{[n]}\right) \right)^{2} } & = \Ex*{ \left( \frac{N-\widehat N}{N\widehat N} \right)^{2} S^{2} } \\ \qquad\leq & \frac{1}{N^{2}} \Ex*{ \frac{(N-\widehat N)^{2}}{\widehat N} \sum_{\iota \in L_{n,s}} \rho_{\iota} h_{s}\left(D_{\iota}; \omega\right)^{2} }. \end{align}\] By exchangeability and independence of \(\rho_{\iota}\) from the kernel, \[\begin{align} & \Ex*{ \left( U_{n,s,N,\omega}\left(D_{[n]}\right) - \overline{U}_{n,s,\rho,\omega}\left(D_{[n]}\right) \right)^{2} } \\ & \quad \leq \frac{\binom{n}{s}\zeta_{s}^{s}}{N^{2}} \Ex*{ \frac{(N-\widehat N)^{2}}{\widehat N} \rho_{[s]} }. \end{align}\] Conditioning on \(\rho_{[s]}=1\), we have \(\widehat N = 1 + B\) with \(B \sim \mathrm{Binomial}\left(\binom{n}{s}-1,p\right)\), so \[\begin{align} & \Ex*{ \frac{(N-\widehat N)^{2}}{\widehat N} \;\middle|\; \rho_{[s]} = 1 } \\ & \quad = \Ex*{\widehat N \;\middle|\;\rho_{[s]} = 1} - 2N + N^{2}\Ex*{\widehat N^{-1} \;\middle|\;\rho_{[s]} = 1}. \end{align}\] The binomial identity \[\Ex*{\frac{1}{1+B}} = \frac{1-(1-p)^{\binom{n}{s}}}{\binom{n}{s}p} \leq \frac{1}{N}\] gives \[\Ex*{ \frac{(N-\widehat N)^{2}}{\widehat N} \;\middle|\; \rho_{[s]} = 1 } \leq 1-p.\] Consequently, \[\label{eq:full95norm95error95bound} \Ex*{ \left( U_{n,s,N,\omega}\left(D_{[n]}\right) - \overline{U}_{n,s,\rho,\omega}\left(D_{[n]}\right) \right)^{2} } \leq \frac{\zeta_{s}^{s}}{N}.\tag{14}\]

The deleted-sample normalization error is handled in exactly the same way. For fixed \(\ell\), write \[S_{0}^{(\ell)} \mathrel{\vcenter{:}}= \sum_{\iota \in \mathcal{O}_{s,0}(\ell)} \rho_{\iota} h_{s}\left(D_{\iota}; \omega\right).\] Repeating the preceding argument with \(\widehat N_{\ell}^{\circ}\) in place of \(\widehat N\) and \(N_{d}^{\circ}\) in place of \(N\) yields \[\label{eq:deleted95norm95error95bound} \Ex*{ \left( U_{n,s,N,\omega}\left(D_{[n],-\ell}\right) - \overline{U}_{n,s,\rho,\omega}\left(D_{[n],-\ell}\right) \right)^{2} } \leq \frac{\zeta_{s}^{s}}{N_{d}^{\circ}}.\tag{15}\] Moreover, \[\frac{N_{d}^{\circ}}{N} = \frac{\binom{n-d}{s}}{\binom{n}{s}} = \prod_{m=0}^{s-1} \frac{n-d-m}{n-m} \longrightarrow 1\] because \(sd = o(n)\). Thus the expected retained count changes only by a \(1+o(1)\) factor, and 15 gives the same normalization bound for deleted samples.

For the normalization transfer, the relevant cancellation comes from separating the avoided and hit subsamples. For fixed \(\ell\), define \[\mathcal{O}_{s,1}(\ell) \mathrel{\vcenter{:}}= \set*{\iota \in L_{n,s}: \iota \cap \ell \neq \emptyset},\] and write \[H_{1} \mathrel{\vcenter{:}}= \sum_{\iota \in \mathcal{O}_{s,1}(\ell)} \rho_{\iota}, \qquad S_{1} \mathrel{\vcenter{:}}= \sum_{\iota \in \mathcal{O}_{s,1}(\ell)} \rho_{\iota} h_{s}\left(D_{\iota}; \omega\right), \qquad N_{1} \mathrel{\vcenter{:}}= \Ex*{H_{1}} = N - N_{d}^{\circ}.\] Thus \(\widehat N = \widehat N_{\ell}^{\circ} + H_{1}\) and \[\label{eq:hit95fraction95bound} \frac{N_{1}}{N} = 1 - \frac{\binom{n-d}{s}}{\binom{n}{s}} = 1 - \prod_{m=0}^{s-1} \left( 1 - \frac{d}{n-m} \right) \lesssim \frac{sd}{n},\tag{16}\] where the last step uses \(sd=o(n)\) and the elementary bound \(1-\prod_{m}(1-a_m)\leq \sum_m a_m\) for \(a_m \in [0,1]\). In particular, 16 gives \(N_{1}=o(N)\), while \(N_{d}^{\circ}\asymp N\).

Now let \[\Delta_{\ell} \mathrel{\vcenter{:}}= \Delta_{\ell}[U_{n,s,N,\omega}] = U_{n,s,N,\omega}\left(D_{[n]}\right) - U_{n,s,N,\omega}\left(D_{[n],-\ell}\right), \qquad R_{\ell} \mathrel{\vcenter{:}}= \Delta_{\ell} - \overline{\Delta}_{\ell}.\] Using \(S=S_{0}^{(\ell)}+S_{1}\) and \(\widehat N=\widehat N_{\ell}^{\circ}+H_{1}\), \[R_{\ell} = \left( \frac{1}{\widehat N} - \frac{1}{N} \right)S_{1} + \left( \frac{N_{1}}{N N_{d}^{\circ}} - \frac{H_{1}}{\widehat N \widehat N_{\ell}^{\circ}} \right)S_{0}^{(\ell)}.\] We bound the two terms on the right separately. First, use \[S_{1}^{2} \leq H_{1}\sum_{\iota \in \mathcal{O}_{s,1}(\ell)} \rho_{\iota}h_{s}(D_{\iota};\omega)^{2} \leq \widehat N\sum_{\iota \in \mathcal{O}_{s,1}(\ell)} \rho_{\iota}h_{s}(D_{\iota};\omega)^{2}.\] Then \[\left( \left( \frac{1}{\widehat N} - \frac{1}{N} \right)S_{1} \right)^{2} \leq \frac{(N-\widehat N)^{2}}{N^{2}\widehat N} \sum_{\iota \in \mathcal{O}_{s,1}(\ell)} \rho_{\iota}h_{s}(D_{\iota};\omega)^{2}.\] By exchangeability within \(\mathcal{O}_{s,1}(\ell)\), conditioning on any fixed hit subset being selected gives the same reciprocal-count factor as in the proof of 14 . Hence \[\Ex*{ \left( \left( \frac{1}{\widehat N} - \frac{1}{N} \right)S_{1} \right)^{2} } \lesssim \frac{N_{1}}{N^{2}}\zeta_{s}^{s} \lesssim \frac{sd}{n}\frac{\zeta_{s}^{s}}{N}.\] For the second term, on the event \(\widehat N_{\ell}^{\circ}>0\), write \[g_{k}(u) \mathrel{\vcenter{:}}= \frac{u}{k(k+u)}, \qquad k,u>0,\] so that \[\frac{N_{1}}{N N_{d}^{\circ}} - \frac{H_{1}}{\widehat N \widehat N_{\ell}^{\circ}} = g_{N_{d}^{\circ}}(N_{1}) - g_{\widehat N_{\ell}^{\circ}}(H_{1}).\] On the complementary zero-count event, \(S_{0}^{(\ell)}=0\) by convention, so the same product is identically zero and there is nothing to bound. Let \[G_\ell \mathrel{\vcenter{:}}= \set*{ \frac{1}{2} N_d^\circ \leq \widehat N_\ell^\circ \leq 2N_d^\circ }.\] On \(G_\ell\), \(N_d^\circ\asymp\widehat N_\ell^\circ\). Since \(N_1=o(N_d^\circ)\), the derivatives \[\partial_u g_k(u)=(k+u)^{-2}, \qquad \partial_k g_k(u)=-\frac{u(2k+u)}{k^2(k+u)^2}\] are uniformly controlled along the two line segments joining \((N_d^\circ,N_1)\), \((\widehat N_\ell^\circ,N_1)\), and \((\widehat N_\ell^\circ,H_1)\). Therefore \[\abs*{ g_{N_d^\circ}(N_1) - g_{\widehat N_\ell^\circ}(H_1) } \1*{G_\ell} \lesssim \left( \frac{\abs*{H_1-N_1}}{(N_d^\circ)^2} + \frac{N_1\abs*{\widehat N_\ell^\circ-N_d^\circ}}{(N_d^\circ)^3} \right) \1*{G_\ell}.\] Using \(\left(S_{0}^{(\ell)}\right)^2\leq\widehat N_\ell^\circ\sum_{\iota\in\mathcal{O}_{s,0}(\ell)}\rho_\iota h_s(D_\iota;\omega)^2\) and \(\widehat N_\ell^\circ\leq2N_d^\circ\) on \(G_\ell\), we obtain \[\begin{align} & \Ex*{ \left[ \left( \frac{N_{1}}{N N_{d}^{\circ}} - \frac{H_{1}}{\widehat N \widehat N_{\ell}^{\circ}} \right)S_{0}^{(\ell)} \right]^2 \1*{G_\ell} } \\ & \quad \lesssim \frac{1}{(N_d^\circ)^3} \Ex*{ (H_1-N_1)^2 \sum_{\iota\in\mathcal{O}_{s,0}(\ell)} \rho_\iota h_s(D_\iota;\omega)^2 } \\ & \qquad + \frac{N_1^2}{(N_d^\circ)^5} \Ex*{ (\widehat N_\ell^\circ-N_d^\circ)^2 \sum_{\iota\in\mathcal{O}_{s,0}(\ell)} \rho_\iota h_s(D_\iota;\omega)^2 }. \end{align}\] Now \(H_1\) depends only on the Bernoulli variables indexed by \(\mathcal{O}_{s,1}(\ell)\), while \(\widehat N_\ell^\circ\) and \(S_0^{(\ell)}\) depend only on those indexed by \(\mathcal{O}_{s,0}(\ell)\), so the hit count is independent of the avoid-layer quantities. Also, \[\Ex*{ \sum_{\iota\in\mathcal{O}_{s,0}(\ell)} \rho_\iota h_s(D_\iota;\omega)^2 } = N_d^\circ\zeta_s^s\] because the centered kernel has second moment \(\zeta_s^s\). For the second expectation, exchangeability within the avoid layer and conditioning on a fixed avoid subset being selected gives \(\widehat N_\ell^\circ=1+B_0\), with \(B_0\sim\mathrm{Binomial}(\abs*{\mathcal{O}_{s,0}(\ell)}-1,p)\); hence \[\Ex*{ (\widehat N_\ell^\circ-N_d^\circ)^2 \sum_{\iota\in\mathcal{O}_{s,0}(\ell)} \rho_\iota h_s(D_\iota;\omega)^2 } \lesssim (N_d^\circ)^2\zeta_s^s.\] Since \(\Varb*{H_1}=N_1(1-p)\leq N_1\), this gives \[\Ex*{ \left[ \left( \frac{N_{1}}{N N_{d}^{\circ}} - \frac{H_{1}}{\widehat N \widehat N_{\ell}^{\circ}} \right)S_{0}^{(\ell)} \right]^2 \1*{G_\ell} } \lesssim \frac{N_1}{(N_d^\circ)^2}\zeta_s^s + \frac{N_1^2}{(N_d^\circ)^3}\zeta_s^s.\] On \(G_\ell^c\), use the crude coefficient bound \[\abs*{ \frac{N_1}{NN_d^\circ} - \frac{H_1}{\widehat N\widehat N_\ell^\circ} } \leq \frac{N_1}{NN_d^\circ} + \frac{1}{\widehat N_\ell^\circ}\] on the positive-count event. Together with \(S_0^{(\ell)2}\leq \widehat N_\ell^\circ\sum_{\mathcal{O}_{s,0}(\ell)}\rho_\iota h_s(D_\iota;\omega)^2\), \(\ref{lem:reciprocal95binomial95good95count}\), and the same exchangeability argument, this yields, after reducing the exponential constant if necessary, \[\Ex*{ \left[ \left( \frac{N_1}{NN_d^\circ} - \frac{H_1}{\widehat N\widehat N_\ell^\circ} \right)S_0^{(\ell)} \right]^2 \1*{G_\ell^c} } \lesssim \zeta_s^s\exp(-cN_d^\circ).\] Combining the good-count and complement bounds, and using \(N_d^\circ\asymp N\), we conclude \[\Ex*{ \left[ \left( \frac{N_{1}}{N N_{d}^{\circ}} - \frac{H_{1}}{\widehat N \widehat N_{\ell}^{\circ}} \right)S_{0}^{(\ell)} \right]^{2} } \lesssim \frac{N_{1}}{\left(N_{d}^{\circ}\right)^{2}}\zeta_{s}^{s} + \frac{N_{1}^{2}}{\left(N_{d}^{\circ}\right)^{3}}\zeta_{s}^{s} + \zeta_s^s\exp(-cN_d^\circ)\] and therefore, by 16 , \[\Ex*{ \left[ \left( \frac{N_{1}}{N N_{d}^{\circ}} - \frac{H_{1}}{\widehat N \widehat N_{\ell}^{\circ}} \right)S_{0}^{(\ell)} \right]^{2} } \lesssim \frac{sd}{n}\frac{\zeta_{s}^{s}}{N} + \zeta_s^s\exp(-cN_d^\circ).\] Combining the last two displays and \(\abs*{a+b}^{2}\leq 2(a^{2}+b^{2})\), we obtain the sharpened coupled bound \[\label{eq:coupled95R95bound} \Ex*{R_{\ell}^{2}} \lesssim \frac{sd}{n}\frac{\zeta_{s}^{s}}{N} + \zeta_s^s\exp(-cN_d^\circ).\tag{17}\] Define \[\overline{Z}_{\ell} \mathrel{\vcenter{:}}= \sqrt{ \frac{n(n-d)}{s^{2}d\binom{n}{d}\zeta_{s,\omega}^{1}} }\,R_{\ell}.\] Using 17 , exchangeability gives \[\sum_{\ell \in L_{n,d}} \Ex*{\overline{Z}_{\ell}^{2}} = \frac{n(n-d)}{s^{2}d\zeta_{s,\omega}^{1}} \Ex*{R_{\ell}^{2}} \lesssim \frac{n-d}{N s\zeta_{s,\omega}^{1}} \zeta_{s}^{s} + \frac{n(n-d)}{s^2d\zeta_{s,\omega}^{1}} \zeta_s^s\exp(-cN_d^\circ) \longrightarrow 0\] by 3, the boundedness of \(\zeta_{s}^{s}\), \((n-d)/n \leq 1\), and the exponential domination used in 5.

Since \[\frac{n\widehat{\sigma}_{JKD}^{2}\left(D_{[n]}; d, \omega\right)}{s^{2}\zeta_{s,\omega}^{1}} = \sum_{\ell \in L_{n,d}} \left( \overline{X}_{\ell} + \overline{Y}_{\ell} + \overline{Z}_{\ell} \right)^{2},\] while \(\sum \Ex*{\overline{X}_{\ell}^{2}} \to 1\), \(\sum \overline{X}_{\ell}^{2} \xrightarrow{\mathbb{P}}1\), and \[\sum_{\ell \in L_{n,d}} \Ex*{ \left( \overline{Y}_{\ell} + \overline{Z}_{\ell} \right)^{2} } \leq 2 \sum_{\ell \in L_{n,d}} \Ex*{\overline{Y}_{\ell}^{2}} + 2 \sum_{\ell \in L_{n,d}} \Ex*{\overline{Z}_{\ell}^{2}} \longrightarrow 0,\] 2 yields \[\label{eq:scaled95incomplete95jkd} \frac{n\widehat{\sigma}_{JKD}^{2}\left(D_{[n]}; d, \omega\right)}{s^{2}\zeta_{s,\omega}^{1}} \xrightarrow{\mathbb{P}}1.\tag{18}\] Thus realized counts add only a second-order perturbation to the Horvitz-Thompson jackknife argument.

Finally, 14 implies \[\Varb*{ U_{n,s,N,\omega}\left(D_{[n]}\right) - \overline{U}_{n,s,\rho,\omega}\left(D_{[n]}\right) } \leq \frac{\zeta_{s}^{s}}{N} = o\left( \frac{s^{2}}{n}\zeta_{s,\omega}^{1} \right),\] again by 3 and bounded \(\zeta_{s}^{s}\). Hence \[\abs*{ \sigma_{n} - \overline{\sigma}_{n} } \leq \sqrt{ \Varb*{ U_{n,s,N,\omega}\left(D_{[n]}\right) - \overline{U}_{n,s,\rho,\omega}\left(D_{[n]}\right) } } = o\left( \sqrt{ \frac{s^{2}}{n}\zeta_{s,\omega}^{1} } \right),\] so 11 implies \[\frac{n\sigma_{n}^{2}}{s^{2}\zeta_{s,\omega}^{1}} \longrightarrow 1.\] The same comparison that transfers the jackknife estimator from \(\overline{U}_{n,s,\rho,\omega}\) to \(U_{n,s,N,\omega}\) also transfers the variance target itself. So both the estimator and the target are asymptotically governed by the same first-projection scale. Combining this with 18 gives, for centered kernels, \[\frac{\widehat{\sigma}_{JKD}^{2}\left(D_{[n]}; d, \omega\right)}{\sigma_{n}^{2}} \xrightarrow{\mathbb{P}}1.\] For a general kernel with bounded \(\theta_s\), the lower Hoeffding projections and \(\zeta_{s,\omega}^1\) are unchanged after replacing \(h_s\) by \(h_s-\theta_s\). Moreover, 3 and the boundedness of \(\zeta_s^s\) imply \(N/n\to\infty\) and \(1/\zeta_{s,\omega}^1=o(Ns/n)\). Since \(N_d^\circ/N\to1\) by \(sd=o(n)\), 5 transfers both the jackknife estimator and the variance target from the centered kernel back to the original kernel. This proves the stated ratio consistency under the zero-count convention. ◻

7 DNN Kernel Tools and Single-Scale Hájek Dominance↩︎

7.1 Notation↩︎

We first isolate the single-scale DNN notation that will be used throughout 7.

For the DNN kernel, define the first-order projection objects by \[\psi_{s}^{1}(x; d_{1}) = \Ex*{h_{s}\left(x; D_{[s]}\right) \;\middle|\;Z_{1} = d_{1}}\] Let \[\theta_{s}(x) \coloneq \Ex*{h_{s}\left(x; D_{[s]}\right)}\] denote the finite-sample DNN mean. Then \[h_{s}^{(1)}\left(x; d_{1}\right) = \psi_{s}^{1}(x; d_{1}) - \theta_{s}(x).\] Since \(\theta_{s}(x)\) and \(\mu(x)\) are both constants in \(d_{1}\), \[\Varb*{h_{s}^{(1)}(x; Z_{1})} = \Varb*{\psi_{s}^{1}(x; Z_{1}) - c}\] is the same under either centering. The two conventions differ only by \(\theta_{s}(x) - \mu(x) \to 0\) as \(s, n \to \infty\); see 7. For higher-order projection kernels, write for \(c = 2, \dotsc, s\), \[\psi_{s}^{c}(x; d_{[c]}) = \Ex*{h_{s}\left(x; D_{[s]}\right) \;\middle|\;D_{[c]} = d_{[c]}},\] and recursively define \[h_{s}^{(c)}\left(x; d_{[c]}\right) = \psi_{s}^{c}(x; d_{[c]}) - \sum_{j = 1}^{c-1}\left(\sum_{\ell \in L_{c,j}}h_{s}^{(j)}(x; d_{\ell})\right) - \theta_{s}(x).\] The corresponding Hoeffding decomposition is \[\tilde{\mu}_{s}\left(x; D_{[n]}\right) = \theta_{s}(x) + \sum_{j = 1}^{s}\binom{s}{j}\binom{n}{j}^{-1} \sum_{\ell \in L_{n,j}} h^{(j)}_{s}(x; D_{\ell}).\] We also write, for any \(1 \leq c \leq s\), \[\begin{align} \zeta_{s}^{1}\left(x\right) & = \Varb*{h_{s}^{(1)}\left(x; Z_{1}\right)} \\ \zeta_{s}^{s}\left(x\right) & = \Varb*{h_{s}\left(x; D_{[s]}\right)} \\ \xi_{s}^{c}\left(x\right) & = \Varb*{\psi_{s}^{c}(x; D_{[c]})} \\ \Omega_{s}\left(x\right) & = \Ex*{h_{s}^{2}\left(x; D_{[s]}\right)} \\ \Omega_{s}^{c}\left(x\right) & = \mathbb{E}\left[h_{s}\left(x; D_{[s]}\right) \cdot h_{s}\left(x; D_{[s]}^{\prime}\right)\right] \end{align}\] where \(D_{[s]} = \{Z_1, \dotsc, Z_{s}\}\) is a vector of i.i.d.random variables drawn from \(P\) and \(D_{[s]}^{\prime} = \{Z_1, \dotsc, Z_{c}, Z_{c+1}^{\prime}, \dotsc, Z_{s}^{\prime}\}\) where \(Z_{c+1}^{\prime}, \dotsc, Z_{s}^{\prime}\) are i.i.d.draws from \(P\) that are independent of \(D_{[s]}\).

7.2 DNN Kernel Lemmas↩︎

Throughout the proofs in this section, we are concerned with the same setup. For the sake of brevity, we introduce this setup here to avoid unnecessary repetition in the following lemmas. Consider a sample size \(n\), a subsampling scale \(s\) growing with \(n\), and \(c\) such that \(0 < c \leq s \leq n\). Let \(D_{[s]} = \left\{Z_1, Z_2, \dotsc, Z_c, Z_{c+1}, \dotsc Z_s \right\}\) be an i.i.d.data set drawn from \(P\) as described in 4. Let \(D^{\prime}_{[s]}= \left\{Z_1, Z_2, \dotsc, Z_c, Z_{c+1}^{\prime}, \dotsc Z_s^{\prime} \right\}\) be a second data set that shares the first \(c\) observations with \(D_{[s]}\). The remaining \(s - c\) observations of \(D^{\prime}_{[s]}\), i.e.\(\left\{Z_{c+1}^{\prime}, \dotsc Z_s^{\prime} \right\}\), are i.i.d.draws from \(P\) that are independent of \(D_{[s]}\). All expectations are with respect to all random elements unless a conditioning bar is displayed; we write \(\Ex*{\cdot \;\middle|\;X}\) for conditional expectation given the sigma-field generated by the displayed variables.

Lemma 6 ([4] - Lemma 12).
* The indicator functions \(\kappa\left(x; Z_{i}, D_{[s]}\right)\) satisfy the following properties.

  1. For any \(i \neq j\), we have \(\kappa\left(x; Z_{i}, D_{[s]}\right) \kappa\left(x; Z_{j}, D_{[s]}\right)=0\) with probability one;

  2. \(\sum_{i=1}^{s} \kappa\left(x; Z_{i}, D_{[s]}\right)=1\);

  3. \(\forall i \in [s]: \quad \Ex*{\kappa\left(x; Z_{i}, D_{[s]}\right)}=s^{-1}\)

  4. \(\Ex*{\kappa\left(x; Z_1, D_{[s]}\right) \;\middle|\;D_{1} = Z_{1}} = \left\{1-\varphi\left(B\left(x,\norm*{X_1-x}\right)\right)\right\}^{s-1}\)

Lemma 7 ([4] - Lemma 13).
* For any \(L^1\) function \(f\) that is continuous at \(x\), it holds that \[\lim _{s \longrightarrow \infty} \Ex*{f\left(X_1\right) s \Ex*{\kappa\left(x; Z_1, D_{[s]}\right) \;\middle|\;X_{1}}} = f(x).\]

We also need product analogues of 7 for the covariance calculations below.

Lemma 8.
* The following three statements hold. \[\forall i \in [c]: \quad \Ex*{\kappa\left(x; Z_{i}, D_{[s]}\right)\kappa\left(x; Z_{i}, D^{\prime}_{[s]}\right)} = (2s - c)^{-1} = \omega(e^{-s})\] \[\begin{align} & \forall i \in [c] \; \forall j \in \{c+1, \dotsc, s\}: \quad \Ex*{\kappa\left(x; Z_{i}, D_{[s]}\right)\kappa\left(x; Z_{j}^{\prime}, D^{\prime}_{[s]}\right)} \\ & \qquad = \frac{1}{s(2s-c)} = \omega(e^{-s}) \end{align}\] \[\begin{align} & \forall i,j \in \{c+1, \dotsc, s\}: \quad \Ex*{\kappa\left(x; Z_{i}, D_{[s]}\right)\kappa\left(x; Z_{j}^{\prime}, D^{\prime}_{[s]}\right)} \\ & \qquad = \frac{2}{s(2s-c)} = \omega(e^{-s}) \end{align}\]

Proof of 8.
* By symmetry, it suffices to consider \(i=1\) and \(j=c+1\) for the first two equations. \[\begin{align} & \Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{1}, D^{\prime}_{[s]}\right)}\\ & \quad = \Ex*{\kappa\left(x; Z_{1}, D_{[c]}\right)\kappa\left(x; Z_{1}, D_{(c+1):s}\right)\kappa\left(x; Z_{1}, D^{\prime}_{(c+1):s}\right)} \\ & \quad = \Ex*{\kappa\left(x; Z_{1}, D_{[2s - c]}\right)} = (2s - c)^{-1} \end{align}\] For the remaining two cases we use the reciprocal-binomial identity \[\label{eq:reciprocal95binomial95identity} \sum_{r = 0}^{a}\binom{a}{r}\binom{N}{r}^{-1} = \frac{N+1}{N+1-a}, \qquad 0 \leq a \leq N.\tag{19}\] This follows by writing \(\binom{a}{r}\binom{N}{r}^{-1} = \binom{N-r}{a-r}\binom{N}{a}^{-1}\) and applying the hockey-stick identity. Considering the second case, we find the following. \[\begin{align} & \Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right)} \\ & = \frac{1}{(2s - c)!} \sum_{r = 0}^{s - c - 1}\binom{s - c - 1}{r} r! \left((s - 1) + (s - c - 1 - r)\right)! \\ & = \frac{1}{(2s - c)!}\sum_{r = 0}^{s - c - 1}\binom{s - c - 1}{r} r! \left(2s - c - 2 - r\right)! \\ & = \frac{1}{(2s - c)(2s - c - 1)} \sum_{r = 0}^{s - c - 1}\binom{s - c - 1}{r}\binom{2s - c - 2}{r}^{-1} \\ & \overset{\text{\ref{eq:reciprocal95binomial95identity}}}{=} \frac{1}{s(2s-c)}. \end{align}\] While unintuitive at first, the terms in this expression have intuitive meaning when we consider this as a combinatorial problem. Consider lining up the observations in order of their distance to the point of interest and counting the cases for which the expression in the expectation is equal to one. First, there are \((2s-c)!\) possible orderings of the observations with probability one, leading to the denominator. Next, notice that only those orderings where \(\norm*{X_{c+1}^{\prime} - x} \leq \norm*{X_{1}-x}\) and \(\norm*{X_{1} - x} \leq \norm*{X_{i} - x}\) for any \(i = 2, \dotsc, c\) can possibly lead to a non-zero realization of the kernel term. Furthermore, out of the \((s-c-1)\) observations in \(D^{\prime}_{(c+2):s}\), it is possible for \(i = 0, \dotsc, s-c-1\) observations to lie at a distance to the point of interest that is smaller than \(\norm*{X_{1}-x}\) but larger than \(\norm*{X_{c+1}^{\prime} - x}\) in any permutation. The sum adjusts for those possible configurations. Because \(c \leq s-1\), the exact value is bounded from below by \((2s^2)^{-1}\) and is therefore \(\omega(e^{-s})\).

Considering the third case, without loss of generality, we consider the case of \(i = j = c+1\). We find the following. \[\begin{align} & \Ex*{\kappa\left(x; Z_{c+1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right)} \\ & = \mathbb{E}\left[ \kappa\left(x; Z_{c+1}, D_{[c]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D_{[c]}\right) \kappa\left(x; Z_{c+1}, D_{(c+1):s}\right)\kappa\left(x; Z_{c+1}^{\prime}, D_{(c+1):s}^{\prime}\right) \right] \\ & = \frac{2}{(2s - c)!} \sum_{r = 0}^{s-c-1} \binom{s - c - 1}{r}(s - 1 + r)!(s-c-1-r)! \\ & = \frac{2(2s - c - 2)!}{(2s-c)!} \sum_{r = 0}^{s-c-1} \binom{s - c - 1}{r}\binom{2s - c - 2}{s-1+r}^{-1} \\ & = \frac{2}{(2s-c)(2s-c-1)}\sum_{r = 0}^{s-c-1} \binom{s - c - 1}{r}\binom{2s - c - 2}{s-1+r}^{-1} \\ & = \frac{2}{(2s-c)(2s-c-1)}\sum_{r = 0}^{s-c-1} \binom{s - c - 1}{r}\binom{2s - c - 2}{s-c-1-r}^{-1} \\ & \overset{\text{\ref{eq:reciprocal95binomial95identity}}}{=} \frac{2}{s(2s-c)}. \end{align}\] The third case follows from a similar combinatorial logic as the second. By symmetry, take \(\norm*{X_{c+1}^{\prime} - x} \leq \norm*{X_{c+1}-x}\) and multiply by two. Any number \(i = 0, \dotsc, s-c-1\) of observations in \(D^{\prime}_{(c+2):s}\) can be farther from \(x\) than \(X_{c+1}\) or lie between \(\norm*{X_{c+1}^{\prime} - x}\) and \(\norm*{X_{c+1}-x}\). The summation counts these configurations. The exact value is again bounded from below by \((s^2)^{-1}\), and hence it is \(\omega(e^{-s})\). ◻

Lemma 9.
* The following statements hold. \[\begin{align} & \forall i \in [c]: \quad \Ex*{\kappa\left(x; Z_{i}, D_{[s]}\right)\kappa\left(x; Z_{i}, D^{\prime}_{[s]}\right) \;\middle|\;X_i} = \left\{1-\varphi\left(B\left(x,\norm*{X_{i}-x}\right)\right)\right\}^{2s-c-1} \end{align}\] \[\begin{align} & \forall i \in [c] \; \forall j \in \{c+1, \dotsc, s\}: \\ & \Ex*{\kappa\left(x; Z_{i}, D_{[s]}\right)\kappa\left(x; Z_{j}^{\prime}, D^{\prime}_{[s]}\right) \;\middle|\;X_i, X_{j}^{\prime}} \\ & \qquad = \1*{\norm*{X_{j}^{\prime} - x} \leq \norm*{X_{i} - x}} \left\{1-\varphi\left(B\left(x,\norm*{X_{i}-x}\right)\right)\right\}^{s-1} \left\{1-\varphi\left(B\left(x,\norm*{X_{j}^{\prime}-x}\right)\right)\right\}^{s-c-1} \end{align}\] \[\begin{align} & \forall i,j \in \{c+1, \dotsc, s\}: \\ & \Ex*{\kappa\left(x; Z_{i}, D_{[s]}\right)\kappa\left(x; Z_{j}^{\prime}, D^{\prime}_{[s]}\right) \;\middle|\;X_{i}, X_{j}^{\prime}} \\ & \qquad = \left\{1-\varphi\left(B\left(x, \min\left(\norm*{X_{i} - x}, \norm*{X_{j}^{\prime}-x}\right)\right)\right)\right\}^{s-c-1} \\ & \qquad\quad \times \left\{1-\varphi\left(B\left(x,\max\left(\norm*{X_{i} - x}, \norm*{X_{j}^{\prime}-x}\right)\right)\right)\right\}^{s-1} \end{align}\]

Proof of 9.
* By symmetry, it suffices to take \(i=1\) for the first equation. \[\begin{align} & \Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{1}, D^{\prime}_{[s]}\right) \;\middle|\;X_1} \\ & = \mathbb{E}\left[\kappa\left(x; Z_{1}, D_{[c]}\right) \kappa\left(x; Z_{1}, D_{(c+1):s}\right) \kappa\left(x; Z_{1}, D^{\prime}_{(c+1):s}\right) \middle| X_1\right] \\ & = \Ex*{\kappa\left(x; Z_{1}, D_{[c]}\right)\;\middle|\;X_1} \Ex*{\kappa\left(x; Z_{1}, D_{(c+1):s}\right)\;\middle|\;X_1} \Ex*{\kappa\left(x; Z_{1}, D_{(c+1):s}^{\prime}\right)\;\middle|\;X_1} \\ & = \left\{1-\varphi\left(B\left(x,\norm*{X_{i}-x}\right)\right)\right\}^{c-1} \left\{1-\varphi\left(B\left(x,\norm*{X_{i}-x}\right)\right)\right\}^{s-c} \left\{1-\varphi\left(B\left(x,\norm*{X_{i}-x}\right)\right)\right\}^{s-c} \\ & = \left\{1-\varphi\left(B\left(x,\norm*{X_{i}-x}\right)\right)\right\}^{2s-c-1} \end{align}\]

By symmetry, it suffices to take \(i=1\) and \(j=c+1\) for the second equation. \[\begin{align} & \Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right) \;\middle|\;X_1, X_{c+1}^{\prime}} \\ & = \mathbb{E}\Bigg[ \mathbb{E}\Bigg[ \kappa\left(x; Z_{1}, D_{[c]}\right) \kappa\left(x; Z_{1}, D_{(c+1):s}\right) \kappa\left(x; Z_{c+1}^{\prime}, D_{[c]}\right) \kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{(c+1):s}\right) \Biggm| X_{[c]}, X_{c+1}^{\prime}\Bigg] \Biggm| X_1, X_{c+1}^{\prime} \Bigg] \\ & = \mathbb{E}\left[ \mathbb{E}\left[ \kappa\left(x; Z_{1}, D_{(c+1):s}\right) \kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{(c+1):s}\right) \middle| X_{[c]}, X_{c+1}^{\prime}\right] \kappa\left(x; Z_{1}, D_{[c]}\right) \kappa\left(x; Z_{c+1}^{\prime}, D_{[c]}\right) \middle| X_1, X_{c+1}^{\prime}\right] \\ & = \mathbb{E}\left[ \mathbb{E}\left[ \kappa\left(x; Z_{1}, D_{(c+1):s}\right) \kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{(c+1):s}\right) \middle| X_1, X_{c+1}^{\prime}\right] \kappa\left(x; Z_{1}, D_{[c]}\right) \kappa\left(x; Z_{c+1}^{\prime}, D_{[c]}\right) \middle| X_1, X_{c+1}^{\prime}\right] \\ & = \mathbb{E}\left[ \kappa\left(x; Z_{1}, D_{(c+1):s}\right) \kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{(c+1):s}\right) \middle| X_1, X_{c+1}^{\prime}\right] \mathbb{E}\left[ \kappa\left(x; Z_{1}, D_{[c]}\right) \kappa\left(x; Z_{c+1}^{\prime}, D_{[c]}\right) \middle| X_1, X_{c+1}^{\prime}\right] \\ & = \Ex*{\kappa\left(x; Z_{1}, D_{(c+1):s}\right)\;\middle|\;X_1} \mathbb{E}\left[\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{(c+1):s}\right) \middle| X_{c+1}^{\prime}\right] \1*{\norm*{X_{c+1}^{\prime} - x} \leq \norm*{X_{1} - x}} \mathbb{E}\left[ \kappa\left(x; Z_{1}, D_{[c]}\right) \middle| X_1\right] \\ & = \1*{\norm*{X_{c+1}^{\prime} - x} \leq \norm*{X_{1} - x}} \Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\;\middle|\;X_1} \mathbb{E}\left[\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{(c+1):s}\right) \middle| X_{c+1}^{\prime}\right] \\ & = \1*{\norm*{X_{c+1}^{\prime} - x} \leq \norm*{X_{1} - x}} \left\{1-\varphi\left(B\left(x,\norm*{X_1-x}\right)\right)\right\}^{s-1} \left\{1-\varphi\left(B\left(x,\norm*{X_{c+1}^{\prime}-x}\right)\right)\right\}^{s-c-1} \end{align}\]

For the third case, without loss of generality, we consider the case of \(i = j = c+1\). \[\begin{align} & \mathbb{E}\left[ \kappa\left(x; Z_{c+1}, D_{[s]}\right) \kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right) \middle| X_{c+1}, X_{c+1}^{\prime}\right] \\ & = \mathbb{E}\left[ \mathbb{E}\left[ \kappa\left(x; Z_{c+1}, D_{[s]}\right) \kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right) \middle| X_{[c]}, X_{c+1}, X_{c+1}^{\prime}\right] \middle| X_{c+1}, X_{c+1}^{\prime} \right] \\ & = \mathbb{E}\left[ \mathbb{E}\left[ \kappa\left(x; Z_{c+1}, D_{[c+1]}\right) \kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[c+1]}\right) \kappa\left(x; Z_{c+1}, D_{(c+1):s}\right) \kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{(c+1):s}\right) \middle| X_{[c]}, X_{c+1}, X_{c+1}^{\prime}\right] \middle| X_{c+1}, X_{c+1}^{\prime} \right] \\ & = \mathbb{E}\left[ \mathbb{E}\left[ \kappa\left(x; Z_{c+1}, D_{[c+1]}\right) \kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[c+1]}\right) \middle| X_{[c]}, X_{c+1}, X_{c+1}^{\prime}\right] \kappa\left(x; Z_{c+1}, D_{(c+1):s}\right) \kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{(c+1):s}\right) \middle| X_{c+1}, X_{c+1}^{\prime} \right] \\ & = \mathbb{E}\left[ \kappa\left(x; Z_{c+1}, D_{[c+1]}\right) \kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[c+1]}\right) \middle| X_{c+1}, X_{c+1}^{\prime}\right] \Ex*{\kappa\left(x; Z_{c+1}, D_{(c+1):s}\right)\;\middle|\;X_{c+1}} \Ex*{\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{(c+1):s}\right)\;\middle|\;X_{c+1}^{\prime}} \end{align}\] Without loss of generality, consider the case that \(\norm*{X_{c+1} - x} \leq \norm*{X_{c+1}^{\prime} - x}\). \[\mathbb{E}\left[ \kappa\left(x; Z_{c+1}, D_{[c+1]}\right) \kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[c+1]}\right) \middle| X_{c+1}, X_{c+1}^{\prime}\right] = \mathbb{E}\left[ \kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[c+1]}\right) \middle| X_{c+1}^{\prime}\right]\] Also, \[\Ex*{\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[c+1]}\right) \;\middle|\;X_{c+1}^{\prime}} \Ex*{\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{(c+1):s}\right)\;\middle|\;X_{c+1}^{\prime}} = \Ex*{\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right)\;\middle|\;X_{c+1}^{\prime}}\] Hence \[\begin{align} & \Ex*{\kappa\left(x; Z_{c+1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right) \;\middle|\;X_{c+1}, X_{c+1}^{\prime}} \\ & = \1*{\norm*{X_{c+1}^{\prime} - x} \leq \norm*{X_{c+1} - x}} \left\{1-\varphi\left(B\left(x,\norm*{X_{c+1} - x}\right)\right)\right\}^{s-1} \left\{1-\varphi\left(B\left(x,\norm*{X_{c+1}^{\prime}-x}\right)\right)\right\}^{s-c-1} \\ & \qquad + \1*{\norm*{X_{c+1}^{\prime} - x} > \norm*{X_{c+1} - x}} \left\{1-\varphi\left(B\left(x,\norm*{X_{c+1} - x}\right)\right)\right\}^{s-c-1} \left\{1-\varphi\left(B\left(x,\norm*{X_{c+1}^{\prime}-x}\right)\right)\right\}^{s-1} \\ & = \left\{1-\varphi\left(B\left(x, \min\left(\norm*{X_{c+1} - x}, \norm*{X_{c+1}^{\prime}-x}\right)\right)\right)\right\}^{s-c-1} \left\{1-\varphi\left(B\left(x,\max\left(\norm*{X_{c+1} - x}, \norm*{X_{c+1}^{\prime}-x}\right)\right)\right)\right\}^{s-1} \end{align}\] ◻

Lemma 10.
* For any \(L^{2}(\mathcal{X})\) function \(f\) that is continuous at \(x\), it holds that \[\lim_{s \longrightarrow \infty} \underbrace{\mathbb{E}\left[f^{2}(X_{1}) (2s - c) \Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{1}, D^{\prime}_{[s]}\right) \;\middle|\;X_{1}} \right]}_{(A)} = f^{2}(x)\] \[\lim_{s \longrightarrow \infty} \underbrace{\mathbb{E}\left[ f(X_{1}) f(X_{c+1}^{\prime}) \frac{\Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right) \;\middle|\;X_{1}, X_{c+1}^{\prime}}}{\Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right)}} \right]}_{(B)} = f^{2}(x)\] \[\lim_{s \longrightarrow \infty} \underbrace{\mathbb{E}\left[ f(X_{c+1}) f(X_{c+1}^{\prime}) \frac{\Ex*{\kappa\left(x; Z_{c+1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right) \;\middle|\;X_{c+1}, X_{c+1}^{\prime}}}{\Ex*{\kappa\left(x; Z_{c+1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right)}} \right]}_{(C)} = f^{2}(x)\]

Proof of 10.
* We will largely argue along the same lines as the original proof in [4]. Thus, consider first the following inequalities. \[\begin{align} & \abs*{(A) - f^{2}(x)} = \left|\mathbb{E}\left[f^{2}(X_{1}) (2s - c) \Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{1}, D^{\prime}_{[s]}\right) \;\middle|\;X_{1}} \right] - f^{2}(x)\right| \\ & \leq \mathbb{E}\left[\abs*{f^{2}(X_{1}) - f^{2}(x)} (2s - c) \Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{1}, D^{\prime}_{[s]}\right) \;\middle|\;X_{1}} \right] \end{align}\] \[\begin{align} & \abs*{(B) - f^{2}(x)} \\ & = \left|\mathbb{E}\left[ f(X_{1}) f(X_{c+1}^{\prime}) \frac{\Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right) \;\middle|\;X_{1}, X_{c+1}^{\prime}}}{\Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right)}} \right] - f^{2}(x)\right| \\ & \leq \mathbb{E}\left[ \abs*{f(X_{1}) f(X_{c+1}^{\prime}) - f^{2}(x)} \frac{\Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right) \;\middle|\;X_{1}, X_{c+1}^{\prime}}}{\Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right)}} \right] \end{align}\] \[\begin{align} & \abs*{(C) - f^{2}(x)} \\ & = \left|\mathbb{E}\left[ f(X_{c+1}) f(X_{c+1}^{\prime}) \frac{\Ex*{\kappa\left(x; Z_{c+1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right) \;\middle|\;X_{c+1}, X_{c+1}^{\prime}}}{\Ex*{\kappa\left(x; Z_{c+1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right)}} \right] - f^{2}(x)\right| \\ & \leq \mathbb{E}\left[ \abs*{f(X_{c+1}) f(X_{c+1}^{\prime}) - f^{2}(x)} \frac{\Ex*{\kappa\left(x; Z_{c+1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right) \;\middle|\;X_{c+1}, X_{c+1}^{\prime}}}{\Ex*{\kappa\left(x; Z_{c+1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right)}} \right] \end{align}\] Now, fix an arbitrary \(\epsilon > 0\). By continuity of \(f\) at \(x\), there exists a \(\delta > 0\), such that the following holds. \[\forall X, X^{\prime} \in B(x, \delta): \quad \abs*{f(X) f(X^{\prime}) - f^{2}(x)} < \epsilon\] We can consider decompositions of these terms in analogy to [4], i.e., by considering cases with observations lying within this sphere or outside of it, and observe the following. \[\begin{align} & \mathbb{E}\Big[\abs*{f^{2}(X_{1}) - f^{2}(x)} (2s - c) \\ & \quad \times \mathbb{E}\left[\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{1}, D^{\prime}_{[s]}\right) \1*{X_1 \in B(x, \delta)} \middle| X_{1} \right] \Big] \\ & \leq \epsilon \mathbb{E}\left[(2s - c) \mathbb{E}\left[\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{1}, D^{\prime}_{[s]}\right) \1*{X_1 \in B(x, \delta)} \middle| X_{1} \right] \right] \\ & \leq \epsilon \mathbb{E}\left[(2s - c) \mathbb{E}\left[\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{1}, D^{\prime}_{[s]}\right) \middle| X_{1} \right] \right] = \epsilon \end{align}\] \[\begin{align} & \mathbb{E}\Bigg[ \abs*{f(X_{1}) f(X_{c+1}^{\prime}) - f^{2}(x)} \1*{X_1, X_{c+1}^{\prime} \in B(x, \delta)} \\ & \qquad \times \left. \frac{\Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right) \;\middle|\;X_{1}, X_{c+1}^{\prime}}}{\Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right)}} \right] \\ & \leq \epsilon \mathbb{E}\left[ \frac{\Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right) \;\middle|\;X_{1}, X_{c+1}^{\prime}}}{\Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right)}} \1*{X_1, X_{c+1}^{\prime} \in B(x, \delta)} \right] \\ & \leq \epsilon \mathbb{E}\left[ \frac{\Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right) \;\middle|\;X_{1}, X_{c+1}^{\prime}}}{\Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right)}} \right] = \epsilon \end{align}\] \[\begin{align} & \mathbb{E}\Bigg[ \abs*{f(X_{c+1}) f(X_{c+1}^{\prime}) - f^{2}(x)} \frac{\Ex*{\kappa\left(x; Z_{c+1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right) \;\middle|\;X_{c+1}, X_{c+1}^{\prime}}}{\Ex*{\kappa\left(x; Z_{c+1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right)}} \\ & \quad \times \1*{X_{c+1}, X_{c+1}^{\prime} \in B(x, \delta)} \Bigg] \\ & \leq \epsilon \mathbb{E}\left[ \frac{\Ex*{\kappa\left(x; Z_{c+1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right) \;\middle|\;X_{c+1}, X_{c+1}^{\prime}}}{\Ex*{\kappa\left(x; Z_{c+1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right)}} \1*{X_{c+1}, X_{c+1}^{\prime} \in B(x, \delta)} \right] \\ & \leq \epsilon \mathbb{E}\left[ \frac{\Ex*{\kappa\left(x; Z_{c+1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right) \;\middle|\;X_{c+1}, X_{c+1}^{\prime}}}{\Ex*{\kappa\left(x; Z_{c+1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right)}} \right] = \epsilon \end{align}\] For the complementary events, use the fact that if \(X\) or \(X^{\prime}\) does not lie in \(B(x,\delta)\), then \[B(x, \delta) \subseteq B(x, \max\left(\norm*{X - x}, \norm*{X^{\prime} - x}\right)).\] Therefore, \[\begin{align} & \mathbb{E}\Bigg[\abs*{f^{2}(X_{1}) - f^{2}(x)} (2s - c) \\ & \qquad \times \mathbb{E}\left[\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{1}, D^{\prime}_{[s]}\right) \left(1 - \1*{X_1 \in B(x, \delta)}\right) \middle| X_{1} \right] \Bigg] \\ & \leq \mathbb{E}\left[\abs*{f^{2}(X_{1}) - f^{2}(x)} (2s - c) \left\{1-\varphi\left(B\left(x,\delta\right)\right)\right\}^{2s-c-1} \left(1 - \1*{X_1 \in B(x, \delta)}\right) \right] \\ & \leq (2s - c) 1 \left\{1-\varphi\left(B\left(x,\delta\right)\right)\right\}^{2s-c-1} \Ex*{\abs*{f^{2}(X_{1}) - f^{2}(x)}} \end{align}\] In the second case, first recall the form of the conditional expectation from 9. \[\begin{align} & \Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right) \;\middle|\;X_1, X_{c+1}^{\prime}} \\ & = \1*{\norm*{X_{c+1}^{\prime} - x} \leq \norm*{X_{1} - x}} \\ & \qquad \times \left\{1-\varphi\left(B\left(x,\norm*{X_1-x}\right)\right)\right\}^{s-1} \left\{1-\varphi\left(B\left(x,\norm*{X_{c+1}^{\prime}-x}\right)\right)\right\}^{s-c-1} \end{align}\] The indicator is nonzero only if \(\max\left(\norm*{X_{1} - x}, \norm*{X_{c+1}^{\prime} - x}\right) = \norm*{X_{1} - x}\), in which case \[B(x, \delta) \subseteq B(x, \norm*{X_{1} - x})\] Hence \[\begin{align} & \mathbb{E}\Bigg[ \abs*{f(X_{1}) f(X_{c+1}^{\prime}) - f^{2}(x)} \frac{\Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right) \;\middle|\;X_{1}, X_{c+1}^{\prime}}}{\Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right)}} \\ & \quad \left(1 - \1*{X_1, X_{c+1}^{\prime} \in B(x, \delta)}\right) \Bigg] \\ & \overset{(\text{\ref{lem:expec95kernel95prod}})}{\leq} (2s-c)^{2} \mathbb{E}\Big[ \abs*{f(X_{1}) f(X_{c+1}^{\prime}) - f^{2}(x)} \\ & \qquad \Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right) \;\middle|\;X_{1}, X_{c+1}^{\prime}} \left(1 - \1*{X_1, X_{c+1}^{\prime} \in B(x, \delta)}\right) \Big] \\ & \overset{(\text{\ref{lem:cond95expec95kernel95prod}})}{=} (2s-c)^{2} \mathbb{E}\Big[ \abs*{f(X_{1}) f(X_{c+1}^{\prime}) - f^{2}(x)} \1*{\norm*{X_{c+1}^{\prime} - x} \leq \norm*{X_{1} - x}} \\ & \qquad \times \left\{1-\varphi\left(B\left(x,\norm*{X_1-x}\right)\right)\right\}^{s-1} \left\{1-\varphi\left(B\left(x,\norm*{X_{c+1}^{\prime}-x}\right)\right)\right\}^{s-c-1} \\ & \qquad \times \left(1 - \1*{X_1, X_{c+1}^{\prime} \in B(x, \delta)}\right) \Big] \\ & \leq (2s-c)^{2} \left\{1-\varphi\left(B\left(x, \delta \right)\right)\right\}^{s-1} \\ & \qquad \times \mathbb{E}\Big[ \abs*{f(X_{1}) f(X_{c+1}^{\prime}) - f^{2}(x)} \1*{\delta < \norm*{X_{c+1}^{\prime} - x} \leq \norm*{X_{1} - x}} \Big] \\ & \leq (2s-c)^{2} \left\{1-\varphi\left(B\left(x, \delta \right)\right)\right\}^{s-1} \Ex*{\abs*{f(X_{1}) f(X_{c+1}^{\prime}) - f^{2}(x)}} \end{align}\]

Similarly, for the third case set \[K_{\Delta} \mathrel{\vcenter{:}}= \kappa\left(x; Z_{c+1}, D_{[s]}\right) \kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right).\] Then \[\begin{align} & \mathbb{E}\Bigg[ \abs*{f(X_{c+1}) f(X_{c+1}^{\prime}) - f^{2}(x)} \frac{\Ex*{K_{\Delta} \;\middle|\;X_{c+1}, X_{c+1}^{\prime}}}{\Ex*{K_{\Delta}}} \left(1 - \1*{X_{c+1}, X_{c+1}^{\prime} \in B(x, \delta)}\right) \Bigg] \\ & \overset{(\text{\ref{lem:expec95kernel95prod}})}{\leq} (2s-c)^{2} \mathbb{E}\Big[ \abs*{f(X_{c+1}) f(X_{c+1}^{\prime}) - f^{2}(x)} \Ex*{K_{\Delta} \;\middle|\;X_{c+1}, X_{c+1}^{\prime}} \left(1 - \1*{X_{c+1}, X_{c+1}^{\prime} \in B(x, \delta)}\right) \Big] \\ & \overset{(\text{\ref{lem:cond95expec95kernel95prod}})}{=} (2s-c)^{2} \mathbb{E}\Big[\abs*{f(X_{c+1}) f(X_{c+1}^{\prime}) - f^{2}(x)} \left\{1-\varphi\left(B\left(x, \min\left(\norm*{X_{c+1} - x}, \norm*{X_{c+1}^{\prime}-x}\right)\right)\right)\right\}^{s-c-1} \\ & \qquad \times \left\{1-\varphi\left(B\left(x,\max\left(\norm*{X_{c+1} - x}, \norm*{X_{c+1}^{\prime}-x}\right)\right)\right)\right\}^{s-1} \left(1 - \1*{X_{c+1}, X_{c+1}^{\prime} \in B(x, \delta)}\right) \Big] \\ & \leq (2s-c)^{2} \mathbb{E}\Big[\abs*{f(X_{c+1}) f(X_{c+1}^{\prime}) - f^{2}(x)} \left\{1-\varphi\left(B\left(x, \min\left(\norm*{X_{c+1} - x}, \norm*{X_{c+1}^{\prime}-x}\right)\right)\right)\right\}^{s-c-1} \\ & \qquad \times \left\{1-\varphi\left(B\left(x,\delta\right)\right)\right\}^{s-1} \\ & \qquad \times \left(1 - \1*{X_{c+1}, X_{c+1}^{\prime} \in B(x, \delta)}\right) \Big] \\ & \leq (2s-c)^{2} \left\{1-\varphi\left(B\left(x,\delta\right)\right)\right\}^{s-1} \\ & \qquad \times \Ex*{\abs*{f(X_{c+1}) f(X_{c+1}^{\prime}) - f^{2}(x)}} \end{align}\] The remaining expectation factors are finite: \[\begin{align} \mathbb{E}\Big[\abs*{f^{2}(X) - f^{2}(x)}\Big] & \leq \mathbb{E}\Big[f^{2}(X)\Big] + f^{2}(x) = \norm*{f}_{L_2}^{2} + f^{2}(x) \end{align}\] \[\begin{align} & \mathbb{E}\Big[\abs*{f(X) f(X^{\prime}) - f^{2}(x)}\Big] \leq \mathbb{E}\Big[\abs*{f(X) f(X^{\prime})}\Big] + f^{2}(x) \\ & \leq \mathbb{E}\Big[\abs*{f(X)} \abs*{f(X^{\prime})}\Big] + f^{2}(x) = \norm*{f}_{L^{1}}^{2} + f^{2}(x) \end{align}\]

Since \(f\in L^{2}(\mathcal{X})\) on a bounded domain, \(\norm*{f}_{L^{1}}<\infty\). Hence \[\begin{align} & \mathbb{E}\Bigg[\abs*{f^{2}(X_{1}) - f^{2}(x)} (2s - c) \mathbb{E}\left[\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{1}, D^{\prime}_{[s]}\right) \left(1 - \1*{X_1 \in B(x, \delta)}\right) \middle| X_{1} \right] \Bigg] \longrightarrow 0 \quad \text{as} \quad s \longrightarrow \infty \end{align}\] \[\begin{align} & \mathbb{E}\Bigg[ \abs*{f(X_{1}) f(X_{c+1}^{\prime}) - f^{2}(x)} \frac{\Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right) \;\middle|\;X_{1}, X_{c+1}^{\prime}}}{\Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right)}} \left(1 - \1*{X_1, X_{c+1}^{\prime} \in B(x, \delta)}\right) \Bigg] \longrightarrow 0 \quad \text{as} \quad s \longrightarrow \infty \end{align}\] \[\begin{align} & \mathbb{E}\Bigg[ \abs*{f(X_{c+1}) f(X_{c+1}^{\prime}) - f^{2}(x)} \frac{\Ex*{\kappa\left(x; Z_{c+1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right) \;\middle|\;X_{c+1}, X_{c+1}^{\prime}}}{\Ex*{\kappa\left(x; Z_{c+1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D^{\prime}_{[s]}\right)}} \\ & \qquad \times \left(1 - \1*{X_{c+1}, X_{c+1}^{\prime} \in B(x, \delta)}\right) \Bigg] \longrightarrow 0 \quad \text{as} \quad s \longrightarrow \infty \end{align}\] Combining these bounds, for large enough \(s\) each of \(\norm*{(A) - f^{2}(x)}\), \(\norm*{(B) - f^{2}(x)}\), and \(\norm*{(C) - f^{2}(x)}\) is bounded by \(2\epsilon\). As \(\epsilon\) was arbitrary this concludes the proof.

 ◻

Lemma 11.
* The following inequalities hold. \[\begin{align} & \forall i \in [c] \; \forall j \in \{c+1, \dotsc, s\}: \\ & \Ex*{\kappa\left(x; Z_{i}, D_{[s]}\right)\kappa\left(x; Z_{j}^{\prime}, D^{\prime}_{[s]}\right)} \leq \frac{1}{s(2s-c)} \end{align}\] \[\begin{align} & \forall i,j \in \{c+1, \dotsc, s\}: \\ & \Ex*{\kappa\left(x; Z_{i}, D_{[s]}\right)\kappa\left(x; Z_{j}^{\prime}, D^{\prime}_{[s]}\right)} \leq \frac{2}{s(2s-c)} \end{align}\]

Proof of 11.
* Both inequalities are immediate from the exact identities in 8. ◻

7.3 DNN Kernel Expectations↩︎

We now collect the expectation calculations for the single-scale DNN kernel and its first projection under the nonparametric regression setup.

Lemma 12 (NPR - DNN Kernel Expectation).
* Let \(x\) denote a point of interest. Then \[\Ex*{h_s\left(x; D_{[s]}\right)} = \Ex*{Y_1 s \Ex*{\kappa\left(x; Z_1, D_{[s]}\right) \;\middle|\;X_{1}}} \longrightarrow \mu\left(x\right) \quad \text{as} \quad s \longrightarrow \infty\]

Proof of 12. This result follows immediately from 7 and the following observation. \[\begin{align} & \Ex*{Y_1 s \Ex*{\kappa\left(x; Z_1, D_{[s]}\right) \;\middle|\;X_{1}}} = \Ex*{\left(\mu\left(X_1\right) + \varepsilon_1\right) s \Ex*{\kappa\left(x; Z_1, D_{[s]}\right) \;\middle|\;X_{1}}} \\ & = \Ex*{\left(\mu\left(X_1\right) + \Ex*{\varepsilon_1 \;\middle|\;X_1}\right) s \Ex*{\kappa\left(x; Z_1, D_{[s]}\right) \;\middle|\;X_{1}}} \\ & = \Ex*{\mu\left(X_1\right) s \Ex*{\kappa\left(x; Z_1, D_{[s]}\right) \;\middle|\;X_{1}}} \\ & \overset{\text{(\ref{lem:dem13})}}{\longrightarrow} \mu\left(x\right) \quad \text{as} \quad s \longrightarrow \infty \end{align}\] ◻

Lemma 13 (NPR - DNN Hájek Kernel Expectation).
* Let \(z_1 = (x_1, y_1)\) denote a specific realization of \(Z\) and \(x\) denote a point of interest. Then \[\begin{align} \psi_{s}^{1}\left(x; z_1\right) & = \mu(x_1)\Ex*{\kappa\left(x; Z_1, D_{[s]}\right)\;\middle|\;X_1 = x_1} \\ & \quad + \varepsilon_1 \Ex*{\kappa\left(x; Z_1, D_{[s]}\right)\;\middle|\;X_1 = x_1} \\ & \quad + \Ex*{\sum_{i = 2}^{s} \kappa\left(x; Z_{i}, D_{[s]}\right) \mu(X_{i})\;\middle|\;X_1 = x_1}. \end{align}\]

Proof of 13. \[\begin{align} & \psi_{s}^{1}\left(x; z_1\right) = \Ex*{h_{s}\left(x; D_{[s]}\right) \;\middle|\;Z_1 = z_1} = \Ex*{\sum_{i = 1}^{s} \kappa\left(x; Z_{i}, D_{[s]}\right) Y_{i} \;\middle|\;Z_1 = z_1} \\ & = \mathbb{E}\Bigg[\left(\mu(x_1) + \varepsilon_1\right)\kappa\left(x; Z_1, D_{[s]}\right) \\ & \qquad\qquad\quad + \sum_{i = 2}^{s} \kappa\left(x; Z_{i}, D_{[s]}\right) \mu(X_{i})\;\Bigg|\; Z_1 = z_1 \Bigg] \\ & = \mu(x_1)\Ex*{\kappa\left(x; Z_1, D_{[s]}\right)\;\middle|\;X_1 = x_1} \\ & \qquad + \varepsilon_1 \Ex*{\kappa\left(x; Z_1, D_{[s]}\right)\;\middle|\;X_1 = x_1} \\ & \qquad + \Ex*{\sum_{i = 2}^{s} \kappa\left(x; Z_{i}, D_{[s]}\right) \mu(X_{i})\;\middle|\;X_1 = x_1} \end{align}\] ◻

7.4 DNN Kernel Variances & Covariances↩︎

We next bound the second moments and overlap covariances needed for the single-scale DNN Hájek-dominance argument.

Lemma 14 (Adapted from [4]).
* Let \(D_{[s]} = \{Z_1, \dotsc, Z_{s}\}\) be a vector of i.i.d.random variables drawn from \(P\). Furthermore, let \[\Omega_{s}\left(x\right) = \Ex*{h_{s}^{2}\left(x; D_{[s]}\right)}.\] Then, \[\Omega_{s}\left(x\right) = \Ex*{\left(\mu\left(X_1\right)+ \varepsilon_1\right)^2 s \Ex*{\kappa\left(x; Z_1, D_{[s]}\right) \;\middle|\;X_{1}}} \lesssim \mu^2(x) + \overline{\sigma}_{\varepsilon}^2 + o(1)\]

Proof of 14.
* This result follows immediately from 7 and the following observation. \[\begin{align} & \Omega_{s}\left(x\right) = \Ex*{h_{s}^{2}\left(x; D_{[s]}\right)} = \Ex*{\left(\sum_{i = 1}^{s}\kappa\left(x; Z_{i}, D_{[s]}\right)Y_{i}\right)^2} \\ & = \Ex*{\sum_{i = 1}^{s}\sum_{j = 1}^{s}\left(\kappa\left(x; Z_{i}, D_{[s]}\right)\kappa\left(x; Z_{j}, D_{[s]}\right)Y_{i}Y_{j}\right)} = \Ex*{s \kappa\left(x; Z_{1}, D_{[s]}\right)Y_{1}^2} \\ & = \Ex*{Y_{1}^2 s \Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right) \;\middle|\;X_{1}}} = \Ex*{\left(\mu\left(X_1\right) + \varepsilon_1\right)^2 s \Ex*{\kappa\left(x; Z_1, D_{[s]}\right) \;\middle|\;X_{1}}} \\ & = \Ex*{\left(\mu^2\left(X_1\right) + 2\mu\left(X_1\right)\varepsilon_1 + \varepsilon_1^2\right)s \Ex*{\kappa\left(x; Z_1, D_{[s]}\right) \;\middle|\;X_{1}}} \\ & = \mathbb{E}\left[\left(\mu^2\left(X_1\right) + 2\mu\left(X_1\right) \Ex*{\varepsilon_1 \;\middle|\;X_1} + \Ex*{\varepsilon_1^2 \;\middle|\;X_1}\right) s \Ex*{\kappa\left(x; Z_1, D_{[s]}\right) \;\middle|\;X_{1}}\right] \\ & = \Ex*{\left(\mu^2\left(X_1\right) +\sigma_{\varepsilon}^{2}(X_1)\right) s \Ex*{\kappa\left(x; Z_1, D_{[s]}\right) \;\middle|\;X_{1}}} \\ & \overset{\text{(\ref{lem:dem13})}}{\longrightarrow} \mu^2\left(x\right) +\sigma_{\varepsilon}^{2}(x) \quad \text{as} \quad s \longrightarrow \infty \end{align}\] Furthermore, we have the following inequality. \[\mu^2(x) + \sigma_{\varepsilon}^2(x) \leq \mu^2\left(x\right) + \overline{\sigma}_{\varepsilon}^{2}\] Thus, we obtain the desired result. ◻

Lemma 15.
* Let \(D_{[s]} = \{Z_1, \dotsc, Z_{s}\}\) be a vector of i.i.d.random variables drawn from \(P\). Let \(D_{[s]}^{\prime} = \{Z_1, \dotsc, Z_{c}, Z_{c+1}^{\prime}, \dotsc, Z_{s}^{\prime}\}\) where \(Z_{c+1}^{\prime}, \dotsc, Z_{s}^{\prime}\) are i.i.d.draws from \(P\) that are independent of \(D_{[s]}\). Furthermore, let \[\Omega_{s}^{c}\left(x\right) = \mathbb{E}\left[h_{s}\left(x; D_{[s]}\right) h_{s}\left(x; D_{[s]}^{\prime}\right)\right].\] Then, \[\Omega_{s}^{c}\left(x\right) \lesssim \mu^2(x) + \overline{\sigma}_{\varepsilon}^2 + o(1)\]

Proof of 15. \[\begin{align} & \Omega_{s}^{c}\left(x\right) = \mathbb{E}\left[h_{s}\left(x; D_{[s]}\right) h_{s}\left(x; D_{[s]}^{\prime}\right)\right] \\ & = \mathbb{E}\left[ \left(\sum_{i = 1}^{s}\kappa\left(x; Z_{i}, D_{[s]}\right)Y_{i}\right) \left(\sum_{j = 1}^{c}\kappa\left(x; Z_{j}, D_{[s]}^{\prime}\right)Y_{j} + \sum_{j = c+1}^{s}\kappa\left(x; Z_{j}^{\prime}, D_{[s]}^{\prime}\right)Y_{j}^{\prime}\right) \right] \\ & = \underbrace{\Ex*{c \kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{1}, D_{[s]}^{\prime}\right)Y_{1}^{2}}}_{(A)} \\ & \quad + 2 \underbrace{\Ex*{c(s-c) \kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D_{[s]}^{\prime}\right)Y_{1}Y_{c+1}^{\prime}}}_{(B)} \\ & \quad + \underbrace{\Ex*{(s-c)^2 \kappa\left(x; Z_{c+1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D_{[s]}^{\prime}\right)Y_{c+1}Y_{c+1}^{\prime}}}_{(C)} \end{align}\] Starting from this decomposition, we will analyze the terms one by one using 10. \[\begin{align} (A) & = \Ex*{c\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{1}, D_{[s]}^{\prime}\right)Y_{1}^{2}} \\ & = \frac{c}{2s-c} \Ex*{\left(\mu(X_1) + \varepsilon_1\right)^2 (2s-c) \Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{1}, D_{[s]}^{\prime}\right) \;\middle|\;X_{1}}} \\ & = \frac{c}{2s-c} \Ex*{\left(\mu^{2}(X_1) + \sigma_{\varepsilon}^{2}(X_1)\right) (2s-c) \Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{1}, D_{[s]}^{\prime}\right) \;\middle|\;X_{1}}} \\ & \overset{\text{(\ref{lem:kernel95prod95dirac95convergence})}}{\lesssim} \frac{c}{2s-c} \left(\mu^2(x) + \sigma_{\varepsilon}^{2}(x)\right) + o(1) \end{align}\] Similarly, we can find the following. \[\begin{align} (B) & = \Ex*{c(s-c) \kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D_{[s]}^{\prime}\right)Y_{1}Y_{c+1}^{\prime}} \\ & \overset{(\text{\ref{lem:expec95kernel95prod}})}{=} \frac{c(s-c)}{s(2s-c)} \mathbb{E}\Bigg[ \left(\mu(X_1) + \varepsilon_{1}\right)\left(\mu(X_{c+1}^{\prime}) + \varepsilon_{c+1}^{\prime}\right) \\ & \qquad \times \left. \frac{\Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D_{[s]}^{\prime}\right) \;\middle|\;X_{1}, X_{c+1}^{\prime}}}{\Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D_{[s]}^{\prime}\right)}} \right] \\ & = \frac{c(s-c)}{s(2s-c)} \mathbb{E}\left[ \mu(X_1) \mu(X_{c+1}^{\prime}) \frac{\Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D_{[s]}^{\prime}\right) \;\middle|\;X_{1}, X_{c+1}^{\prime}}}{\Ex*{\kappa\left(x; Z_{1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D_{[s]}^{\prime}\right)}} \right] \\ & \overset{(\text{\ref{lem:kernel95prod95dirac95convergence}})}{\lesssim} \frac{c(s-c)}{s(2s-c)} \mu^2(x) + o(1) \end{align}\] The third term can be asymptotically bounded in the following way. \[\begin{align} (C) & = \Ex*{(s-c)^2 \kappa\left(x; Z_{c+1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D_{[s]}^{\prime}\right)Y_{c+1}Y_{c+1}^{\prime}} \\ & \overset{(\text{\ref{lem:expec95kernel95prod}})}{=} \frac{2(s-c)^2}{s(2s-c)} \mathbb{E}\Bigg[ \mu(X_{c+1}) \mu(X_{c+1}^{\prime}) \\ & \qquad\qquad\qquad \times \frac{\Ex*{\kappa\left(x; Z_{c+1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D_{[s]}^{\prime}\right) \;\middle|\;X_{c+1}, X_{c+1}^{\prime}}}{\Ex*{\kappa\left(x; Z_{c+1}, D_{[s]}\right)\kappa\left(x; Z_{c+1}^{\prime}, D_{[s]}^{\prime}\right)}} \Bigg] \\ & \overset{(\text{\ref{lem:kernel95prod95dirac95convergence}})}{\lesssim} \frac{2(s-c)^2}{s(2s-c)} \mu^2(x) + o(1) \end{align}\] The coefficients \(c/(2s-c)\), \(c(s-c)/(s(2s-c))\), and \(2(s-c)^2/(s(2s-c))\) are uniformly bounded over \(1 \leq c \leq s-1\). The result of 15 follows immediately by summing up the asymptotic bounds for the individual terms. ◻

7.5 Single-Scale DNN Results↩︎

This subsection establishes the single-scale DNN results used in the TDNN analysis.

Lemma 16 (DNN selector profile).
* Under 5 [asm:tdnn95design95density], for each scale \(s\) define \[q_s(X_1) \mathrel{\vcenter{:}}= \Ex*{ \kappa\left(x; Z_1, D_{[s]}\right) \;\middle|\;X_1 }.\] Let \(R_1 \mathrel{\vcenter{:}}=\norm*{X_1-x}_2\) and \(F_x(t) \mathrel{\vcenter{:}}=\Pb*{\norm*{X-x}_2 \leq t}\). Then \(U_1 \mathrel{\vcenter{:}}= F_x(R_1)\) is uniformly distributed on \((0,1)\), \[\label{eq:dnn95selector95profile} q_s(X_1) = \left(1-U_1\right)^{s-1},\tag{20}\] and, for any \(a>0\), \[\label{eq:dnn95selector95moment} \Ex*{q_s(X_1)^a} = \frac{1}{a(s-1)+1}.\tag{21}\] More generally, for any \(a,b \geq 0\) with \(a+b>0\), \[\label{eq:dnn95selector95cross95moment} \Ex*{q_s(X_1)^a q_t(X_1)^b} = \frac{1}{a(s-1)+b(t-1)+1}.\tag{22}\]

The quantity \(q_s(X_1)\) is the conditional probability that, once \(X_1\) is fixed, none of the remaining \(s-1\) observations lands closer to \(x\) than \(X_1\). This geometric interpretation is what later turns selector moments into the \(s^{-1}\) first-projection rate.

Proof. Observation \(1\) is the nearest neighbor among \(s\) draws exactly when the remaining \(s-1\) observations all lie outside \(B(x,R_1)\). Hence 20 follows. The density condition implies that \(U_1=F_x(R_1)\) is uniform on \((0,1)\), so the displayed moment identities follow by integrating powers of \(1-U_1\) over the unit interval. ◻

Lemma 17 (Single-scale DNN Hájek dominance).
* Consider a data-generating process as outlined in 4 and 5. Let \(s = o(n)\). Then the DNN estimator satisfies the asymptotic Hájek dominance condition. In particular, \[\label{eq:dnn95hajek95input95rates} \zeta_s^s(x) \lesssim 1, \qquad \zeta_s^1(x) \asymp s^{-1}.\tag{23}\]

Proof. By 14, \[\zeta_s^s(x) \leq \Omega_s(x) \lesssim 1.\] For the lower bound on the first-projection variance, use the selector profile \(q_s\) from 16. The decomposition in 13 gives \(h_s^{(1)}(x; Z_1) = q_s(X_1)\varepsilon_1 + b_s(X_1)\) for an \(X_1\)-measurable remainder \(b_s\). Since \(\Ex{\varepsilon_1 \;\middle|\;X_1} = 0\), the law of total variance gives \[\zeta_s^1(x) \geq \Ex*{q_s(X_1)^2\sigma_\varepsilon^2(X_1)} \geq \underline{\sigma}_{\varepsilon}^{2}\, \Ex*{q_s(X_1)^2}.\] 21 with \(a=2\) gives \(\Ex*{q_s(X_1)^2} = (2(s-1)+1)^{-1}\), and hence \(\zeta_s^1(x) \gtrsim s^{-1}\). The cross-scale identity in 22 is used later in the TDNN non-cancellation argument. The reverse bound \(\zeta_s^1(x) \lesssim s^{-1}\) is the first-projection variance rate from [4], adapted to the present heteroskedastic setup by using the uniform variance bound in 5 [asm:tdnn95design95variance]. Consequently, \[\zeta_s^1(x) \asymp s^{-1}.\] Therefore, \[\frac{s}{n} \left( \frac{\zeta_s^s(x)}{s\zeta_s^1(x)} - 1 \right) = O\left(\frac{s}{n}\right) \longrightarrow 0\] as \(s=o(n)\), which is exactly 1 for the single-scale DNN estimator. ◻

Remark 8. The two rates in 23 depend on separate properties of the kernel. The bound \(\zeta_s^s(x) \lesssim 1\) follows from the fact that the nearest-neighbor selector picks at most one observation per subsample, so the full-kernel second moment is controlled by the second moment of whatever function is being selected. The rate \(\zeta_s^1(x) \asymp s^{-1}\) follows from the selection probability: a fixed observation is the nearest neighbor among \(s\) draws with probability of order \(s^{-1}\), independently of the function value at that observation. Consequently, both rates hold for any kernel of the form \(\kappa(x;\,Z_i,D_{[s]})\,f(Z_i)\) in which \(f\) has finite second moment — the regression response \(Y\) is one instance, but the rates are the same whenever the same nearest-neighbor selector is applied to a square-integrable function of the observation.

Lemma 18 (Single-scale DNN verifies the row-wise \(L^r\) condition).
* Consider a data-generating process as outlined in 4, 5, and 8. Then the single-scale DNN first projection satisfies 2. More precisely, with \(r_\eta \mathrel{\vcenter{:}}= 1 + \frac{1}{2}\min\{\eta,2\}\), where \(\eta\) is from 8, \[\Ex*{ \left( \frac{h_s^{(1)}(x; Z_1)^2}{\zeta_s^1(x)} \right)^{r_\eta} } \lesssim s^{r_\eta-1},\] so the ratio in 2 vanishes whenever \(s=o(n)\).

Proof. Write \(r=r_\eta\) for the exponent fixed in the statement. The \(b_s\) bound below uses the upper rate \(\zeta_s^1(x) \lesssim s^{-1}\) from 17, which rests on the first-projection variance calculation of [4]. Selector moment. For \(q_s\) from 16, 21 with \(a=2r\) gives \[\Ex*{q_s(X_1)^{2r}} = \frac{1}{2r(s-1)+1} \lesssim s^{-1}.\] By 13, the first projection decomposes as \(h_s^{(1)}(x; Z_1) = q_s(X_1)\,\varepsilon_1 + b_s(X_1)\), where \[b_s(u) \coloneq \mu(u)\,q_s(u) + \mathbb{E}\!\left[ \sum_{i=2}^{s}\kappa\!\left(x; Z_i, D_{[s]}\right)\mu(X_i) \,\middle|\, X_1 = u \right] - \theta_s(x).\] Bias remainder. Since \(\mathcal{X}\) is compact and \(\mu\) is continuous under 5 [asm:tdnn95design95compact] [asm:tdnn95design95regression], write \(M_\mu \mathrel{\vcenter{:}}=\sup_{v\in\mathcal{X}}\abs*{\mu(v)} < \infty\). For the \(b_s\) remainder: since the selector weights \(\kappa(x; Z_i, D_{[s]})\) sum to one almost surely (6), \[\mu(u)\,q_s(u) + \Ex*{\sum_{i=2}^{s}\kappa\mu(X_i) \;\middle|\;X_1=u} = \Ex*{\sum_{i=1}^{s}\kappa(x;Z_i,D_{[s]})\mu(X_i) \;\middle|\;X_1=u}.\] The right-hand side is bounded by \(M_\mu\) in absolute value. Together with \(\abs*{\theta_s(x)} \leq M_\mu\) (since \(\theta_s(x)\) is itself an expectation of the same convex combination of \(\mu\) values, bounded by \(M_\mu\) via 6), we have \(\abs*{b_s(u)} \leq 2M_\mu\) uniformly in \(u\). Because \(\Ex{\varepsilon_1 \;\middle|\;X_1} = 0\) kills the cross term, \[\zeta_s^1(x) = \Varb*{h_s^{(1)}(x; Z_1)} = \Ex*{q_s(X_1)^2\,\sigma_\varepsilon^2(X_1)} + \Ex*{b_s(X_1)^2} \geq \Ex*{b_s(X_1)^2},\] so \(\Ex{b_s(X_1)^2} \leq \zeta_s^1(x) \lesssim s^{-1}\) by 17. The \(L^\infty\) bound then gives \[\Ex*{\abs*{b_s(X_1)}^{2r}} \leq (2M_\mu)^{2(r-1)}\,\Ex*{b_s(X_1)^2} \lesssim s^{-1}.\] Conditional error moment. By 8, the choice of \(r=r_\eta\), and the boundedness of \(\mu\) on \(\mathcal{X}\), there is a constant \(C_{\varepsilon,r}<\infty\) such that \[\label{eq:dnn95conditional95error95moment} \sup_{u\in\mathcal{X}} \Ex*{\abs*{\varepsilon}^{2r}\;\middle|\;X=u} \leq C_{\varepsilon,r}.\tag{24}\] Since \(q_s(X_1)\) is \(X_1\)-measurable, iterated expectation and 21 24 give \[\label{eq:dnn95selector95error95moment} \begin{align} \Ex*{ q_s(X_1)^{2r}\abs*{\varepsilon_1}^{2r} } & = \Ex*{ q_s(X_1)^{2r} \Ex*{\abs*{\varepsilon_1}^{2r}\;\middle|\;X_1} } \\ & \leq C_{\varepsilon,r} \Ex*{q_s(X_1)^{2r}} \lesssim s^{-1}. \end{align}\tag{25}\] Combining 25 with the \(b_s\) bound and \(\abs*{a+b}^{2r}\leq 2^{2r-1}\left(\abs*{a}^{2r}+\abs*{b}^{2r}\right)\) yields \[\label{eq:dnn95first95projection95moment} \Ex*{\abs*{h_s^{(1)}(x; Z_1)}^{2r}} \lesssim s^{-1}.\tag{26}\] By 17, \(\zeta_s^1(x) \asymp s^{-1}\). Consequently, \[\Ex*{ \left( \frac{h_s^{(1)}(x; Z_1)^2}{\zeta_s^1(x)} \right)^r } = \frac{\Ex*{\abs*{h_s^{(1)}(x; Z_1)}^{2r}}}{\left(\zeta_s^1(x)\right)^r} \lesssim s^{r-1}.\] Since \(r>1\) and \(s=o(n)\), \[\frac{1}{n^{r-1}} \Ex*{ \left( \frac{h_s^{(1)}(x; Z_1)^2}{\zeta_s^1(x)} \right)^r } \lesssim \left(\frac{s}{n}\right)^{r-1} \longrightarrow 0.\] This proves 2. ◻

8 Extension to the TDNN Estimator↩︎

8.1 Notation↩︎

For the TDNN extension, we only need the first TDNN projection, its associated variance terms, and the effective-localizer notation.

\[\psi_{\mathfrak{S}}^{1}(x; d_{1}) = \Ex*{h_{\mathfrak{S}}\left(x; D_{[s_{2}]}\right) \;\middle|\;Z_{1} = d_{1}}\] Let \[\theta_{\mathfrak{S}}(x) \coloneq \Ex*{h_{\mathfrak{S}}\left(x; D_{[s_{2}]}\right)}\] denote the finite-sample TDNN mean, and set \[h_{\mathfrak{S}}^{(1)}\left(x; d_{1}\right) = \psi_{\mathfrak{S}}^{1}(x; d_{1}) - \theta_{\mathfrak{S}}(x).\] For the TDNN kernel, write \[\begin{align} \zeta_{\mathfrak{S}}^{1}\left(x\right) & = \Varb*{h_{\mathfrak{S}}^{(1)}\left(x; Z_{1}\right)} \\ \zeta_{\mathfrak{S}}^{s_2}\left(x\right) & = \Varb*{h_{\mathfrak{S}}\left(x; D_{[s_2]}\right)}. \end{align}\] The two-scale kernel also admits an observation-level representation that will be useful in the Hájek-dominance argument. Define the effective TDNN weights by \[\begin{align} \tilde{w}_i\left(x; D_{[s_2]}\right) & \coloneq w_1^* \binom{s_2}{s_1}^{-1} \sum_{\ell \in L_{s_2,s_1}} \1*{i \in \ell}\kappa\left(x; Z_i, D_{\ell}\right) \\ & \qquad + w_2^* \kappa\left(x; Z_i, D_{[s_2]}\right), \qquad i = 1,\dotsc,s_2. \end{align}\] Then \[h_{\mathfrak{S}}\left(x; D_{[s_2]}\right) = \sum_{i=1}^{s_2} \tilde{w}_i\left(x; D_{[s_2]}\right) Y_i.\] The induced effective TDNN localizer is \[\tilde{\tau}_{\mathfrak{S}}(u) \coloneq s_2 \Ex*{\tilde{w}_1\left(x; D_{[s_2]}\right) \;\middle|\;X_1 = u}.\]

8.2 Two-Scale TDNN Argument↩︎

This subsection establishes the two-scale TDNN result from the corresponding single-scale DNN bounds. The two-scale kernel decomposes into an embedded \(s_1\)-scale DNN average inside an \(s_2\)-sample plus the ordinary \(s_2\)-scale DNN kernel, so both pieces can be analyzed through the single-scale results derived above.

In what follows, we use the first-projection objects \(\psi_s^1\), \(\psi_{\mathfrak S}^1\), \(h_s^{(1)}\), and \(h_{\mathfrak S}^{(1)}\), together with the variance terms \(\zeta_s^1\), \(\zeta_s^s\), \(\zeta_{\mathfrak S}^1\), and \(\zeta_{\mathfrak S}^{s_2}\). By 17, the corresponding single-scale DNN estimator satisfies 1, with \[\zeta_s^s(x) \lesssim 1, \qquad \zeta_s^1(x) \asymp s^{-1},\] for every scale \(s=o(n)\).

The first structural step is to separate the signed TDNN combination from the positive averaging operator over \(s_1\)-subsets. Define \[\label{eq:embedded95dnn95average} \bar h_{s_1 \mid s_2}\left(x; D_{[s_2]}\right) \coloneq \binom{s_2}{s_1}^{-1} \sum_{\ell \in L_{s_2,s_1}} h_{s_1}\left(x; D_{\ell}\right).\tag{27}\] Then the TDNN kernel can be written as \[\label{eq:tdnn95embedded95decomp} h_{\mathfrak S}\left(x; D_{[s_2]}\right) = w_1^* \bar h_{s_1 \mid s_2}\left(x; D_{[s_2]}\right) + w_2^* h_{s_2}\left(x; D_{[s_2]}\right).\tag{28}\] Under 7, the coefficient \(w_1^*\) is negative and \(w_2^*>1\), so the final TDNN observation-level weights are signed. Accordingly, Jensen’s inequality is applied to the positive uniform average in 27 inside the decomposition 28 , not to the signed TDNN combination itself.

8.2.0.1 Step 1: Control the embedded \(s_1\)-scale second moment.

Because \(\bar h_{s_1 \mid s_2}\) is the uniform average of the \(s_1\)-scale kernels over all \(s_1\)-subsets of a fixed \(s_2\)-sample, Jensen’s inequality gives \[\label{eq:embedded95jensen} \bar h_{s_1 \mid s_2}^2\left(x; D_{[s_2]}\right) \leq \binom{s_2}{s_1}^{-1} \sum_{\ell \in L_{s_2,s_1}} h_{s_1}^2\left(x; D_{\ell}\right).\tag{29}\] Taking expectations and using exchangeability, \[\label{eq:embedded95second95moment95bound} \Ex*{\bar h_{s_1 \mid s_2}^2\left(x; D_{[s_2]}\right)} \leq \Ex*{h_{s_1}^2\left(x; D_{[s_1]}\right)} \lesssim 1,\tag{30}\] where the last step is precisely the single-scale DNN input at scale \(s_1\). The averaging bound in 29 smooths rather than amplifies the embedded \(s_1\)-scale contribution: averaging over many \(s_1\)-subsets cannot have larger second moment than the average of their individual second moments. Combining 30 with the analogous DNN bound at scale \(s_2\), and using that \(w_1^*\) and \(w_2^*\) stay bounded under 7, we obtain \[\label{eq:tdnn95full95kernel95strategy} \begin{align} \zeta_{\mathfrak S}^{s_2}(x) & \leq \Ex*{h_{\mathfrak S}^2\left(x; D_{[s_2]}\right)} \\ & \lesssim \left(w_1^*\right)^2 \Ex*{\bar h_{s_1 \mid s_2}^2\left(x; D_{[s_2]}\right)} + \left(w_2^*\right)^2 \Ex*{h_{s_2}^2\left(x; D_{[s_2]}\right)} \\ & \lesssim 1. \end{align}\tag{31}\] Conceptually, this shows that the embedded \(s_1\)-piece can be controlled directly at its own scale before it is recombined with the \(s_2\)-piece.

8.2.0.2 Step 2: Identify the first projection of the embedded \(s_1\)-scale term.

The key combinatorial identity is simpler at the projection level than at the raw second-moment level. Conditionally on \(Z_1 = z_1\), a uniformly drawn \(s_1\)-subset of an \(s_2\)-sample contains observation \(1\) with probability \(s_1/s_2\). If the subset contains \(1\), its conditional law matches the usual \(s_1\)-scale DNN setup with one observation fixed. If the subset does not contain \(1\), its conditional expectation is just the finite-sample DNN mean \(\theta_{s_1}(x) = \Ex*{h_{s_1}\left(x; D_{[s_1]}\right)}\). Thus, with \[\psi_{s_1 \mid s_2}^{1}(x; z_1) \coloneq \Ex*{\bar h_{s_1 \mid s_2}\left(x; D_{[s_2]}\right) \;\middle|\;Z_1 = z_1},\] we obtain the exact identity \[\label{eq:embedded95projection95identity} \psi_{s_1 \mid s_2}^{1}(x; z_1) = \frac{s_1}{s_2}\psi_{s_1}^{1}(x; z_1) + \left(1-\frac{s_1}{s_2}\right)\theta_{s_1}(x),\tag{32}\] Subtracting \(\theta_{s_1}(x)\) from 32 gives \[\label{eq:embedded95first95projection95identity} \bar h_{s_1 \mid s_2}^{(1)}(x; z_1) \coloneq \psi_{s_1 \mid s_2}^{1}(x; z_1) - \theta_{s_1}(x) = \frac{s_1}{s_2} h_{s_1}^{(1)}(x; z_1).\tag{33}\] This identity separates the selection step from the nearest-neighbor competition at scale \(s_1\). Observation \(1\) must first be selected into the relevant \(s_1\)-subset and only then can it affect the ordinary DNN nearest-neighbor competition at that scale.

8.2.0.3 Step 3: Close the TDNN first-projection rate.

By linearity of conditional expectation and 33 , \[\label{eq:tdnn95first95projection95linear95combination} h_{\mathfrak S}^{(1)}(x; z_1) = w_1^* \frac{s_1}{s_2} h_{s_1}^{(1)}(x; z_1) + w_2^* h_{s_2}^{(1)}(x; z_1).\tag{34}\] For each scale \(s\), 13 yields the exact decomposition \[\label{eq:dnn95first95projection95qb} h_s^{(1)}(x; Z_1) = q_s(X_1)\varepsilon_1 + b_s(X_1),\tag{35}\] where \[\begin{align} q_s(u) & \coloneq \Ex*{ \kappa\left(x; Z_1, D_{[s]}\right) \;\middle|\;X_1 = u }, \\ b_s(u) & \coloneq \mu(u)\,q_s(u) \\ & \quad + \Ex*{ \sum_{i = 2}^{s} \kappa\left(x; Z_i, D_{[s]}\right)\mu(X_i) \;\middle|\;X_1 = u } \\ & \quad - \Ex*{h_s\left(x; D_{[s]}\right)}. \end{align}\] Here \(b_s(X_1)\) is \(X_1\)-measurable, so \[\label{eq:dnn95first95projection95variance95decomp} \zeta_s^1(x) = \Ex*{q_s(X_1)^2 \sigma_\varepsilon^2(X_1)} + \Ex*{b_s(X_1)^2}\tag{36}\] because \(\Ex*{\varepsilon_1 \;\middle|\;X_1} = 0\) kills the cross term. Combining 34 with 35 , define \[\label{eq:tdnn95first95projection95qb} h_{\mathfrak S}^{(1)}(x; Z_1) = \tilde{q}_{\mathfrak S}(X_1)\varepsilon_1 + \tilde{b}_{\mathfrak S}(X_1),\tag{37}\] where \[\tilde{q}_{\mathfrak S}(u) \coloneq a_{\mathfrak S} q_{s_1}(u) + b_{\mathfrak S} q_{s_2}(u), \qquad \tilde{b}_{\mathfrak S}(u) \coloneq a_{\mathfrak S} b_{s_1}(u) + b_{\mathfrak S} b_{s_2}(u),\] with \[a_{\mathfrak S} \coloneq w_1^* \frac{s_1}{s_2}, \qquad b_{\mathfrak S} \coloneq w_2^*.\] Therefore, \[\label{eq:tdnn95first95projection95variance95decomp} \zeta_{\mathfrak S}^{1}(x) = \Ex*{\tilde{q}_{\mathfrak S}(X_1)^2 \sigma_\varepsilon^2(X_1)} + \Ex*{\tilde{b}_{\mathfrak S}(X_1)^2}.\tag{38}\]

This is the point where bias correction and variance part company: the coefficients are chosen to cancel leading deterministic bias terms, but they must still leave a stochastic first-order signal on the \(s_2^{-1/2}\) scale.

Lemma 19 (Uniform non-cancellation of the TDNN selector coefficient).
* Suppose 7 holds. Then there exist constants \(0 < c_q < C_q < \infty\) such that \[\frac{c_q}{s_2} \leq \Ex*{\tilde{q}_{\mathfrak S}(X_1)^2} \leq \frac{C_q}{s_2}.\]

Proof. Write \(\rho \coloneq s_1 / s_2\). Since \[w_1^* = \frac{1}{1-\rho^{-2/k}} = -\frac{\rho^{2/k}}{1-\rho^{2/k}}, \qquad w_2^* = 1-w_1^* = \frac{1}{1-\rho^{2/k}},\] we have \[a_{\mathfrak S} = -\frac{\rho^{1+2/k}}{1-\rho^{2/k}}, \qquad b_{\mathfrak S} = \frac{1}{1-\rho^{2/k}}.\] Under 7, both coefficients are uniformly bounded and \(b_{\mathfrak S} \geq 1\).

The selector-profile moments from 16 give \[\begin{align} \Ex*{q_{s_1}(X_1)^2} & = \frac{1}{2s_1-1}, \\ \Ex*{q_{s_2}(X_1)^2} & = \frac{1}{2s_2-1}, \\ \Ex*{q_{s_1}(X_1) q_{s_2}(X_1)} & = \frac{1}{s_1+s_2-1}. \end{align}\] The following matrix records the \(L^2\) geometry of the two selector profiles \(q_{s_1}\) and \(q_{s_2}\). Its determinant lower bound says that, under the scale-separation condition, these profiles do not become asymptotically collinear, so the signed TDNN coefficients cannot cancel the stochastic first-projection signal. Let \[\Gamma_{\mathfrak S} \coloneq s_2 \begin{pmatrix} \Ex*{q_{s_1}(X_1)^2} & \Ex*{q_{s_1}(X_1) q_{s_2}(X_1)} \\ \Ex*{q_{s_1}(X_1) q_{s_2}(X_1)} & \Ex*{q_{s_2}(X_1)^2} \end{pmatrix}.\] Then \[s_2 \Ex*{\tilde{q}_{\mathfrak S}(X_1)^2} = \begin{pmatrix} a_{\mathfrak S} & b_{\mathfrak S} \end{pmatrix} \Gamma_{\mathfrak S} \begin{pmatrix} a_{\mathfrak S} \\ b_{\mathfrak S} \end{pmatrix}.\] Moreover, \[\det\left(\Gamma_{\mathfrak S}\right) = \frac{s_2^2 (s_2-s_1)^2}{(2s_1-1)(2s_2-1)(s_1+s_2-1)^2}.\] Using 7, \[s_1 \geq \mathfrak c s_2, \qquad s_2-s_1 \geq \mathfrak c s_2, \qquad s_1 \leq (1-\mathfrak c)s_2,\] so \[\det\left(\Gamma_{\mathfrak S}\right) \geq \frac{\mathfrak c^2}{16(1-\mathfrak c)}.\] Since \[\operatorname{tr}\left(\Gamma_{\mathfrak S}\right) = \frac{s_2}{2s_1-1} + \frac{s_2}{2s_2-1} \leq \frac{1}{\mathfrak c} + 1,\] the eigenvalues of \(\Gamma_{\mathfrak S}\) are uniformly bounded above and away from zero. In particular, there exist constants \(0 < \underline\lambda \leq \overline{\lambda} < \infty\) such that \[\underline\lambda I_2 \preceq \Gamma_{\mathfrak S} \preceq \overline{\lambda} I_2.\] Because \(a_{\mathfrak S}\) and \(b_{\mathfrak S}\) are uniformly bounded and \(b_{\mathfrak S} \geq 1\), there exists \(C_v < \infty\) such that \[1 \leq a_{\mathfrak S}^2 + b_{\mathfrak S}^2 \leq C_v.\] Combining the last two displays gives \[\underline\lambda \leq s_2 \Ex*{\tilde{q}_{\mathfrak S}(X_1)^2} \leq \overline{\lambda} C_v,\] which is exactly the claimed bound. ◻

By 36 and 17, \[\Ex*{b_s(X_1)^2} \leq \zeta_s^1(x) \lesssim s^{-1}.\] Since 7 keeps \(a_{\mathfrak S}\) and \(b_{\mathfrak S}\) uniformly bounded and \(s_1 \asymp s_2\), \[\Ex*{\tilde{b}_{\mathfrak S}(X_1)^2} \lesssim \Ex*{b_{s_1}(X_1)^2} + \Ex*{b_{s_2}(X_1)^2} \lesssim s_2^{-1}.\] Because 5 [asm:tdnn95design95variance] places \(\sigma_\varepsilon^2(\cdot)\) on the compact support \(\mathcal{X}\) as a strictly positive continuous function, there exist constants \[0 < \underline{\sigma}_\varepsilon^2 \leq \sigma_\varepsilon^2(u) \leq \overline{\sigma}_\varepsilon^2 < \infty, \qquad u \in \mathcal{X}.\] Combining these bounds with 38 and 19, we obtain \[\label{eq:tdnn95first95projection95target95rate} \zeta_{\mathfrak S}^{1}(x) \asymp s_2^{-1}.\tag{39}\]

8.2.0.4 Step 3A: Verify the row-wise \(L^r\) condition.

With 39 established, the TDNN first projection satisfies the row-wise square-LLN by the same moment calculation used for the single-scale DNN first projection.

Lemma 20 (TDNN verification of the row-wise \(L^r\) condition).
* Consider a data-generating process as outlined in 4, 5, and 8, and suppose 7 holds. Then the TDNN first projection satisfies 2. More precisely, with \(r_\eta \mathrel{\vcenter{:}}= 1 + \frac{1}{2}\min\{\eta,2\}\), where \(\eta\) is from 8, \[\Ex*{ \left( \frac{h_{\mathfrak S}^{(1)}(x; Z_1)^2}{\zeta_{\mathfrak S}^{1}(x)} \right)^{r_\eta} } \lesssim s_2^{r_\eta-1},\] so the ratio in 2 vanishes whenever \(s_2=o(n)\).

Proof. Write \(r=r_\eta\) for the exponent fixed in the statement. By the exact identity 34 and Minkowski’s inequality in \(L^{2r}\), \[\begin{align} \left( \Ex*{\abs*{h_{\mathfrak S}^{(1)}(x; Z_1)}^{2r}} \right)^{1/(2r)} & \leq \abs*{w_1^*}\frac{s_1}{s_2} \left( \Ex*{\abs*{h_{s_1}^{(1)}(x; Z_1)}^{2r}} \right)^{1/(2r)} \\ & \quad + \abs*{w_2^*} \left( \Ex*{\abs*{h_{s_2}^{(1)}(x; Z_1)}^{2r}} \right)^{1/(2r)}. \end{align}\] The bounded-ratio condition keeps \(w_1^*\) and \(w_2^*\) uniformly bounded, and 26 gives \[\Ex*{\abs*{h_{s_k}^{(1)}(x; Z_1)}^{2r}} \lesssim s_k^{-1}, \qquad k \in \{1,2\}.\] Since \(s_1 \leq s_2\), \[\left( \frac{s_1}{s_2} \right) s_1^{-1/(2r)} \leq s_2^{-1/(2r)}.\] Therefore, \[\Ex*{\abs*{h_{\mathfrak S}^{(1)}(x; Z_1)}^{2r}} \lesssim s_2^{-1}.\] Dividing by 39 yields \[\Ex*{ \left( \frac{h_{\mathfrak S}^{(1)}(x; Z_1)^2}{\zeta_{\mathfrak S}^{1}(x)} \right)^r } = \frac{\Ex*{\abs*{h_{\mathfrak S}^{(1)}(x; Z_1)}^{2r}}}{\left(\zeta_{\mathfrak S}^{1}(x)\right)^r} \lesssim s_2^{r-1}.\] Since \(r>1\) and \(s_2=o(n)\), \[\frac{1}{n^{r-1}} \Ex*{ \left( \frac{h_{\mathfrak S}^{(1)}(x; Z_1)^2}{\zeta_{\mathfrak S}^{1}(x)} \right)^r } \lesssim \left(\frac{s_2}{n}\right)^{r-1} \longrightarrow 0.\] This proves 2. ◻

8.2.0.5 Step 4: Conclude Hájek dominance.

Combining 31 39 , the TDNN kernel satisfies 1 because \[\frac{s_2}{n} \left( \frac{\zeta_{\mathfrak S}^{s_2}(x)}{s_2 \zeta_{\mathfrak S}^{1}(x)} - 1 \right) \longrightarrow 0\] whenever \(s_2=o(n)\). This follows from the same variance scaling as in the single-scale DNN case. The full kernel remains of constant order, while the first projection is diluted by competition among \(s_2\) candidate observations.

8.2.0.6 Summary.

At the level of ideas, the TDNN bound uses three ingredients beyond the single-scale DNN case. The first is the Jensen bound for the embedded \(s_1\)-scale average. The second is the projection identity 33 , which turns the first-order TDNN analysis into a scaled combination of single-scale DNN first projections. The third is the exact stochastic-coefficient decomposition 37 together with the uniform non-cancellation bound in 19. Together, these ingredients reduce the TDNN analysis to the single-scale DNN bounds plus the selector non-cancellation step.

Remark 9 (Multiscale extension). The preceding argument should extend to a \(K\)-scale bias-corrected combination with kernel orders \(1 \leq s_1 < \dotsb < s_K\). The open ingredient is a non-cancellation lower bound for the effective localizer \(\tau_{\mathbf{S},x}^{\mathrm{eff}}(u) \coloneq \sum_m a_m (s_m/s_K) \tau_{s_m,x}(u)\): one needs to show that the bias-correction coefficients \(a_m\), while deliberately chosen to cancel deterministic bias terms, do not simultaneously annihilate the first-order stochastic signal. This is the only step without a direct analogue in the two-scale case and would require a separate combinatorial argument. The remaining ingredients extend directly: the full-kernel second moment inherits the \(O(1)\) bound from the Jensen argument applied scale by scale; and the projection identity \[\bar h_{s_m \mid s_K}^{(1)}(x; z_1) = \frac{s_m}{s_K} h_{s_m}^{(1)}(x; z_1)\] follows from the embedded-average structure, giving each active scale a \(s_K^{-1}\) first-projection variance contribution. If the non-cancellation statement can be established, Hájek dominance at rate \(s_K = o(n)\) would follow without further structural changes to the proof.

References↩︎

[1]
J. N. Arvesen, “Jackknifing U-Statistics,” The Annals of Mathematical Statistics, vol. 40, no. 6, pp. 2076–2100, Dec. 1969, doi: 10.1214/aoms/1177697287.
[2]
L. A. Jaeckel, Dated June 30, 1972“The Infinitesimal Jackknife,” Bell Laboratories, Murray Hill, NJ, Memorandum for File MM 72-1215-11, 1972.
[3]
B. Efron, “Jackknife-After-Bootstrap Standard Errors and Influence Functions,” Journal of the Royal Statistical Society Series B: Statistical Methodology, vol. 54, no. 1, pp. 83–111, Sep. 1992, doi: 10.1111/j.2517-6161.1992.tb01866.x.
[4]
E. Demirkaya, Y. Fan, L. Gao, J. Lv, P. Vossler, and J. Wang, “Optimal Nonparametric Inference with Two-Scale Distributional Nearest Neighbors,” Journal of the American Statistical Association, vol. 119, no. 545, pp. 297–307, Jan. 2024, doi: 10.1080/01621459.2022.2115375.
[5]
W. Hoeffding, “A Class of Statistics with Asymptotically Normal Distribution,” The Annals of Mathematical Statistics, vol. 19, no. 3, pp. 293–325, Sep. 1948, doi: 10.1214/aoms/1177730196.
[6]
A. J. Lee, U-Statistics, 0th ed. Routledge, 2019.
[7]
B. Efron and C. Stein, “The Jackknife Estimate of Variance,” The Annals of Statistics, vol. 9, no. 3, May 1981, doi: 10.1214/aos/1176345462.
[8]
J. Shao and C. F. J. Wu, “A General Theory for Jackknife Variance Estimation,” The Annals of Statistics, vol. 17, no. 3, Sep. 1989, doi: 10.1214/aos/1176347263.
[9]
M. A. Arcones and E. Gine, “On the Bootstrap of U and V Statistics,” The Annals of Statistics, vol. 20, no. 2, Jun. 1992, doi: 10.1214/aos/1176348650.
[10]
W. R. Schucany and D. M. Bankson, _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1467-842X.1989.tb00986.x“Small Sample Variance Estimators for U-Statistics,” Australian Journal of Statistics, vol. 31, no. 3, pp. 417–426, 1989, doi: 10.1111/j.1467-842X.1989.tb00986.x.
[11]
E. W. Frees, “Infinite Order U-Statistics,” Scandinavian Journal of Statistics, vol. 16, no. 1, pp. 29–45, 1989, Accessed: Mar. 31, 2026. [Online]. Available: https://www.jstor.org/stable/4616120.
[12]
L. Mentch and G. Hooker, “Quantifying Uncertainty in Random Forests via Confidence Intervals and Hypothesis Tests,” Journal of Machine Learning Research, vol. 17, no. 26, pp. 1–41, 2016, [Online]. Available: https://jmlr.org/papers/v17/14-168.html.
[13]
Z. Zhou, L. Mentch, and G. Hooker, “V-statistics and Variance Estimation,” Journal of Machine Learning Research, vol. 22, no. 287, pp. 1–48, 2021, Accessed: Sep. 03, 2025. [Online]. Available: https://jmlr.org/papers/v22/21-0575.html.
[14]
W. Peng, T. Coleman, and L. Mentch, “Rates of convergence for random forests via generalized U-statistics,” Electronic Journal of Statistics, vol. 16, no. 1, Jan. 2022, doi: 10.1214/21-EJS1958.
[15]
Q. Wang and B. G. Lindsay, “Variance estimation of a general u-statistic with application to cross-validation,” Statistica Sinica, 2014, doi: 10.5705/ss.2012.215.
[16]
Q. Wang and B. Lindsay, “Pseudo-Kernel Method in U-Statistic Variance Estimation with Large Kernel Size,” Statistica Sinica, vol. 27, no. 3, pp. 1155–1174, 2017, Accessed: Mar. 31, 2026. [Online]. Available: https://www.jstor.org/stable/26383248.
[17]
J. Sexton and P. Laake, “Standard errors for bagged and random forest estimators,” Computational Statistics & Data Analysis, vol. 53, no. 3, pp. 801–811, Jan. 2009, doi: 10.1016/j.csda.2008.08.007.
[18]
S. Wager, T. Hastie, and B. Efron, “Confidence Intervals for Random Forests: The Jackknife and the Infinitesimal Jackknife,” Journal of Machine Learning Research, vol. 15, no. 48, pp. 1625–1651, 2014, Accessed: Jan. 19, 2025. [Online]. Available: https://jmlr.org/papers/v15/wager14a.html.
[19]
Q. Wang and Y. Wei, _eprint: https://doi.org/10.1080/00949655.2022.2081969“Quantifying uncertainty of subsampling-based ensemble methods under a U-statistic framework,” Journal of Statistical Computation and Simulation, vol. 92, no. 17, pp. 3706–3726, Nov. 2022, doi: 10.1080/00949655.2022.2081969.
[20]
T. Xu, R. Zhu, and X. Shao, “On variance estimation of random forests with Infinite-order U-statistics,” Electronic Journal of Statistics, vol. 18, no. 1, pp. 2135–2207, Jan. 2024, doi: 10.1214/24-EJS2247.
[21]
W. Peng, L. Mentch, and L. Stefanski, “Bias, Consistency, and Alternative Perspectives of the Infinitesimal Jackknife,” Statistica Sinica, 2026, doi: 10.5705/ss.202022.0131.
[22]
Q. Wang and X. Cai, _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/sta4.70070“A New Perspective on U-Statistic Variance Estimation,” Stat, vol. 14, no. 2, p. e70070, 2025, doi: 10.1002/sta4.70070.
[23]
Y. Lin and Y. Jeon, “Random Forests and Adaptive Nearest Neighbors,” Journal of the American Statistical Association, vol. 101, no. 474, pp. 578–590, Jun. 2006, doi: 10.1198/016214505000001230.
[24]
S. Athey, J. Tibshirani, and S. Wager, “Generalized random forests,” The Annals of Statistics, vol. 47, no. 2, pp. 1148–1178, Apr. 2019, doi: 10.1214/18-AOS1709.
[25]
[26]
B. Von Bahr and C.-G. Esseen, “Inequalities for the $r$th Absolute Moment of a Sum of Random Variables, $1 \leqq r \leqq 2$,” The Annals of Mathematical Statistics, vol. 36, no. 1, pp. 299–303, 1965, doi: 10.1214/aoms/1177700291.
[27]
B. M. Steele, “Exact bootstrap k-nearest neighbor learners,” Machine Learning, vol. 74, no. 3, pp. 235–255, Mar. 2009, doi: 10.1007/s10994-008-5096-0.
[28]
G. Biau and L. Devroye, “On the layered nearest neighbour estimate, the bagged nearest neighbour estimate and the random forest method in regression and classification,” Journal of Multivariate Analysis, vol. 101, no. 10, pp. 2499–2518, Nov. 2010, doi: 10.1016/j.jmva.2010.06.019.