Private Rate-Double-Robust Inference


We reconcile privacy protection and rate-double-robust inference. The privacy of individuals is protected by a local privacy mechanism: injecting noise into their sensitive data, revealing only the noisy data for inference. Hence, privacy protection hinders inference. In contrast, the inference of a target parameter is rate-double-robust when the large-sample bias of an estimator of the parameter is characterised by a trade-off between the estimation errors of two other, nuisance, parameters. Hence, rate-double-robustness facilitates inference. Our starting point of reconciliation is a class of rate-double-robust target parameters indexed linearly by an infinite-dimensional and nonlinearly by a low-dimensional regression. Among others, this includes causal parameters. To infer these targets privately, we show how suitable privacy mechanisms transfer the semiparametric properties of the sensitive-data model to the private setting. Rate-double-robustness is transferred, enabling locally-private, unbiased and semiparametrically efficient inference of our target parameters. Finally, we transform general nonparametric nuisance estimators into private ones, which inherit convergence properties of their nonprivate counterparts. For parametric nuisance models, we develop a private method-of-moments estimator and its large-sample inference theory.

1 Introduction↩︎

Sensitive data of units in a sample are desirable to protect. This may be accomplished by a privacy mechanism, which disguises the sensitive data by deliberately injecting noise into them. Next, only the noisy — and not the sensitive — data are revealed, preserving the privacy of sampled units, but hindering inference.

This paper is concerned with the inference of a parameter \(\chi(P_{VX})\) of the distribution \(P_{VX}\) of the data \((V,X)\) when \(X\) is privacy-protected. Hence, we wish to infer \(\chi(P_{VX})\) from the data \((V,Z)\), where \(Z\) is the noisy version of \(X\) disguised by a given privacy mechanism.

We guarantee privacy by local mechanisms: the noise is injected to the sensitive data of each unit. Then we consider inference under a fixed mechanism, employed uniformly for each unit. While some of our results hold for general local mechanisms, specialising the mechanisms yields more interesting results. Our specialised mechanisms leave the sensitive data intact with probability \(\alpha\), and output pure noise with probability \(1-\alpha\), guaranteeing total-variation privacy [1]. It is this \(\alpha\)-identity which enables inference. Under this specialisation, \(X\) can take values in any measurable space, such as metric spaces. Thus, these mechanisms are much more flexible than those using additive noise [2].

We focus on the inference of parameters \(\chi(P_{VX})\) with a rate-double-robustness (or mixed-bias) property. A parameter has this property if the large-sample bias of an estimator thereof is characterised by the product of the estimation errors of two other parameters, which are then called nuisance parameters. The product is attractive as the errors can compensate each other. This is favourable for infinite-dimensional nuisance parameters with large estimation errors. If the product vanishes, \(\chi(P_{VX})\) can be inferred unbiasedly in the large-sample limit. Examples of rate-double-robust parameters include average treatment effects.

Our contribution is threefold. First, we propose a novel class of rate-double-robust parameters in the nonprivate setting, motivated by [3] and [4]. The class comprises parameters which depend linearly on an infinite-dimensional regression and nonlinearly on a low-dimensional regression. While [3] consider dependence parameters more general than regressions, they only allow for linear dependencies.1 While [4] show that the average treatment effect on the treated — which falls into our class — is rate-double-robust, they do not generalise this result to nonlinear dependencies on low-dimensional regressions. Generalisation is straightforward as low-dimensional parameters are estimable at a fast rate.

Second, turning to the private setting, we provide conditions for the privacy mechanism to infer \(\chi(P_{VX})\) from the observed noisy data \((V,Z)\sim P_{VZ}\), and show how our specialised mechanisms satisfy them. Namely, we connect the semiparametric properties of the statistical models for \(P_{VX}\) and \(P_{VZ}\). Under our specialised mechanisms, if a parameter is rate-double-robust in the nonprivate setting, then so it remains in the private setting. This leads to privacy-protected large-sample unbiased inference of \(\chi(P_{VX})\) from \((V,Z)\), for fast-enough nuisance estimators. The limiting variance increases with the noise level of the privacy mechanism, but it is semiparametrically efficient in the private model induced by a nonparametric model for \(P_{VX}\) and by the specialised mechanism.

Third, we study private estimation of the nuisance parameters. Expressing them as expected-loss minimisers in the private setting paves the way for estimation through empirical risk minimisation. Alternatively, given a nonprivate “source” estimator in a general class of nonparametric estimators, we transform it to an estimator which uses only the noisy data \((V,Z)\). We show how the transformed estimator inherit the guarantees of its nonprivate source. For example, the convergence rates of privatised kernel and orthogonal series estimators remain the same as those of their nonprivate counterparts inflated by the noise level of the privacy mechanism. For parametric nuisance models, we develop a private method-of-moments estimator for \(\mathbb{R}^K\)-valued parameters identified from moment conditions [5], [6]. We also derive its limiting distribution — an apparently new result in private inference.

In summary, to the best of our knowledge, our work is the first achieving locally private, efficient and unbiased rate-double-robust inference for parameters as general as the ones in our proposed class, and with data taking values in generic spaces.

In 2, we situate our work in the literature. In 3, we introduce our rate-double-robust class without privacy. 4 adds privacy. 5 discusses private estimation of the double-robust parameters. 6 focuses on the private estimation of nuisance parameters. 7 concludes.

2 Literature↩︎

Privacy-preserving inference [2], [7], [8] can offer central or local privacy guarantees; see [9] for a survey.2 Central mechanisms inject noise into sample aggregates, hence are less stringent than the local ones adopted by us, which noise individual data.

In the central paradigm, parametric (regression) models are studied by [13], [14], [15], [16], and the nonparametric median by [17]. In the local paradigm, convergence rates [18][20] and Fisher-information bounds [21] are derived.

Causal parameters are important instances in our class. Their private inference is addressed by [22], [23], [24], and [25]. Compared to them, we support more general private covariate adjustment or parameters, under less stringent assumptions about \(X\).

Private efficiency theory, also optimising over the privacy mechanism, is pioneered by [26] for parametric models. We only consider efficient inference for a given mechanism, but we adopt nonparametric models like [27] and [28] do. They provide minimax rates up to constants, whereas our results are asymptotically exact. [29] considers minimax rates in nonparametric models, specialised to the expected density; we infer parameters in a broad class.

Private M-estimators have been well studied for parametric models [30][36]. Yet, only [10] and [37] appear to derive limiting distributions. Compared to the former, we offer more stringent, local privacy; unlike the latter, our method-of-moments estimator is asymptotically unbiased.

3 Rate-Double-Robust Inference↩︎

In this section, we construct our target parameters \(\chi(P_{VX})\) and derive their inferential properties without privacy. This serves as a basis for private inference in 4. Thus, for now, we consider the nonprivate setting when the sensitive data \((V,X)\) are observable.

Here, \((V,X)\) is a random element defined on the probability space \((\Omega, \mathscr{F}_\Omega, \mathbb{P})\), with distribution \(P_{VX}\) belonging to \(\mathcal{P}_{VX}\subseteq \mathcal{P}_{\mathfrak{V}\mathfrak{X}}\), where \(\mathcal{P}_{\mathfrak{V}\mathfrak{X}}\) is the set of all possible distributions on the measurable space \((\mathfrak{V}\times\mathfrak{X},\mathscr{F}_{\mathfrak{V}\times\mathfrak{X}})\), where \(\mathfrak{V}=\mathfrak{V}_\mathfrak{1}\times\mathfrak{V}_\mathfrak{2}\times\ldots\). Our primary interest is in the nonparametric model \[\begin{align} \mathcal{P}_{VX}=\mathcal{P}_{\mathfrak{V}\mathfrak{X}}.\label{priv:eq:nonparamodel95set} \end{align}\tag{1}\]

Each functional \(\chi:\mathcal{P}_{VX}\to\mathbb{R}\) yields a parameter \(\chi(P_{VX})\). 3.1 introduces our class of \(\chi\) which yield rate-double-robust parameters; 3.2 studies their inferential properties.

Notation.For \(h:\mathfrak{V}\times\mathfrak{X}\to\mathbb{R}\), we let \(P_{VX}h\mathrel{\vcenter{:}}= P_{VX}h(V,X)\mathrel{\vcenter{:}}=\int_{\mathfrak{V}\times\mathfrak{X}}h(v,x)\,\mathrm{d}P_{VX}(v,x)\). Fix \(p\in[1,\infty)\). With \(\left\lVert{h}\right\rVert_{L_p(P_{VX})}\mathrel{\vcenter{:}}=\left(P_{VX}|h|^p\right)^{\frac{1}{p}}\), write \(L_p(P_{VX})\) for all \(h:\mathfrak{V}\times\mathfrak{X}\to\mathbb{R}\) with \(\left\lVert{h}\right\rVert_{L_p(P_{VX})}^p<\infty\), and \(L_p^0(P_{VX})\) for all \(h\in L_p(P_{VX})\) with \(P_{VX}h=0\). Let \(\left\lVert{h}\right\rVert_{\infty}\mathrel{\vcenter{:}}=\sup_{(v,x)\in\mathfrak{V}\times\mathfrak{X}}|h(v,x)|\), and \(\rho((h,a),(h',a'))\mathrel{\vcenter{:}}=\left\lVert{h-h'}\right\rVert_{L_2(P_{VX})}+|a-a'|\) be a metric on \(L_2(P_{VX})\times\mathbb{R}\ni(h,a)\). Let \(\delta_{(v,x)}\) be the Dirac measure at \((v,x)\in\mathfrak{V}\times\mathfrak{X}\). We call a \((\pi_{1}\circ V,\pi_{2}\circ V,\ldots)\) a collection of coordinates of \(V\), if all the \(\pi_j\circ v\in \mathfrak{V}_\mathfrak{j'}\) for all \(v\in\mathfrak{V}\) for some \(\mathfrak{V}_\mathfrak{j'}\in \{\mathfrak{V}_\mathfrak{1},\mathfrak{V}_\mathfrak{2},\ldots\}\); for example, if \(\mathfrak{V}=\mathfrak{V}_\mathfrak{1}\times\mathfrak{V}_\mathfrak{2}\times\mathfrak{V}_\mathfrak{3}\) with corresponding \(V=(V_1,V_2,V_3)\), then \((V_2,V_1)\) is a collection of coordinates of \(V\).

3.1 Rate-Double-Robust Parameter Class↩︎

Let \(V_1\) and \(V_2\) be two arbitrary collections of coordinates of \(V\) with values in \(\mathfrak{V}_{\mathit{1}}\) and \(\mathfrak{V}_{\mathit{2}}\), respectively, where, importantly, \(\mathfrak{V}_{\mathit{2}}\) is finite. For given \(m,g:\mathfrak{V}\times\mathfrak{X}\to\mathbb{R}\), define the regressions \[\begin{align} \label{priv:eq:regressions95def}\begin{aligned} \mu_{\mathcal{X}}(v_{\mathit{1}},x)&\mathrel{\vcenter{:}}=\mathbb{E}\left[\left. m(V,X)\,\right\vert\, V_{\mathit{1}}=v_{\mathit{1}},X=x\right],\quad (v_{\mathit{1}},x)\in \mathfrak{V}_{\mathit{1}}\times\mathfrak{X}, \\ \gamma_{\mathcal{V}}(v_{\mathit{2}})&\mathrel{\vcenter{:}}=\mathbb{E}\left[\left. g(V,X)\,\right\vert\, V_{\mathit{2}}=v_{\mathit{2}}\right],\quad v_{\mathit{2}}\in\mathfrak{V}_{\mathit{2}}, \end{aligned} \end{align}\tag{2}\] assuming \(\mu_{\mathcal{X}}\in L_2(P_{V_{\mathit{1}}X}),\gamma_{\mathcal{V}}\in L_2(P_{V_{\mathit{2}}})\). For a given \(f:\mathfrak{V}\times\mathfrak{X}\times L_2(P_{V_{\mathit{1}}X})\times L_2(P_{V_{\mathit{2}}})\to\mathbb{R}\), our targets are \(\chi(P_{VX})=\mathbb{E}f(V,X, \mu_{\mathcal{X}}, \gamma_{\mathcal{V}})\). As \(\mathfrak{V}_{\mathit{2}}\) is finite, we can define without loss of generality our target parameter as \[\begin{align} \chi(P_{VX})&\mathrel{\vcenter{:}}=\mathbb{E}f(V,X, \mu_{\mathcal{X}}, \gamma_{\mathcal{V}}(c)) \label{priv:eq:dr95para}\end{align}\tag{3}\] for a fixed \(c\in\mathfrak{V}_{\mathit{2}}\), and \(f:\mathfrak{V}\times\mathfrak{X}\times L_2(P_{V_{\mathit{1}}X})\times \Gamma \to\mathbb{R}\), for \(\Gamma\supseteq g(\mathfrak{X},\mathfrak{V})\). We constrain \(f\), requiring that \[\begin{align} L_2(P_{V_{\mathit{1}}X})\ni\mu \mapsto \mathbb{E}f(V,X,\mu,\gamma) &\text{ be \left\lVert{\cdot}\right\rVert_{L_2(P_{VX})}-continuous for all }\gamma\in\Gamma, \tag{4} \\ L_2(P_{V_{\mathit{1}}X})\ni\mu\mapsto f(V,X,\mu,\gamma)& \text{ be linear } P_{VX}\text{-a.s. for all } \gamma\in\Gamma,\tag{5} \\ \Gamma \ni\gamma\mapsto f(V,X,\mu,\gamma)& \text{ be twice continuously differentiable }P_{VX}\text{-a.s. for all }\nonumber \\ &\text{ } \mu\in L_2(P_{V_{\mathit{1}}X})\text{ with } \text{derivatives \partial_\gamma f, \partial_\gamma^2 f.} \tag{6}\end{align}\] Conditions 4 , 5 , and 6 restrict the structure of the dependencies on the regressions to enforce rate-double-robustness.

Further, we impose the integrability conditions \[\begin{align} \label{priv:eq:integrability}\begin{aligned} \mathbb{E}f(V,X, \mu_{\mathcal{X}}, \gamma_{\mathcal{V}}(c))^2 <\infty, \quad \mathbb{E}\mathbb{1}_{V_{\mathit{2}}=c}g(V,X)^2<\infty, \\ \mathbb{E}\left[\left. m(V,X)^2\,\right\vert\, V_{\mathit{1}},X\right]<\infty\quad \text{ P_{V_{\mathit{1}}X}-a.s..} \end{aligned} \end{align}\tag{7}\]

Different choices of \((V_{\mathit{1}},V_{\mathit{2}},m,g,f)\) satisfying the above conditions give rise to a class of parameters of the form 3 . In 9.1, we present examples such as average treatment effects ([3]; [4]), “geometric parameters,” and parameters from economics. We also demonstrate how dependence on multiple regressions can be accommodated. In 3.2, we show how parameters in this class lend themselves to rate-double-robust inference.

3.2 Inferential Properties↩︎

Consider the one-step estimator [38][40] of \(\chi(P_{VX})\). Starting from an arbitrary “plug-in” estimator \(\chi(\hat{P}_{VX})\) constructed from a random sample \(\mathcal{S}\mathrel{\vcenter{:}}=((V_i,X_i))_{i\in[n]}\) from \(P_{VX}\), the one-step estimator \[\begin{align} \hat{\chi}_n\mathrel{\vcenter{:}}=\chi(\hat{P}_{VX})+\mathbb{P}_n\hat{\tilde{\chi}} = \chi(\hat{P}_{VX})+\frac{1}{n}\sum_{i\in[n]}\hat{\tilde{\chi}}(V_i,X_i) \label{priv:eq:bias95corrected95chihat} \end{align}\tag{8}\] corrects for the plug-in bias via a directional-derivative expansion of the functional \(\chi\) [39]. This derivate is representable with the so-called efficient influence function \(\tilde{\chi}\) of \(\chi(P_{VX})\) for the model \(\mathcal{P}_{VX}\). Hence, informally, \[\begin{align} \chi(\hat{P}_{VX})-\chi(P_{VX})\approx -P_{VX}\hat{\tilde{\chi}}, \label{priv:eq:vonmises95approx} \end{align}\tag{9}\] expanding by the estimate \(\hat{\tilde{\chi}}\) of \(\tilde{\chi}\) in the direction \(P_{VX}-\hat{P}_{VX}\). This motivates 8 with the unknown \(P_{VX}\hat{\tilde{\chi}}\) estimated with \(\mathbb{P}_n\hat{\tilde{\chi}}\).

To construct \(\hat{\tilde{\chi}}\) in 8 , we need \(\tilde{\chi}\). Let \(r\in L_2(P_{V_{\mathit{1}}X})\) be the function satisfying \[\begin{align} \mathbb{E}f(V,X, \mu, \gamma_{\mathcal{V}}(c))=\mathbb{E}r(V_{\mathit{1}},X)\mu(V_{\mathit{1}},X)\quad \text{ for all } \mu\in L_2(P_{V_{\mathit{1}}X}). \label{priv:eq:riesz95rep} \end{align}\tag{10}\] The existence and uniqueness of \(r\) follows from the Riesz representation theorem by 4 and 5 , whence \(r\) is called the Riesz representer. The representer \(r\), whose dependence on \((P_{VX},f,c)\) is silent in our notation, is obtained by manipulating the left-hand side of 10 until the right-hand side is reached (see 9.1). With \(r\), \(\tilde{\chi}\) is derived in 1. The dependence on the regressions manifests itself in ?? ; under no dependence, the (generalised) derivates \(r,\partial_\gamma f\) are zero.

Proposition 1 (Efficient Influence Function of \(\chi(P_{VX})\)). In the nonparametric model 1 for \(P_{VX}\), the efficient influence function \(\tilde{\chi}: \mathfrak{V}\times\mathfrak{X}\to\mathbb{R}\) of \(\chi(P_{VX})\) in 3 is, at \(P_{VX}\), \[\begin{align} \label{priv:eq:eif95chi} \begin{aligned} \tilde{\chi}(v,x)\mathrel{\vcenter{:}}=&\, r(v_{\mathit{1}},x)(m(v,x)-\mu_{\mathcal{X}}(v_{\mathit{1}},x))\\ &+\frac{\mathbb{1}_{v_{\mathit{2}}=c}}{p_{V_{\mathit{2}}}(c)}(g(v,x)-\gamma_{\mathcal{V}}(c))\mathbb{E}\partial_\gamma f(V,X, \mu_{\mathcal{X}}, \gamma_{\mathcal{V}}(c)) \\ &+ f(v,x, \mu_{\mathcal{X}}, \gamma_{\mathcal{V}}(c))-\chi(P_{VX}), \quad (v,x)\in\mathfrak{X}\times\mathfrak{V}, \end{aligned} \end{align}\qquad{(1)}\] where we denote by \(v_{\mathit{1}},v_{\mathit{2}}\) the coordinates of \(v\) that correspond to \(V_{\mathit{1}},V_{\mathit{2}}\) of \(V\). When \(V_\mathit{2}=\varnothing\), it is understood that \(\frac{\mathbb{1}_{v_{\mathit{2}}=c}}{p_{V_{\mathit{2}}}(c)}(g(v,x)-\gamma_{\mathcal{V}}(c))=g(v,x)-\mathbb{E}g(V,X)\).

Proof. All proofs are in the appendix. ◻

Now we verify that \(\chi(P_{VX})\) in 3 is rate-double-robust. 1 below controls the approximation error in 9 . Because \(\mathfrak{V}_{\mathit{2}}\) is finite, the first, product term dominates in ?? for reasonable “estimators” \(p_{V_{\mathit{2}}}'(c),\gamma_{\mathcal{V}}'(c)\). This proves the one-step estimator 8 of \(\chi(P_{VX})\) rate-double-robust with bias characterised by the product of estimation errors of the nuisance parameters \(\mu_{\mathcal{X}},r\).

Theorem 1 (Rate-Double-Robustness). Let \(\mu_{\mathcal{X}}',r'\in L_2(P_{V_{\mathit{1}}X})\) and \(p_{V_{\mathit{2}}}'(c),\gamma_{\mathcal{V}}'(c),\chi',e''\in\mathbb{R}\) all be arbitrary. Set \[\begin{align} \label{priv:eq:eif95chi95prime} \begin{aligned} \tilde{\chi}'(v,x)\mathrel{\vcenter{:}}=&\, r'(v_{\mathit{1}},x)(m(v,x)-\mu_{\mathcal{X}}'(v_{\mathit{1}},x))+\frac{\mathbb{1}_{v_{\mathit{2}}=c}}{p_{V_{\mathit{2}}}'(c)}(g(v,x)-\gamma_{\mathcal{V}}'(c))e'' \\ &+ f(v,x, \mu_{\mathcal{X}}', \gamma_{\mathcal{V}}'(c))-\chi'. \end{aligned} \end{align}\qquad{(2)}\] Then \[\begin{align} \label{priv:eq:chihat95biasdecomp} \begin{aligned} \chi'-\chi(P_{VX})+P_{VX}\tilde{\chi}' =&\, -P_{VX}(r-r')(\mu_{\mathcal{X}}-\mu_{\mathcal{X}}') \\ &+(\gamma_{\mathcal{V}}(c)-\gamma_{\mathcal{V}}'(c))\left(\frac{p_{V_{\mathit{2}}}(c)}{p_{V_{\mathit{2}}}'(c)}e''-e'\right) \\ &-(\gamma_{\mathcal{V}}(c)-\gamma_{\mathcal{V}}'(c))^2\frac{P_{VX}\partial_\gamma^2 f(V,X,\mu_{\mathcal{X}}',\widetilde{\gamma_{\mathcal{V}}(c)})}{2} \end{aligned} \end{align}\qquad{(3)}\] for some \(\widetilde{\gamma_{\mathcal{V}}(c)}\) between \(\gamma_{\mathcal{V}}(c)\) and \(\gamma_{\mathcal{V}}'(c)\), and \(e'\mathrel{\vcenter{:}}= P_{VX}\partial_\gamma f(V,X,\mu_{\mathcal{X}}',\gamma_{\mathcal{V}}'(c))\).

The function \(\tilde{\chi}\) is a key object in semiparametric efficiency theory. By definition, it is an element of the \(L_2(P_{VX})\)-completion \(\overline{\mathrm{lin}\,}\mathcal{T}_{VX}\) of the linear span of the tangent set of the model \(\mathcal{P}_{VX}\) at \(P_{VX}\), \(\mathcal{T}_{VX}\): the set of all directions in which \(P_{VX}\) can be perturbed so that the perturbed \(P_{VX}\) remains in \(\mathcal{P}_{VX}\). For the nonparametric model \(\mathcal{P}_{\mathfrak{V}\mathfrak{X}}\) in 1 , \(\overline{\mathrm{lin}\,}\mathcal{T}_{VX}=\mathcal{T}_{VX}=L_2^0(P_{VX})\). Among all elements of \(\overline{\mathrm{lin}\,}\mathcal{T}_{VX}\), it is \(\tilde{\chi}\) whose squared norm \(P_{VX}\tilde{\chi}^2\) is the limiting variance of asymptotically efficient estimators of \(\chi(P_{VX})\) ([38]). In the nonprivate setting, we show that 8 is asymptotically efficient (12). In the following, we prove an analogue in the private setting.

4 Privacy↩︎

To preserve the privacy of each unit in the sample \(\mathcal{S}=((V_i,X_i))_{i\in[n]}\), a noisy version \(Z_i\) of \(X_i\) is generated, and only \(\bar{\mathcal{S}}\mathrel{\vcenter{:}}=((V_i,Z_i))_{i\in[n]}\) is revealed to infer \(\chi(P_{VX})\). In 4.1, we describe how \(Z_i\) is generated to enable inference from \(\bar{\mathcal{S}}\), which is discussed in 4.2.

4.1 Privacy Mechanism↩︎

The \(Z_i\) are generated from \(X_i\) via a privacy mechanism, which can be thought of as a random map. Given a measurable space \((\mathfrak{Z},\mathscr{F}_{\mathfrak{Z}})\), the \(Z_i\) are generated as random draws \(Z_i{\vert\,}((V_j,X_j))_{j\in[n]}\sim Q(\cdot{\vert\,}X_i)\) for all \(i\in[n]\) for a \(Q\in\mathcal{Q}(\mathfrak{X}\to\mathfrak{Z})\), where \(\mathcal{Q}(\mathfrak{X}\to\mathfrak{Z})\) is the set of all Markov kernels \(Q:\mathscr{F}_{\mathfrak{Z}}\times\mathfrak{X}\to[0,1]\), so that \(B\mapsto Q(B{\vert\,}x)\) is a probability measure for all \(x\in\mathfrak{X}\), and \(x\mapsto Q(B{\vert\,}x)\) is measurable for every \(B\in\mathscr{F}_{\mathfrak{Z}}\). The kernel \(Q\) is called a local noninteractive privacy mechanism: local, because it noises the data of each unit \(i\), and noninteractive, because it does not use the data of units \(j\neq i\) to generate \(Z_i\) [26]. Therefore, \(((V_i,Z_i))_{i\in[n]}\) is a random sample from the distribution of \((V,Z)\), the mixture \[\begin{align} P_{VZ}(B_{\mathsf{v}}, B_{\mathsf{z}})= \int_{B_{\mathsf{v}}}\int_{\mathfrak{X}}Q(B_{\mathsf{z}}{\vert\,}x) \,\mathrm{d}P_{VX}(v,x),\quad B_{\mathsf{v}}\in\mathscr{F}_{\mathfrak{V}}, B_{\mathsf{z}}\in\mathscr{F}_{\mathfrak{Z}}. \label{priv:eq:distribution95vz} \end{align}\tag{11}\] The distribution \(P_{VZ}\), determined by \(Q\) and \(P_{VX}\), thus belongs to the model \[\begin{align} \mathcal{P}_{VZ}(Q,\mathcal{P})\mathrel{\vcenter{:}}=\left\{P\in\mathcal{P}_{\mathfrak{V}\mathfrak{Z}}: P(B_{\mathsf{v}}, B_{\mathsf{z}})= \int_{B_{\mathsf{v}}}\int_{\mathfrak{X}}Q(B_{\mathsf{z}}{\vert\,}x) \,\mathrm{d}\tilde{P}(v,x) \right. \nonumber \\ \text{ holds for all } \left. B_{\mathsf{v}}\in\mathscr{F}_{\mathfrak{V}}, B_{\mathsf{z}}\in\mathscr{F}_{\mathfrak{Z}}, \text{ as } \tilde{P} \text{ runs through } \mathcal{P}\subset \mathcal{P}_{\mathfrak{V}\mathfrak{X}} \vphantom{\int_{B_{\mathsf{v}}}\int_{\mathfrak{X}}} \right\}, \label{priv:eq:vzq95model} \end{align}\tag{12}\] the set of all possible distributions of \((V,Z)\) induced by the mechanism \(Q\) as the distribution of \((V,X)\) varies across \(\mathcal{P}\); here, \(\mathcal{P}_{\mathfrak{V}\mathfrak{Z}}\) is the set of all probability distributions on \((\mathfrak{V}\times\mathfrak{Z},\mathscr{F}_{\mathfrak{V}\times\mathfrak{Z}})\). Fixing \(\mathcal{P}=\mathcal{P}_{VX}\) and \(\tilde{P}=P_{VX}\) in 12 yields \(P=P_{VZ}\).

How to choose the mechanism? First, the output space \(\mathfrak{Z}\) has to be specified. The choice of \(\mathfrak{Z}\) is an unexplored topic in privacy literature, beyond our current scope. Hence, for some of our results to follow, \(\mathfrak{Z}\) can be any given measurable space; however, \(\mathfrak{Z}=\mathfrak{X}\) — the usual choice in the literature — shall yield more insightful results.

Second, given \(\mathfrak{Z}\), a mechanism \(Q\in\mathcal{Q}(\mathfrak{X}\to\mathfrak{Z})\) has to be chosen. Regard 11 as the flow of information from \(P_{VX}\) to \(P_{VZ}\). If \(Q(\cdot{\vert\,}x)\) does not depend on \(x\), then \(Z\) carries no information about \(X\), constituting maximal privacy but precluding inference. In the other extreme, if \(\mathfrak{Z}=\mathfrak{X}\), and \(Q(\cdot{\vert\,}x)=\delta_{x}\) concentrates on \(X\), then \(P_{VZ}=P_{VX}\), leading to the opposite effect. Thus, a sufficient and necessary condition for the identification of every parameter \(\chi(P_{VX})\) from \(P_{VZ}\) is the existence of a map \(L_Q:\mathcal{P}_{VZ}(Q, \mathcal{P}_{VX})\to\mathcal{P}_{VX}\) such that \[\begin{align} P_{VX}=L_Q(P_{VZ}). \label{priv:eq:linopL} \end{align}\tag{13}\] The map \(L_Q\) inverts 11 to recover \(P_{VX}\) from every \(P_{VZ}\) generated by a given \(Q\). With \(L_Q\), every parameter3 of the sensitive-data distribution is identifiable from the noisy-data distribution as \[\begin{align} \psi(P_{VZ})\mathrel{\vcenter{:}}=\chi \circ L_Q(P_{VZ}) = \chi(P_{VX}). \label{priv:eq:identification95psi} \end{align}\tag{14}\] The dependence of \(\psi\) on \(Q\) remains implicit in our notation. To infer \(\chi(P_{VX})\), we wish to choose \(Q\) such that \(L_Q\) exists. Consider first a discrete \(X\).

Example 1 (Finitely discretely distributed \(X\)). Suppose that \(\mathfrak{X}=\left\{x_1,\ldots,x_{|\mathfrak{X}|}\right\}\) and \(\mathfrak{Z}=\left\{z_1,\ldots,z_{|\mathfrak{Z}|}\right\}\) are finite sets, and that \(P_{VX}\) has a \(\nu_V\times\nu_X\)-density \(p_{VX}\) for the counting measure \(\nu_X\). Then \(P_{VZ}\) in 11 admits a \(\nu_V\times\nu_X\)-density \[\begin{align} p_{VZ}(v,z)=\sum_{x\in\mathfrak{X}} Q(\left\{z\right\}{\vert\,}x)p_{VX}(v,x),\quad (v,z)\in\mathfrak{V}\times\mathfrak{Z}. \label{priv:eq:vz95disrete95dens} \end{align}\qquad{(4)}\] Representing \(Q\) as the \(|\mathfrak{Z}|\)-by-\(|\mathfrak{X}|\) matrix \[\begin{align} Q = \begin{bmatrix} (Q(\left\{z_j\right\}\mid x_1))_{j\in[|\mathfrak{Z}|]} & (Q(\left\{z_j\right\}\mid x_2))_{j\in[|\mathfrak{Z}|]} & \cdots & (Q(\left\{z_j\right\}\mid x_{|\mathfrak{X}|}))_{j\in[|\mathfrak{Z}|]} \end{bmatrix}, \label{priv:eq:q95as95matrix} \end{align}\qquad{(5)}\] the display ?? is equivalent to \[\begin{align} [0,1]^{|\mathfrak{Z}|\times 1}\ni \bar{p}_{VZ}(v) \mathrel{\vcenter{:}}= (p_{VZ}(v,z_j))_{j\in[|\mathfrak{Z}|]} = Q (p_{VX}(v,x_j))_{j\in[{|\mathfrak{X}|}]} \eqqcolon Q \bar{p}_{VX}(v),\quad v\in\mathfrak{V}. \end{align}\] For \(z\in\mathbb{R}^{|\mathfrak{Z}|\times 1}, x\in\mathbb{R}^{|\mathfrak{X}|\times 1}\) and \(Q\in\mathcal{Q}(\left\{x_1,\ldots,x_{|\mathfrak{X}|}\right\}\to\left\{z_1,\ldots,z_{|\mathfrak{Z}|}\right\})\), consider the system of linear equations \(z=Qx\) in \(x\). Only if \(|\mathfrak{Z}|\geq|\mathfrak{X}|\), can this system have a unique solution. Let us impose \(|\mathfrak{Z}|\mathrel{\vcenter{:}}=|\mathfrak{X}|\eqqcolon J\), and set \[\begin{align} \mathcal{Q}_J\mathrel{\vcenter{:}}=\mathcal{Q}(\left\{x_1,\ldots,x_J\right\}\to\left\{z_1,\ldots,z_J\right\}). \label{priv:eq:discreteQdef} \end{align}\qquad{(6)}\] Then the system has a unique solution for all \(z\in\mathbb{R}^{J\times 1}\) if and only if \(Q\in\mathcal{Q}_J\) viewed as a matrix is invertible with inverse \(Q^{-1}\), in which case the solution is \(Q^{-1}z\) (e.g.[41]). Conclude that if \(Q\) is invertible, then \(\bar{p}_{VX}(v)=Q^{-1}\bar{p}_{VZ}(v)\) for all \(v\in\mathfrak{V}\). Hence, if \[\begin{align} Q\in \mathcal{Q}^{\mathrm{I}}_J\mathrel{\vcenter{:}}=\left\{Q\in \mathcal{Q}_J: Q^{-1}\text{ exists}\right\} \label{priv:eq:discreteQinv} \end{align}\qquad{(7)}\] then \(L_Q\) in 13 exists, is unique, and is completely determined by the matrix \(Q^{-1}\). An example ([26]) of \(Q\in \mathcal{Q}^{\mathrm{I}}_J\) is \[\begin{align} Q=c_{J,\alpha} \begin{bmatrix} e^\alpha & 1 & \cdots &1 \\ 1 & e^\alpha & \cdots & 1 \\ \vdots & \vdots & \ddots & \vdots \\ 1 & 1& \cdots & e^\alpha \end{bmatrix}, Q^{-1}= c_{J,\alpha}^{(-1)} \begin{bmatrix} e^\alpha+J-2 & -1 & \cdots &-1 \\ -1 & e^\alpha+J-2 & \cdots & -1 \\ \vdots & \vdots & \ddots & \vdots \\ -1 & -1& \cdots & e^\alpha+J-2 \end{bmatrix}\label{priv:eq:discrete95covar95Qexample} \end{align}\qquad{(8)}\] with \(c_{J,\alpha}\mathrel{\vcenter{:}}=\frac{1}{e^\alpha+J-1}\) and \(c_{J,\alpha}^{(-1)}\mathrel{\vcenter{:}}=\frac{e^\alpha+J-1}{e^{2\alpha}+(J-2)e^\alpha-J+1}\) for any \(\alpha>0\).

1 shows in ?? the role of the parameter \(\alpha\) determining the privacy level: an \(\alpha\approx 0\) equalises the entries of \(Q\), with no information flowing from \(X\) to \(Z\). Formally, \(Q\) in ?? satisfies \((e^\alpha-1)\)-total-variational privacy for any \(0<\alpha\leq\log(2)\):

Definition 1 (Local \(\alpha\)-Total-Variation Privacy (\(\alpha\)-LTVP) [1]). For \(0\leq\alpha\leq1\), a mechanism \(Q\in\mathcal{Q}(\mathfrak{X}\to\mathfrak{Z})\) is locally \(\alpha\)-total-variationally private if \(\sup_{B\in\mathscr{F}_{\mathfrak{Z}}}|Q(B{\vert\,}x)-Q(B{\vert\,}x')|\leq\alpha\) for all \(x,x'\in\mathfrak{X}\).

For generic \(X\) as well, total-variational privacy proves to be a suitable paradigm to ensure the existence of \(L_Q\). Consider \[\begin{align} \mathcal{Q}_{\delta}\mathrel{\vcenter{:}}=\left\{ Q\in \mathcal{Q}(\mathfrak{X}\to\mathfrak{X}): Q(B {\vert\,}x)= \alpha \delta_x(B)+(1-\alpha)\bar{Q}(B), \alpha\in(0,1), \bar{Q}\in \mathcal{P}_{\mathfrak{X}}\right\}, \label{priv:eq:qtv} \end{align}\tag{15}\] where \(\mathcal{P}_{\mathfrak{X}}\) is the set of all probability measures on \((\mathfrak{X},\mathscr{F}_{\mathfrak{X}})\). The \(Z\) drawn from a mechanism in 15 for a unit with \(X=x\) equals \(x\) itself with probability \(\alpha\), and is pure noise drawn from \(\bar{Q}\) with probability \(1-\alpha\). Hence, the smaller \(\alpha\), the stricter the privacy. Trivially, any \(Q\) in 15 is \(\alpha\)-LTVP, and, by 4 in 10.3, it ensures the existence of \[\begin{align} (L_QP_{VZ})(B_{\mathsf{v}},B_{\mathsf{x}})\mathrel{\vcenter{:}}=\frac{1}{\alpha}P_{VZ}(B_{\mathsf{v}},B_{\mathsf{x}})-\frac{1-\alpha}{\alpha}P_V(B_{\mathsf{v}})\bar{Q}(B_{\mathsf{x}}) \\ =\frac{1}{\alpha}P_{VZ}(B_{\mathsf{v}},B_{\mathsf{x}})-\frac{1-\alpha}{\alpha}P_{VZ}(B_{\mathsf{v}}, \mathfrak{X})\bar{Q}(B_{\mathsf{x}})=P_{VX}(B_{\mathsf{v}},B_{\mathsf{x}}),\quad B_{\mathsf{v}}\in\mathscr{F}_{\mathfrak{V}},B_{\mathsf{x}}\in\mathscr{F}_{\mathfrak{X}}, \end{align}\] a linear map, whereby the identification 14 of \(\chi(P_{VX})\) from \(P_{VZ}\) readily follows.

It may seem restrictive that for discrete \(X\) we allow for any invertible mechanism, but for generic \(X\) we confine ourselves to 15 . However, for generic \(X\), it appears difficult to obtain \(L_Q\) without a Dirac measure; see 10.5 for further discussion. In summary, we collect in \[\begin{align} \mathcal{Q}_{\psi}\mathrel{\vcenter{:}}=\mathcal{Q}^{\mathrm{I}}_J\cup\mathcal{Q}_{\delta} \end{align}\] the set of mechanisms implying 14 , understanding that \(Q\in\mathcal{Q}^{\mathrm{I}}_J\) only if \(\mathfrak{X}\) and \(P_{VX}\) are as in 1.

4.2 Private Inferential Properties↩︎

Now we derive the private analogue of the inferential properties in 3.2. This entails the efficient influence function of \(\psi(P_{VZ})\) in 14 at \(P_{VZ}\) in the model \(\mathcal{P}_{VZ}(Q, \mathcal{P}_{VX})\) of 12 , given a mechanism \(Q\in\mathcal{Q}(\mathfrak{X}\to\mathfrak{Z})\).4 In line with private-inference practice, \(Q\) is treated as common knowledge available for inference. Some of our results hold for any model \(\mathcal{P}_{VX}\), but the main interest is in the nonparametric model \(\mathcal{P}_{\mathfrak{V}\mathfrak{X}}\) of 1 .

For \(Q\in\mathcal{Q}(\mathfrak{X}\to\mathfrak{Z})\), define the linear operator \(Q_\mathcal{X}: L_2(P_{VZ})\to L_2(P_{VX})\) as \[\begin{align} (Q_\mathcal{X}k)(v,x)&\mathrel{\vcenter{:}}=\int_\mathfrak{Z}k(v,z)Q(\mathrm{d}z{\vert\,}x)=\mathbb{E}\left[\left. k(V,Z)\,\right\vert\, V=v,X=x\right], \label{priv:eq:linopQX} \end{align}\tag{16}\] whose properties are derived in 3 in 10.3. In particular, it has adjoint \(Q_{\mathcal{X}}^{*}:L_2(P_{VX})\to L_2(P_{VZ}), h\mapsto\mathbb{E}\left[\left. h(V,X)\,\right\vert\, V=\cdot,Z=\cdot\right]\), and is invertible: when \(Q\in\mathcal{Q}_J\), then \(Q_\mathcal{X}^{-1}\) exists and is unique if and only if \(Q\in\mathcal{Q}^{\mathrm{I}}_J\) with inverse \((Q^\intercal)^{-1}\) viewed as a transposed matrix; more generally, when \(Q\in\mathcal{Q}_{\delta}\), then the range-restricted \(Q_\mathcal{X}:L_2(P_{VZ})\to L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})\) and its inverse \(Q_\mathcal{X}^{-1}:L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})\to L_2(P_{VZ})\) are \[\begin{align} (Q_\mathcal{X}k)(v,x)&=\alpha k(v,x)+(1-\alpha)\int_\mathfrak{X}k(v,z)\bar{Q}(\mathrm{d}z), \quad (v,x)\in\mathfrak{V}\times\mathfrak{X}, \\ (Q_\mathcal{X}^{-1}h)(v,z)&=\frac{1}{\alpha}h(v,z)-\frac{1-\alpha}{\alpha}\int_\mathfrak{X}h(v,x)\bar{Q}(\mathrm{d}x), \quad (v,z)\in\mathfrak{V}\times\mathfrak{X}. \end{align}\]

Theorem 2 (Efficient Influence Function of \(\psi(P_{VZ})\)). Suppose that \(\chi(P_{VX})\) has efficient influence function \(\varphi\in \overline{\mathrm{lin}\,}\mathcal{T}_{VX}\) at \(P_{VX}\) in some model \(\mathcal{P}_{VX}\subset \mathcal{P}_{\mathfrak{V}\mathfrak{X}}\) with tangent set \(\mathcal{T}_{VX}\), and that \(Q\in\mathcal{Q}_{\psi}\). If \(\varphi \in Q_\mathcal{X}Q_{\mathcal{X}}^{*}\mathcal{T}_{VX}\), then the efficient influence function of \(\psi(P_{VZ})\) at \(P_{VZ}\) in the model \(\mathcal{P}_{VZ}(Q, \mathcal{P}_{VX})\) of 12 is \(Q_\mathcal{X}^{-1}\varphi\). If \(\varphi=\tilde{\chi}\) in ?? in the nonparametric model \(\mathcal{P}_{VX}=\mathcal{P}_{\mathfrak{V}\mathfrak{X}}\) satisfies \(\tilde{\chi}\in L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})\), then \[\begin{align} \tilde{\psi}\mathrel{\vcenter{:}}= Q_\mathcal{X}^{-1}\tilde{\chi}\label{priv:eq:eif95chi95vz} \end{align}\qquad{(9)}\] is the efficient influence function of \(\psi(P_{VZ})\) at \(P_{VZ}\) in the model \(\mathcal{P}_{VZ}(Q, \mathcal{P}_{\mathfrak{V}\mathfrak{X}})\).

In 2, the conditions \(\varphi \in Q_\mathcal{X}Q_{\mathcal{X}}^{*}\mathcal{T}_{VX}\) and \(\tilde{\chi}\in L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})\) are important. For instance, if \(\bar{Q}\) of \(Q\in\mathcal{Q}_{\delta}\) has large mass at extreme locations of \(\tilde{\chi}\), the latter may fail. If they hold, then an asymptotically efficient estimator of \(\psi(P_{VZ})\) based on a random sample from \(P_{VZ}\in \mathcal{P}_{VZ}(Q,\) \(\mathcal{P}_{\mathfrak{V}\mathfrak{X}})\) with a given \(Q\in\mathcal{Q}_{\psi}\) has limiting variance \[\begin{align} P_{VZ}\tilde{\psi}^2=P_{VX}[Q_\mathcal{X}(\tilde{\psi}^2)]=P_{VX}[Q_\mathcal{X}[(Q_\mathcal{X}^{-1}\tilde{\chi})(Q_\mathcal{X}^{-1}\tilde{\chi})]] \label{priv:eq:eff95var95vz} \end{align}\tag{17}\] by the properties of \(Q_\mathcal{X}\). When there is no privacy, so \(Q_\mathcal{X}\) and \(Q_\mathcal{X}^{-1}\) are the identity, 17 equals the nonprivate efficiency bound \(P_{VX}\tilde{\chi}^2\) in 3.2, as expected. Specifically, if \(Q\in\mathcal{Q}_{\delta}\), then we have the bounds \[\begin{align} \begin{aligned}\label{priv:eq:eff95var95vz95bound} P_{VX}\tilde{\chi}^2+\frac{1-\alpha}{\alpha}\left((P_V\otimes\bar{Q})\tilde{\chi}^2- P_V\left(\int \tilde{\chi}(V,x)\bar{Q}(\mathrm{d}x)\right)^2\right) \leq P_{VZ}\tilde{\psi}^2 \\ \leq \frac{2-\alpha}{\alpha} P_{VX}\tilde{\chi}^2+\frac{2(2-\alpha)(1-\alpha)}{\alpha^2}(P_V\otimes\bar{Q})\tilde{\chi}^2 \end{aligned} \end{align}\tag{18}\] by 3; hence, the private efficiency bound \(P_{VZ}\tilde{\psi}^2\) is never smaller than the nonprivate bound \(P_{VX}\tilde{\chi}^2\), and the stricter the privacy, the larger this gap in general.

5 Private Estimation↩︎

In this section, we construct a private analogue of the one-step estimator 8 of 3 , and show that it reaches the efficient limit 17 in the nonparametric model 1 for \(P_{VX}\), given a mechanism \(Q\in\mathcal{Q}_{\psi}\). We assume that three, mutually independent, random samples \(\bar{\mathcal{S}}=((V_i,Z_i))_{i\in[n]}\), \(\bar{\mathcal{S}}'=((V_i',Z_i'))_{i\in[n]},\) \(\bar{\mathcal{S}}''=((V_i'',Z_i''))_{i\in[n]}\) from \(P_{VZ}\) are available for inference.5

Analogously to 8 , we begin with an initial estimator \(\psi(\hat{P}_{VZ})\) and correct it as \[\begin{align} \hat{\psi}_n &\mathrel{\vcenter{:}}=\psi(\hat{P}_{VZ})+\bar{\mathbb{P}}_n\hat{\tilde{\psi}}= \psi(\hat{P}_{VZ})+\frac{1}{n}\sum_{i\in[n]} \hat{\tilde{\psi}}(V_i,Z_i), \end{align}\] where \(\bar{\mathbb{P}}_n\) is the empirical measure of \(\bar{\mathcal{S}}\), and \[\begin{align} \hat{\tilde{\psi}}\mathrel{\vcenter{:}}=&\, Q_\mathcal{X}^{-1}\check{\tilde{\chi}}, \tag{19} \\ \check{\tilde{\chi}}(v,x)\mathrel{\vcenter{:}}=&\, \check{r}(v_{\mathit{1}},x)(m(v,x)-{\check{\mu}_{\mathcal{X}}}(v_{\mathit{1}},x))+\frac{\mathbb{1}_{v_{\mathit{2}}=c}}{\check{p}_{V_{\mathit{2}}}(c)}(g(v,x)-\check{\gamma}_{\mathcal{V}}(c))\check{e}\nonumber \\ &+ f(v,x, {\check{\mu}_{\mathcal{X}}}, \check{\gamma}_{\mathcal{V}}(c))-\psi(\hat{P}_{VZ}) \eqqcolon\check{\tilde{\chi}}_0(v,x)- \psi(\hat{P}_{VZ}). \tag{20} \end{align}\] are estimates of the private and nonprivate influence functions \(\tilde{\psi}\) and \(\tilde{\chi}\) in ?? and ?? , respectively. The inverse \(Q_\mathcal{X}^{-1}\) of 16 is known as \(Q\) is known. As \(Q_\mathcal{X}^{-1}h=h\) for a constant function \(h\), we have \(\hat{\psi}_n=\frac{1}{n}\sum_{i\in[n]} (Q_\mathcal{X}^{-1}\check{\tilde{\chi}}_0)(V_i,Z_i)\), so \(\psi(\hat{P}_{VZ})\in\mathbb{R}\) can be arbitrary.

In contrast, all the estimates \[\begin{align} \label{priv:eq:nuisance95vz} \begin{aligned} \check{\eta}&\mathrel{\vcenter{:}}=(\check{r},{\check{\mu}_{\mathcal{X}}},\check{\gamma}_{\mathcal{V}}(c),\check{p}_{V_{\mathit{2}}}(c), \check{e})\in L_2(P_{V_{\mathit{1}}X})\times L_2(P_{V_{\mathit{1}}X}) \times\Gamma\times \mathbb{R}\times\mathbb{R}\,\text{ of } \\ \eta&\mathrel{\vcenter{:}}=(r,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c), p_{V_{\mathit{2}}}(c), e) \end{aligned} \end{align}\tag{21}\] are based on the noisy samples \(\bar{\mathcal{S}}'\) and \(\bar{\mathcal{S}}''\) as clarified in ¿tbl:priv:tab:priv95dr95samples?. Specifically, by 3, \[\begin{align} P_{VZ}Q_\mathcal{X}^{-1}h=P_{VX}h\,\text{ for all } h\in L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q}). \label{priv:eq:changemeasure} \end{align}\tag{22}\] This motivates the estimators \[\begin{align} \check{e}\mathrel{\vcenter{:}}=\frac{1}{n}\sum_{i\in[n]} \partial_\gamma \bar{f}(V_i'',Z_i'',{\check{\mu}_{\mathcal{X}}},\check{\gamma}_{\mathcal{V}}(c)), \tag{23} \\ \check{\gamma}_{\mathcal{V}}(c)\mathrel{\vcenter{:}}=\frac{1}{\check{p}_{V_{\mathit{2}}}(c)}\frac{1}{n}\sum_{i\in[n]}\bar{g}_c(V_i',Z_i'),\quad \check{p}_{V_{\mathit{2}}}(c)\mathrel{\vcenter{:}}= N_c/n,\, N_c\mathrel{\vcenter{:}}=\sum_{i\in[n]}\mathbb{1}_{V_{\mathit{2}i}'=c}, \tag{24}\\ \bar{f}(v,z,\mu,\gamma)\mathrel{\vcenter{:}}=(Q_\mathcal{X}^{-1}(v,x)\mapsto f(v,x,\mu,\gamma))(v,z), \, (v,z,\mu,\gamma)\in \mathfrak{V}\times\mathfrak{Z}\times L_2(P_{V_{\mathit{1}}X})\times\Gamma, \tag{25} \\ \bar{g}_c(v,z)\mathrel{\vcenter{:}}=(Q_\mathcal{X}^{-1}g_c)(v,z),\quad g_c(v,x)\mathrel{\vcenter{:}}=\mathbb{1}_{v_{\mathit{2}}=c}g(v,x),\quad (v,x,z)\in\mathfrak{V}\times\mathfrak{X}\times\mathfrak{Z}. \end{align}\] Indeed, \(P_{VZ}\partial_\gamma \bar{f}=P_{VX}\partial_\gamma f\) and \(\mathbb{E}\check{\gamma}_{\mathcal{V}}(c)=\gamma_{\mathcal{V}}(c)\), by \(\partial_\gamma \bar{f}=Q_\mathcal{X}^{-1}\partial_\gamma f\) and the definition 2 of \(\gamma_{\mathcal{V}}(c)\) as \(\frac{1}{p_{V_{\mathit{2}}}(c)}\mathbb{E}g_c(V,X)\). It follows that \(\gamma_{\mathcal{V}}(c)\) and \(p_{V_{\mathit{2}}}(c)\) are estimable by 24 at rate \(O_{{P_{VZ}}}\left(n^{-1/2}\right)\) because of their low dimensionality.

The estimation of \((\mu_{\mathcal{X}},r)\) is addressed in 6. If the estimators 21 are consistent and \(f\) is continuous in an appropriate norm accounting for the noise measure \(\bar{Q}\) in 15 as specified in 2 in 8, then the second, empirical process term in the decomposition \[\begin{align} \sqrt{n}(\hat{\psi}_n - \psi(P_{VZ})) &= \sqrt{n}\bar{\mathbb{P}}_n\tilde{\psi}+\sqrt{n}(\bar{\mathbb{P}}_n-P_{VZ})(\hat{\tilde{\psi}}-\tilde{\psi})+\sqrt{n}\bar{R}_n, \tag{26} \\ \bar{R}_n&\mathrel{\vcenter{:}}=\psi(\hat{P}_{VZ})-\psi(P_{VZ})+P_{VZ} \hat{\tilde{\psi}}, \tag{27} \end{align}\] is \(o_{P_{VZ}}\left(1\right)\). The first term \(\sqrt{n}\bar{\mathbb{P}}_n\tilde{\psi }\overset{P_{VZ}}{\rightsquigarrow}\) \(\mathcal{N}(0,P_{VZ}\tilde{\psi}^2)\) by the standard central limit theorem as \(\tilde{\psi}\in L_2^0(P_{VZ})\) is the efficient influence function ?? . The change-of-measure property 22 combined with 2 and 20 implies that the third term \[\begin{align} \label{priv:eq:bias95r95vz} \begin{aligned} \bar R_n =&\;\psi(\hat{P}_{VZ})-\chi(P_{VX})+P_{VX}\check{\tilde{\chi}} \\ =&\;-P_{VX}(r-\check{r})(\mu_{\mathcal{X}}-{\check{\mu}_{\mathcal{X}}}) +(\gamma_{\mathcal{V}}(c)-\check{\gamma}_{\mathcal{V}}(c))\left(\frac{p_{V_{\mathit{2}}}(c)}{\check{p}_{V_{\mathit{2}}}(c)}\check{e}-\bar{e}'\right) \\ &-(\gamma_{\mathcal{V}}(c)-\check{\gamma}_{\mathcal{V}}(c))^2\frac{P_{VX}\partial_\gamma^2 f(V,X,{\check{\mu}_{\mathcal{X}}},\check{\gamma}_{\mathcal{V}}(c))}{2}, \end{aligned} \end{align}\tag{28}\] for some \(\tilde{\gamma}_{\mathcal{V}}(c)\) between \(\gamma_{\mathcal{V}}(c)\) and \(\check{\gamma}_{\mathcal{V}}(c)\), and \[\begin{align} \bar{e}'\mathrel{\vcenter{:}}= P_{VX}\partial_\gamma f(V,X,{\check{\mu}_{\mathcal{X}}},\check{\gamma}_{\mathcal{V}}(c)) = P_{VZ}\partial_\gamma \bar{f}(V,Z,{\check{\mu}_{\mathcal{X}}},\check{\gamma}_{\mathcal{V}}(c)). \label{priv:eq:derivhatprime95vz} \end{align}\tag{29}\] Whence, \(\sqrt{n}\bar{R}_n=o_{P_{VZ}}\left(1\right)\) under vanishing product \(({\check{\mu}_{\mathcal{X}}}-\mu_{\mathcal{X}})(\check{r}-r)\) of estimation errors and regularity conditions. This yields our main result whereby \(\hat{\psi}_n\) is rate-double-robust and asymptotically efficient in the nonparametric model 1 for \(P_{VX}\).

Assumption 1 (Rates of Private Estimators). The estimators 21 satisfy \(P_{VX}(({\check{\mu}_{\mathcal{X}}}-\mu_{\mathcal{X}})(\check{r}-r))=o_{P_{VZ}}\left(n^{-1/2}\right)\), \(\check{\gamma}_{\mathcal{V}}(c)-\gamma_{\mathcal{V}}(c)=O_{{P_{VZ}}}\left(n^{-1/2}\right)\), \(\check{p}_{V_{\mathit{2}}}(c)-p_{V_{\mathit{2}}}(c)=O_{{P_{VZ}}}\left(n^{-1/2}\right)\), and \(P_{VX} \partial_\gamma^2 f(V,X,{\check{\mu}_{\mathcal{X}}},\tilde{\gamma}_{\mathcal{V}}(c))=O_{{P_{VZ}}}\left(1\right)\) for \(\tilde{\gamma}_{\mathcal{V}}(c)\) between \(\check{\gamma}_{\mathcal{V}}(c)\) and \(\gamma_{\mathcal{V}}(c)\).

Corollary 1 (Asymptotic Efficiency of \(\hat{\psi}_n\)). Suppose that \(P_{VZ}\in\mathcal{P}_{VZ}(Q, \mathcal{P}_{\mathfrak{V}\mathfrak{X}})\) for a fixed mechanism \(Q\in\mathcal{Q}_{\psi}\). If [priv:ass:consistent_nuisance_priv,priv:ass:rates_nuisance_priv] hold, then \(\sqrt{n}(\hat{\psi}_n-\psi(P_{VZ}))\overset{P_{VZ}}{\rightsquigarrow}\mathcal{N}(0, P_{VZ}\tilde{\psi}^2)\) as \(n\to\infty\).

Use of Samples for Estimation
Estimators Samples
\(\bar{\mathcal{S}}\) \(\bar{\mathcal{S}}'\) \(\bar{\mathcal{S}}''\)
\(\hat\psi_n\)
\(\expderivchk\)
\(\rieszXchk,\muXchk,\gammaVchk(c), \chk p_{V_{\mhit{2}}}(c)\)

A sample is used in the construction of an estimator if and only if \(✔\) is present in their corresponding cell.

6 Private Estimation of Nuisance Parameters↩︎

1 essentially shows that if the product of the estimation errors of the regression \(\mu_{\mathcal{X}}\) and of the Riesz representer \(r\) is small enough, then \(\hat{\psi}_n\) is asymptotically efficient. In this section, we consider the estimation of \((\mu_{\mathcal{X}},r)\) from the random sample \(\bar{\mathcal{S}}'\) of 5 from \(P_{VZ}\in\mathcal{P}_{VZ}(Q, \mathcal{P}_{VX})\) given a fixed mechanism \(Q\in\mathcal{Q}_{\psi}\).

By the Cauchy–Schwarz inequality, the product is bounded by individual errors as \(P_{VX}({\check{\mu}_{\mathcal{X}}}-\mu_{\mathcal{X}})(\check{r}-r)\leq \sqrt{P_{VX}({\check{\mu}_{\mathcal{X}}}-\mu_{\mathcal{X}})^2}\sqrt{P_{VX}(\check{r}-r)^2}\). If \(\mu_{\mathcal{X}}\) belongs to a finite-dimensional model smoothly indexed by \(\theta\in\mathbb{R}^K\), then the product vanishes fast enough for 1 to apply under regularity conditions, even if \(\check{r}\) is an infinite-dimensional nonparametric thus slower estimator, and vice versa. To accommodate this tradeoff, we present results for finite- and infinite-dimensional models.

For finite-dimensional smooth models, a common estimator is the method-of-moments ([5], [6]). In 2, we devise a locally-private method-of-moments estimator enabled by property 22 , and derive its limiting distribution under the standard regularity conditions of 3 in 11.1. As in 18 , the worst-case dependence of the limiting variance \(\bar{\Sigma}\) on \(\alpha\) is \(1/\alpha^2\) driven by \(\bar{\Phi}\).

Proposition 2 (Private Method-of-Moments). Let \(\theta_0\mathrel{\vcenter{:}}=\arg\min_{\theta\in\Theta} P_{VX}\Xi_\theta\) for a fixed \(\Xi_{\theta}:\mathfrak{V}\times\mathfrak{X}\to\mathbb{R}\), \(\theta\in\Theta\subset\mathbb{R}^K\), for a fixed \(K\). Let \(\phi_{\tilde{\theta}}(v,x)\mathrel{\vcenter{:}}=\mathrm{D}_\theta \Xi_{\tilde{\theta}}(v,x)^\intercal\) and \(\dot{\phi}_{\tilde{\theta}}(v,x)\mathrel{\vcenter{:}}=\mathrm{D}_\theta \phi_{\tilde{\theta}}(v,x)\) be the derivates as maps to \(\mathbb{R}^{K\times 1}\) and to \(\mathbb{R}^{K\times K}\), respectively, for \((v,x,\tilde{\theta})\in\mathfrak{V}\times\mathfrak{X}\times\Theta\). Let \(\bar{A}_n\in\mathbb{R}^{K\times K}\) be an arbitrary sequence of (possibly random and then \(\sigma(\bar{\mathcal{S}}')\)-measurable) matrices with \(\bar{A}_n\overset{P_{VZ}}{\to}\bar{A}_0\) as \(n\to\infty\) for a symmetric positive definite \(\bar{A}_0\), and \(\check{\theta}\) be the solution to \(\theta\mapsto \bar{\Lambda}_n(\theta)\mathrel{\vcenter{:}}=\big(\bar{\mathbb{P}}_n'\bar{\phi}_{\theta}^\intercal\big)\bar{A}_n(\bar{\mathbb{P}}_n'\bar{\phi}_{\tilde{\theta}}\big) \equiv 0\) up to \(\bar{\Lambda}_n(\check{\theta})=o_{P_{VZ}}\left(n^{-1/2}\right)\), where \(\bar{\phi}_{\tilde{\theta}}\mathrel{\vcenter{:}}=\mathrm{D}_\theta \bar{\Xi}_{\tilde{\theta}}\) with \(\bar{\Xi}_\theta\mathrel{\vcenter{:}}= Q_\mathcal{X}^{-1}\Xi_\theta\) for the inverse \(Q_\mathcal{X}^{-1}\) of 16 . Let 3 in 11.1 hold. If \(P_{VZ}\in\mathcal{P}_{VZ}(Q, \mathcal{P}_{VX})\), for a fixed \(Q\in\mathcal{Q}_{\psi}\) and \(P_{VX}\in\mathcal{P}_{VX}\) satisfying the given assumptions, then \(\sqrt{n}(\check{\theta}-\theta_0)\overset{P_{VZ}}{\rightsquigarrow} \mathcal{N}(0,\bar{\Sigma})\) as \(n\to\infty\), where \(\bar{\Sigma }\mathrel{\vcenter{:}}=({\dot{\Phi}^\intercal}\bar{A}_0{\dot{\Phi}})^{-1}{\dot{\Phi}^\intercal}\bar{A}_0 \bar{\Phi }\bar{A}_0{\dot{\Phi}}({\dot{\Phi}^\intercal}\bar{A}_0{\dot{\Phi}})^{-1}\), \(\dot{\Phi} \mathrel{\vcenter{:}}= P_{VX} \dot{\phi}_{\theta_0}\), \(\bar{\Phi }\mathrel{\vcenter{:}}= P_{VZ}\bar{\phi}_{\theta_0}\bar{\phi}_{\theta_0}^\intercal\) with \(\dot{\Phi} =P_{VX}\dot{\phi}_{\theta_0}\) and \(\phi_{\theta_0}=\mathrm{D}_\theta\Xi_{\theta_0}^\intercal\). Further, let \(\xi_{\theta}:\mathfrak{V}\times\mathfrak{X}\to\mathbb{R}\), \(\theta\in\Theta\), be (possibly random and then \(\sigma(\bar{\mathcal{S}}')\)-measurable) functions satisfying 3. Then \(\left\lVert{\xi_{\check{\theta}}-\xi_{\theta_0}}\right\rVert_{L_2(P_{VX})}=O_{{P_{VZ}}}\left(n^{-1/2}\right)\).

If \(\mu_{\mathcal{X}}\) or \(r\) follows a smooth model satisfying the conditions of 2, then they are estimable at parametric rate \(O_{{P_{VZ}}}\left(n^{-1/2}\right)\). Indeed, the correctly parametrised versions of the (well-known) relations (8 in 11) \[\begin{align} \mu_{\mathcal{X}}= \arg\min_{\mu\in L_2(P_{V_\mathit{1}}X)}P_{VX}\bigg[\Delta_{\mu}^2(V,X)\bigg], \quad \Delta_{\mu}^2(v,x)\mathrel{\vcenter{:}}=(m(v,x) - \mu(v_{\mathit{1}},x))^2, \tag{30} \\ r_\gamma = \arg\min_{h\in L_2(P_{V_\mathit{1}}X)}P_{VX}\Upsilon_{\gamma,h}, \quad \Upsilon_{\gamma, h}(v,x) \mathrel{\vcenter{:}}= h(v_\mathit{1},x)^2-2f(v,x,h,\gamma), \gamma\in\Gamma, \tag{31} \end{align}\] for the Riesz representer \(r_\gamma\) of \(\mu\mapsto \mathbb{E}f(V,X,\mu,\gamma)\), lead to \(\left\lVert{\xi_{\check{\theta}}-\xi_{\theta_0}}\right\rVert_{L_2(P_{VX})}=O_{{P_{VZ}}}\left(n^{-1/2}\right)\) for \(\xi_{\theta_0}\in\{\mu_{\mathcal{X}},r\}\); see [priv:cor:muXpara_rate_priv,priv:cor:rieszXpara_rate_priv] in 11.

For infinite-dimensional models \(\theta_0\in\{\mu_{\mathcal{X}},r\}\), we transform nonprivate estimators \(\hat{\theta}\) into private ones for estimation from the noisy data \(\bar{\mathcal{S}}'\). We consider estimators and their private transformation of the form \[\begin{align} \hat{\theta}(v_\mathit{1},x)\mathrel{\vcenter{:}}=\frac{1}{n}\sum_{i\in[n]}w_{n,i}(v_\mathit{1}, x, V_{i}',X_i',\hat{\vartheta}), \tag{32} \\ \tag{33} \begin{aligned} \check\theta(v_\mathit{1},x)&\mathrel{\vcenter{:}}=\frac{1}{n}\sum_{i\in[n]}\bar{w}_{n,i}(v_\mathit{1},x,V_i',Z_i',\check{\vartheta}), \\ \bar{w}_{n,i}(v_\mathit{1},x,v',z',\vartheta)&\mathrel{\vcenter{:}}=(Q_\mathcal{X}^{-1}(v',x')\mapsto w_{n,i}(v_\mathit{1}, x, v',x',\vartheta))(v',z'), \end{aligned} \end{align}\] where \(w_{n,i}:\mathfrak{V}_{\mathit{1}}\times\mathfrak{X}\times\mathfrak{V}\times\mathfrak{X}\times\mathcal{T}\to\mathbb{R}\), \(i\in[n]\), is a triangular array of known, nonrandom functions. Here, \(\hat{\vartheta}\) estimates \(\vartheta\in\mathcal{T}\), which allows \(\theta_0\) to depend on a “secondary-nuisance” parameter \(\vartheta\). We remain agnostic about \(\mathcal{T}\), affording flexible formulations. The operator \(Q_\mathcal{X}^{-1}\) is the inverse of 16 , and \(\check{\vartheta}\) is an estimator of \(\vartheta\) computed from \(\bar{\mathcal{S}}'\).

3 bounds the error of \(\check\theta\). In ?? , the first term is the secondary-nuisance error (e.g.the error of the conditioning density in kernel regression estimates), which is trivially zero if there is no dependence on \(\vartheta\); the second is a variance term, which usually scales inversely with the effective sample size (e.g. \(1/(nh^d)\) for \(d\)-dimensional kernel estimators with bandwidth \(h\)) modulo the noise level of the privacy mechanism. The third term in ?? is the squared bias, which, thanks to 22 , is identical to the nonprivate bias (e.g.\(h^{2\beta}\) for said kernel estimates of \(\beta\)-smooth parameters). These terms generally require a case-by-case analysis, but arguments in the nonprivate setting can carry over to the private setting under the mechanism \(Q\in\mathcal{Q}_{\delta}\). The noise level generally affects the first two terms, but never the third. This can translate into “usual,” nonprivate convergence rates of private versions of well-known estimators such as kernel or orthogonal series, modulo the noise level of the mechanism \(Q\); see 11.2.

Proposition 3 (Private Error Bounds). Fix a \(Q\in\mathcal{Q}_{\psi}\), and let \(\check\theta\) be defined according to 33 . Suppose that for a sequence of constants \((a_n)\), \[\begin{align} \label{priv:eq:secnubound} \begin{aligned} P_{VX}T_{n}^2&=O_{{P_{VZ}}}\left(a_n^2\right), \\ T_{n}(v_\mathit{1}, x)&\mathrel{\vcenter{:}}=\frac{1}{n}\sum_{i\in[n]}\left(\bar{w}_{n,i}(v_\mathit{1},x,V_i',Z_i',\check{\vartheta})-\bar{w}_{n,i}(v_\mathit{1},x,V_i',Z_i',\vartheta)\right),\quad (v_\mathit{1},x)\in \mathfrak{V}_{\mathit{1}}\times\mathfrak{X}. \end{aligned} \end{align}\qquad{(10)}\] Let \(\bar{\sigma}_i^2(v_\mathit{1},x)\mathrel{\vcenter{:}}=\mathop{\mathrm{\mathbb{V}}}\left[\bar{w}_{n,i}(v_\mathit{1},x,V,Z,\vartheta)\right]\), \((i,v_\mathit{1},x)\in [n]\times \mathfrak{V}_{\mathit{1}}\times\mathfrak{X}\). Then for any \(\theta\in L_2(P_{V_{\mathit{1}}X})\), \[\begin{align} \label{priv:eq:nonparaestim95priv95bound} \begin{aligned} P_{VX}(\check\theta-\theta)^2 \leq &\, O_{{P_{VZ}}}\left(a_n^2\right)+O_{{P_{VZ}}}\left(\frac{1}{n^2}\sum_{i\in[n]}\int \bar{\sigma}_i^2(v_\mathit{1},x) \,\mathrm{d}P_{V_{\mathit{1}}X}(v_{\mathit{1}},x)\right) \\ &+4\int\left\{\frac{1}{n}\sum_{i\in[n]}P_{VX}[w_{n,i}(v_\mathit{1},x,V,X,\vartheta)]-\theta(v_\mathit{1},x)\right\}^2\,\mathrm{d}P_{V_{\mathit{1}}X}(v_{\mathit{1}},x). \end{aligned} \end{align}\qquad{(11)}\] Morover, if \(Q\in\mathcal{Q}_{\delta}\), then \[\begin{align} \bar{\sigma}_i^2(v_\mathit{1},x)\leq \frac{4}{\alpha^2}\mathbb{E}w_{n,i,}^2(v_\mathit{1},x, V, Z,\vartheta) \label{priv:eq:sigmacondbound} \end{align}\qquad{(12)}\] for all \((i,v_\mathit{1},x)\in [n]\times\mathfrak{V}_{\mathit{1}}\times\mathfrak{X}\).

7 Conclusion↩︎

We introduced a class of rate-double-robust target parameters indexed linearly by an infinite-dimensional and nonlinearly by a low-dimensional regression. The inference of these targets was considered in a private setting, when the data — with values in general spaces — of each individual are protected by a local privacy mechanism. We derived the semiparametric properties of the private model, and constructed asymptotically efficient estimators under a fixed privacy mechanism. This efficiency was shown attainable under the same rate-double-robustness condition as in the nonprivate setting, involving two infinite-dimensional parameters. Therefore, we concluded with the private estimation of these two parameters, preserving their nonprivate rates modulo the noise level of the privacy mechanism.

8 Assumptions↩︎

Results in [priv:sec:estim_dr_priv,priv:sec:estim_nuisance_priv] rely on the following assumptions, respectively.

Assumption 2 (Consistent Private Estimators). Either the privacy mechanism \(Q\in\mathcal{Q}_{\delta}\) in 15 ; or \(Q\in\mathcal{Q}^{\mathrm{I}}_J\) and \((V,X)\) is distributed on a finite set with density \(p_{VX}\) with respect to the counting measure. Let \(\iota\mathrel{\vcenter{:}}= 1\) in the former and \(\iota\mathrel{\vcenter{:}}= 0\) in the latter case. Define the norm \(\left\lVert{\cdot}\right\rVert_{\mathrm{L}_2}\mathrel{\vcenter{:}}=\left\lVert{\cdot}\right\rVert_{L_2(P_{VX})}+\iota\left\lVert{\cdot}\right\rVert_{L_2(P_V\otimes \bar{Q})}\) and the measure \(P_{\mathrm{L}}\mathrel{\vcenter{:}}= P_{VX}+\iota P_V \otimes \bar{Q}\), and let \(\mathrm{L}_2\mathrel{\vcenter{:}}=\left\{f: \mathfrak{V}\times\mathfrak{X}\to \mathbb{R}: \left\lVert{f}\right\rVert_{\mathrm{L}_2}<\infty\right\}\). It holds that \(\tilde{\chi},\check{\tilde{\chi}},r\in \mathrm{L}_2\) and \[\begin{align} \left\lVert{f(\cdot,\mu,\gamma)-f(\cdot,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c))}\right\rVert_{\mathrm{L}_2}\to 0&\, \text{ as }\, \rho((\mu,\gamma),(\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c)))\to 0,\label{priv:eq:f95sq95continu95priv} \\ \left\lVert{\partial_\gamma f(\cdot,\mu,\gamma)-\partial_\gamma f(\cdot,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c))}\right\rVert_{\mathrm{L}_2}\to 0&\, \text{ as }\, \rho((\mu,\gamma),(\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c)))\to 0,\label{priv:eq:f95diffcontinu95priv} \end{align}\] {#eq: sublabel=eq:priv:eq:f95sq95continu95priv,eq:priv:eq:f95diffcontinu95priv} and \[\begin{align} \left\lVert{\check{r}-r}\right\rVert_{\mathrm{L}_2}&=o_{P_{VZ}}\left(1\right), \label{priv:eq:consistency95riesz95l295vz} \\ \check{\gamma}_{\mathcal{V}}(c)-\gamma_{\mathcal{V}}(c)&=o_{P_{VZ}}\left(1\right), \label{priv:eq:consistency95gammaVhat95vz} \\ \check{p}_{V_2}(c)-p_{V_2}(c)&=o_{P_{VZ}}\left(1\right). \label{priv:eq:consistency95phat95vz} \end{align}\] {#eq: sublabel=eq:priv:eq:consistency95riesz95l295vz,eq:priv:eq:consistency95gammaVhat95vz,eq:priv:eq:consistency95phat95vz} Further, it either holds that \[\begin{align} \left\lVert{m-\mu_{\mathcal{X}}}\right\rVert_{\infty}&=O\left(1\right), \label{priv:eq:consistency95bigobound95vz} \\ \left\lVert{\mu_{\mathcal{X}}-{\check{\mu}_{\mathcal{X}}}}\right\rVert_{\infty}&=o_{P_{VZ}}\left(1\right), \label{priv:eq:consistency95mu95supn95vz} \end{align}\] {#eq: sublabel=eq:priv:eq:consistency95bigobound95vz,eq:priv:eq:consistency95mu95supn95vz} or that \[\begin{align} \left\lVert{m-{\check{\mu}_{\mathcal{X}}}}\right\rVert_{\infty}&=O_{{P_{VZ}}}\left(1\right), \label{priv:eq:consistency95bigobound95hat95vz} \\ \left\lVert{\mu_{\mathcal{X}}-{\check{\mu}_{\mathcal{X}}}}\right\rVert_{\mathrm{L}_2}&=o_{P_{VZ}}\left(1\right), \label{priv:eq:consistency95mu95l295vz} \\ P_{\mathrm{L}}\left(\left\{(V_{\mathit{1}},X)\in \mathfrak{V}_{\mathit{1}}\times\mathfrak{X}: |r(V_{\mathit{1}},X)|>\bar{R} \right\}\right)&=0 \label{priv:eq:consistency95riesz95bound95vz} \end{align}\] {#eq: sublabel=eq:priv:eq:consistency95bigobound95hat95vz,eq:priv:eq:consistency95mu95l295vz,eq:priv:eq:consistency95riesz95bound95vz} for some constant \(\bar{R}<\infty\). We may replace \(\left\lVert{\cdot}\right\rVert_{\mathrm{L}_2}\) with \(\left\lVert{.}\right\rVert_{\infty}\) and ?? with \(\left\lVert{r}\right\rVert_{\infty}<\infty\).

Assumption 3 (Private Method-of-Moments). The set \(\Theta\subset\mathbb{R}^K\) is compact, and \(\theta_0\in\mathrm{Int}\,\Theta\) is the unique minimiser of \(\theta\mapsto P_{VX}\Xi_\theta\) for \(\Xi_{\theta}:\mathfrak{V}\times\mathfrak{X}\to\mathbb{R}\), \(\theta\in\Theta\). The derivative \(\phi_{\tilde{\theta}}(v,x)\mathrel{\vcenter{:}}=\mathrm{D}_\theta \Xi_{\tilde{\theta}}(v,x)^\intercal\) as a map to \(\mathbb{R}^{K\times 1}\) exists at all \((v,x,\tilde{\theta})\in\mathfrak{V}\times\mathfrak{X}\times\mathrm{Nb}(\theta_0)\), for a neighbourhood \(\mathrm{Nb}(\theta_0)\) of \(\theta_0\), and satisfy \(\left\lVert{\left\lVert{\phi_{\theta_0}}\right\rVert_{2}^2}\right\rVert_{L_1(P_{VZ})}<\infty\), where, for a fixed \((v,x)\in\mathfrak{V}\times\mathfrak{X}\), \(\left\lVert{\phi_{\theta_0}}\right\rVert_{2}^2(v,x)\) is sum of the \(K\) squared entries of \(\phi_{\theta_0}(v,x)\). The derivative \(\dot{\phi}_{\tilde{\theta}}(v,x)\mathrel{\vcenter{:}}=\mathrm{D}_\theta \phi_{\tilde{\theta}}(v,x)\) as a map to \(\mathbb{R}^{K\times K}\) exists at all \((v,x,\tilde{\theta})\in\mathfrak{V}\times\mathfrak{X}\times\mathrm{Nb}(\theta_0)\), and \(\theta\mapsto \dot{\phi}_{\theta}(v,x)\) is continuous at all \((v,x,\theta)\in\mathfrak{V}\times\mathfrak{X}\times\mathrm{Nb}(\theta_0)\), with the expectation \(P_{VX}\dot{\phi}_{\theta_0}\) existent and invertible as a matrix. Further, \(\left\lVert{\sup_{\theta\in\mathrm{Nb}(\theta_0)}\left\lVert{\dot{\phi}_\theta}\right\rVert_{1}}\right\rVert_{L_1(P_{VZ})}<\infty\), where, for a fixed \((v,x)\in\mathfrak{V}\times\mathfrak{X}\), \(\left\lVert{\dot{\phi}_\theta}\right\rVert_{1}(v,x)\) is the sum of the absolute values of the \(K^2\) entries of \(\dot{\phi}_\theta(v,x)\). The \(\xi_\theta\) are such that the derivative \(\mathrm{D}_\theta\xi_{\tilde{\theta}}(v,x)\) exists at all \(\tilde{\theta}\in\mathrm{Nb}(\theta_0)\) for all \((v,x)\in\mathfrak{V}\times\mathfrak{X}\), and \(\left\lVert{\sup_{\tilde{\theta}\in\mathrm{Nb}(\theta_0)}\left\lVert{\mathrm{D}_\theta \xi_{\tilde{\theta}}}\right\rVert_{2}}\right\rVert_{L_2(P_{VX})}=O_{{P_{VZ}}}\left(1\right)\).

9 Rate-Double-Robust Inference↩︎

9.1 contains examples of parameters in our rate-double-robust class. 9.2 proves the main results in 3.

9.1 Examples↩︎

In this section, we present examples of parameters \(\chi(P_{VX})=\mathbb{E}f(V,X, \mu_{\mathcal{X}}, \gamma_{\mathcal{V}}(c))\) defined in 3.1. In [priv:ex:ate_dr,priv:ex:att_dr] we discuss causal estimands; in [priv:ex:deriv,priv:ex:ray_int,priv:ex:line_int], “geometric” parameters. In 7, we study the special case of a known Riesz representer illustrated by 8 with a simple model from economics. Finally, in 9, we consider the extension of our class to feature dependence on multiple regressions.

In [priv:ex:ate_dr,priv:ex:att_dr], we adopt the potential outcome framework of [42] and [43]. Specifically, we consider a binary treatment \(D\in\{0,1\}\) and an observed outcome \(Y=DY^1+(1-D)Y^0\) for partially unobserved potential outcomes \(Y^0,Y^1\) with values in \(\mathbb{R}\). Let \(X\) be covariates taking values in a measurable space \((\mathfrak{X},\mathscr{F}_{\mathfrak{X}})\) with law \(P_X\) and satisfying unconfoundedness \(Y^d\!\perp\!\!\!\perp D\mid X\) for \(d\in\{0,1\}\). Let \[\begin{align} \label{priv:eq:causal95definitions} \begin{aligned} \mu_{\mathcal{X}}(d,x)&\mathrel{\vcenter{:}}=\mathbb{E}\left[\left. Y\,\right\vert\, D=d,X=x\right] \\ \pi_{\mathcal{X}}(d{\vert\,}x)&\mathrel{\vcenter{:}}=\mathbb{E}\left[\left. \mathbb{1}_{D=d}\,\right\vert\, X=x\right] \end{aligned} \end{align}\tag{34}\] for \((d,x)\in\{0,1\}\times\mathfrak{X}\). We assume that \(\pi_{\mathcal{X}}(d{\vert\,}X)\geq \epsilon\) for all \(d\in\{0,1\}\) \(P_{VX}\)-a.s., for some \(\epsilon>0\). In [priv:ex:ate_dr,priv:ex:att_dr], 1 recovers the familiar efficient influence functions for average treatment effects, under the nonparametric model with unknown propensity score. Moreover, the bias formulae in 3 show the average treatment effect on the treated rate-double-robust. This latter result is aligned with [4], and is an improvement on [3], who too, establish asymptotic normality, but not efficiency, as this parameter is not natively included in their class.

Example 2 (Average Treatment Effect). In the causal context 34 , let \(V\mathrel{\vcenter{:}}=(Y,D)\) and \(V_{\mathit{1}}\mathrel{\vcenter{:}}= D\), \(m(V,X)\mathrel{\vcenter{:}}= Y\); \(V_{\mathit{2}}\), \(g\), and \(\gamma_{\mathcal{V}}\) are not used. Then \(\mathbb{E}Y^d=\mathbb{E}\mu_{\mathcal{X}}(d,X)\) for \(d\in\{0,1\}\). Hence, \(f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c))\mathrel{\vcenter{:}}=\mu_{\mathcal{X}}(d,X)\) gives \(\chi(P_{VX})= \mathbb{E}Y^d\), while \(f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c))\mathrel{\vcenter{:}}=\mu_{\mathcal{X}}(1,X)-\mu_{\mathcal{X}}(0,X)\) gives \(\chi(P_{VX})= \mathbb{E}Y^1-\mathbb{E}Y^0\), with 4 , 5 , 6 , 56 holding.

The Riesz representer for \(\mathbb{E}Y^d\) is \(r(d',x)=\frac{\mathbb{1}_{d'=d}}{\pi_{\mathcal{X}}(d {\vert\,}x)}\) and for \(\mathbb{E}Y^1-\mathbb{E}Y^0\), it is \(r(d,x)=\frac{d}{\pi_{\mathcal{X}}(1{\vert\,}x)}-\frac{1-d}{1-\pi_{\mathcal{X}}(1{\vert\,}x)}.\) None of these representers \(r\) depends on \(\gamma_{\mathcal{V}}(c)\).

The efficient influence function ?? , in the nonparametric model (with unknown \(\pi_{\mathcal{X}}\)), of \(\mathbb{E}Y^d\), with \(\partial_\gamma f = 0\), is \[\begin{align} \tilde{\chi}(\tilde{v},\tilde{x}) =\frac{\mathbb{1}_{\tilde{d}=d}}{\pi_{\mathcal{X}}(d\mid \tilde{x})}(\tilde{y}-\mu_{\mathcal{X}}(d,\tilde{x}))+\mu_{\mathcal{X}}(d,\tilde{x}) - \chi(P_{VX}); \end{align}\] of \(\chi(P_{VX})=\mathbb{E}Y^1 - \mathbb{E}Y^0\), with \(\partial_\gamma f = 0\), is \[\begin{align} \tilde{\chi}(v,x)=&\, \frac{d}{\pi_{\mathcal{X}}(1{\vert\,}x)}(y-\mu_{\mathcal{X}}(1,x)) -\frac{1-d}{1-\pi_{\mathcal{X}}(1{\vert\,}x)}(y-\mu_{\mathcal{X}}(0,x)) \\ &+\mu_{\mathcal{X}}(1,X)-\mu_{\mathcal{X}}(0,X) - \chi(P_{VX}), \end{align}\] as in [44], since \(\mathbb{1}_{\tilde{d}=d}h(\tilde{d},\tilde{x})=\mathbb{1}_{\tilde{d}=d}h(d,\tilde{x})\) for any \(h\). Note that while the model for \((Y^0,Y^1,D,\) \(X)\) is not nonparametric, because it is constrained by \(Y^d\!\perp\!\!\!\perp D{\vert\,}X\), the model for \((Y,D,X)\) is nonparametric if the models for \(Y^0{\vert\,}X\), \(Y^1{\vert\,}X\), \(D{\vert\,}X\) and \(X\) are all nonparametric.

Consequently, the bias ?? established by 1 is \[\begin{align} R_n \mathrel{\vcenter{:}}=&\;-P_{VX}(r-\hat{r})(\mu_{\mathcal{X}}-{\hat{\mu}_{\mathcal{X}}}) \\ =&\, P_{DX}\left[\mathbb{1}_{D=d}\frac{\hat{\pi}_{\mathcal{X}}(d{\vert\,}X)-\pi_{\mathcal{X}}(d{\vert\,}X)}{\hat{\pi}_{\mathcal{X}}(d{\vert\,}X)\pi_{\mathcal{X}}(d{\vert\,}X)}\big({\hat{\mu}_{\mathcal{X}}}(D,X)-\mu_{\mathcal{X}}(D,X)\big)\right], \end{align}\] for \(\mathbb{E}Y^d\) where \({\hat{\mu}_{\mathcal{X}}},\hat{r}\) are some estimators of its Riesz representer and regression, respectively. For \(\mathbb{E}Y^1-\mathbb{E}Y^0\), the same quantity is \[\begin{align} R_n \mathrel{\vcenter{:}}=-P_{VX}(r-\hat{r})(\mu_{\mathcal{X}}-{\hat{\mu}_{\mathcal{X}}}) \\ = P_{DX}\left[D\frac{\hat{\pi}_{\mathcal{X}}(1{\vert\,}X)-\pi_{\mathcal{X}}(1{\vert\,}X)}{\hat{\pi}_{\mathcal{X}}(1{\vert\,}X)\pi_{\mathcal{X}}(1{\vert\,}X)}\big({\hat{\mu}_{\mathcal{X}}}(D,X)-\mu_{\mathcal{X}}(D,X)\big)\right] \\ +P_{DX}\left[(1-D)\frac{\hat{\pi}_{\mathcal{X}}(1{\vert\,}X)-\pi_{\mathcal{X}}(1{\vert\,}X)}{(1-\hat{\pi}_{\mathcal{X}}(1{\vert\,}X))(1-\pi_{\mathcal{X}}(1{\vert\,}X))}\big({\hat{\mu}_{\mathcal{X}}}(D,X)-\mu_{\mathcal{X}}(D,X)\big)\right] \end{align}\] for the corresponding regression and representer (estimators). By the tower property of expectation conditioning on \(X\) and by the definition of \(\pi_{\mathcal{X}}\), we arrive to the usual bias formulae \[\begin{align} R_n =&\, P_{X}\left[\frac{\hat{\pi}_{\mathcal{X}}(d{\vert\,}X)-\pi_{\mathcal{X}}(d{\vert\,}X)}{\hat{\pi}_{\mathcal{X}}(d{\vert\,}X)}\big({\hat{\mu}_{\mathcal{X}}}(d,X)-\mu_{\mathcal{X}}(d,X)\big)\right] & \text{ for } \mathbb{E}Y^d, \\ R_n =&\, P_{X}\left[\frac{\hat{\pi}_{\mathcal{X}}(1{\vert\,}X)-\pi_{\mathcal{X}}(1{\vert\,}X)}{\hat{\pi}_{\mathcal{X}}(1{\vert\,}X)}\big({\hat{\mu}_{\mathcal{X}}}(1,X)-\mu_{\mathcal{X}}(1,X)\big)\right] & \\ &+P_{X}\left[\frac{\hat{\pi}_{\mathcal{X}}(1{\vert\,}X)-\pi_{\mathcal{X}}(1{\vert\,}X)}{1-\hat{\pi}_{\mathcal{X}}(1{\vert\,}X)}\big({\hat{\mu}_{\mathcal{X}}}(0,X)-\mu_{\mathcal{X}}(0,X)\big)\right] & \text{ for } \mathbb{E}Y^1-\mathbb{E}Y^0. \end{align}\]

If we did not allow for the dependence on the parameter \(\gamma_{\mathcal{V}}\), our class of parameters 3 and conditions 4 , 5 would be a strict subset of those of [3]. Allowing for such a nonlinear but smooth dependence enables us to capture more parameters, such as the average treatment effect on the treated, which was indeed shown double robust by [4].

Example 3 (Average Treatment Effect on the Treated). Consider the setting of 2, but now take \(g(V,X)\mathrel{\vcenter{:}}= D\), and \(\gamma_{\mathcal{V}}(c)\mathrel{\vcenter{:}}= p_1\) for \(p_1\mathrel{\vcenter{:}}=\mathbb{E}D\); \(V_\mathit{2}\) is again unused and taken to be empty. Then \(\mathbb{E}\left[\left. Y^0\,\right\vert\, D=1\right]=\mathbb{E}D\mu_{\mathcal{X}}(0,X)/p_1\). Hence \(f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}})\mathrel{\vcenter{:}}= D\mu_{\mathcal{X}}(0,X)/p_1\) gives \(\chi(P_{VX})= \mathbb{E}\left[\left. Y^0\,\right\vert\, D=1\right]\), while \(f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}})\mathrel{\vcenter{:}}= D(\mu_{\mathcal{X}}(1,X)-\mu_{\mathcal{X}}(0,X))/p_1\) gives \(\chi(P_{VX})= \mathbb{E}\left[\left. Y^1-Y^0\,\right\vert\, D=1\right]\), with 4 , 5 , 6 , 56 holding.

The Riesz representer for \(\mathbb{E}\left[\left. Y^0\,\right\vert\, D=1\right]\) is \(r(d,x)=\frac{1-d}{p_1}\frac{\pi_{\mathcal{X}}(1{\vert\,}x)}{1-\pi_{\mathcal{X}}(1{\vert\,}x)},\) and for \(\mathbb{E}\left[\left. Y^1-Y^0\,\right\vert\, D=1\right]\), it is \(r(d,x)=\frac{d}{p_1}-\frac{1-d}{p_1}\frac{\pi_{\mathcal{X}}(1{\vert\,}x)}{1-\pi_{\mathcal{X}}(1{\vert\,}x)}.\) Both representers \(r\) depend on \(\gamma_{\mathcal{V}}(c)=p_1\).

The efficient influence function ?? , in the nonparametric model (with unknown \(\pi_{\mathcal{X}}\)), of \(\mathbb{E}\left[\left. Y^0\,\right\vert\, D=1\right]\), with \(\partial_\gamma f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c))=-D\mu_{\mathcal{X}}(0,X)/p_1^2\), is \[\begin{align} \tilde{\chi}(v,x)=&\, \frac{1-d}{p_1}\frac{\pi_{\mathcal{X}}(1{\vert\,}x)}{1-\pi_{\mathcal{X}}(1{\vert\,}x)}(y-\mu_{\mathcal{X}}(d,x))-\frac{d-p_1}{p_1}\chi(P_{VX})+D\mu_{\mathcal{X}}(0,X)/p_1 \\ &- \chi(P_{VX}) \\ =& \frac{1-d}{p_1}\frac{\pi_{\mathcal{X}}(1{\vert\,}x)}{1-\pi_{\mathcal{X}}(1{\vert\,}x)}(y-\mu_{\mathcal{X}}(0,x))+\frac{d}{p_1}\mu_{\mathcal{X}}(0,X) -\frac{d}{p_1}\chi(P_{VX}); \end{align}\] of \(\mathbb{E}\left[\left. Y^1-Y^0\,\right\vert\, D=1\right]\), with \(\partial_\gamma f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c))=-D(\mu_{\mathcal{X}}(1,X)-\mu_{\mathcal{X}}(0,X))/p_1^2\), is \[\begin{align} \tilde{\chi}(v,x)=&\,\left( \frac{d}{p_1}-\frac{1-d}{p_1}\frac{\pi_{\mathcal{X}}(1{\vert\,}x)}{1-\pi_{\mathcal{X}}(1{\vert\,}x)}\right) (y-\mu_{\mathcal{X}}(d,x))-\frac{d-p_1}{p_1}\chi(P_{VX}) \\ &+ \frac{d}{p_1}(\mu_{\mathcal{X}}(1,x)-\mu_{\mathcal{X}}(0,x))-\chi(P_{VX}) \\ =&\,\frac{d}{p_1}(y-\mu_{\mathcal{X}}(1,x))-\frac{1-d}{p_1}\frac{\pi_{\mathcal{X}}(1{\vert\,}x)}{1-\pi_{\mathcal{X}}(1{\vert\,}x)}(y-\mu_{\mathcal{X}}(0,x)) \\ &+ \frac{d}{p_1}(\mu_{\mathcal{X}}(1,x)-\mu_{\mathcal{X}}(0,x)) + \frac{d}{p_1}\chi(P_{VX}) \end{align}\] as in [44]. The remark on the nonparametric nature of the \((Y,D,X)\)-model in 2 applies here equally.

Consequently, the first term of the bias in 64 established by 1 for \(\mathbb{E}\left[\left. Y^0\,\right\vert\, D=1\right]\) is \[\begin{align} -P_{VX}(r-\hat{r})(\mu_{\mathcal{X}}-{\hat{\mu}_{\mathcal{X}}}) \\ = \frac{\hat{p}_1- p_1}{\hat{p}_1 p_1}P_X\left[\frac{(1-\hat{\pi}_{\mathcal{X}}(1{\vert\,}X))\pi_{\mathcal{X}}(1{\vert\,}X)}{1-\hat{\pi}_{\mathcal{X}}(1{\vert\,}X)}({\hat{\mu}_{\mathcal{X}}}(0,X)-\mu_{\mathcal{X}}(0,X))\right] \\ + \frac{p_1}{p_1\hat{p}_1}P_{X}\left[\frac{(\pi_{\mathcal{X}}(1{\vert\,}X)-\hat{\pi}_{\mathcal{X}}(1{\vert\,}X))({\hat{\mu}_{\mathcal{X}}}(0,X)-\mu_{\mathcal{X}}(0,X))}{1-\hat{\pi}_{\mathcal{X}}(1{\vert\,}X)}\right], \end{align}\] by the tower property of expectation conditioning on \(X\); for \(\mathbb{E}\left[\left. Y^1-Y^0\,\right\vert\, D=1\right]\), it is \[\begin{align} -P_{VX}(r-\hat{r})(\mu_{\mathcal{X}}-{\hat{\mu}_{\mathcal{X}}}) = \frac{\hat{p}_1 - p_1}{\hat{p}_1 p_1}P_X\left[\pi_{\mathcal{X}}(1{\vert\,}X)({\hat{\mu}_{\mathcal{X}}}(1,X)-\mu_{\mathcal{X}}(1,X))\right] \\ -\frac{\hat{p}_1- p_1}{\hat{p}_1 p_1}P_X\left[\frac{(1-\hat{\pi}_{\mathcal{X}}(1{\vert\,}X))\pi_{\mathcal{X}}(1{\vert\,}X)}{1-\hat{\pi}_{\mathcal{X}}(1{\vert\,}X)}({\hat{\mu}_{\mathcal{X}}}(0,X)-\mu_{\mathcal{X}}(0,X))\right] \\ - \frac{p_1}{p_1\hat{p}_1}P_{X}\left[\frac{(\pi_{\mathcal{X}}(1{\vert\,}X)-\hat{\pi}_{\mathcal{X}}(1{\vert\,}X))({\hat{\mu}_{\mathcal{X}}}(0,X)-\mu_{\mathcal{X}}(0,X))}{1-\hat{\pi}_{\mathcal{X}}(1{\vert\,}X)}\right]. \end{align}\] Suppose \(\hat{e}-e'=o_{P_{VX}}\left(1\right)\) and that \(P_{YDX} \partial_\gamma^2 f(Y,D,X,{\hat{\mu}_{\mathcal{X}}},\tilde{p}_1)=O_{{P_{VX}}}\left(1\right)\). Then for the bias to vanish as \(o_{P_{VX}}\left(n^{-1/2}\right)\), it suffices that \(\left\lVert{\Delta}\right\rVert_{L_1(P_{VX})}=o_{P_{VX}}\left(n^{-1/2}\right)\), where \(\Delta(x)\mathrel{\vcenter{:}}=({\hat{\mu}_{\mathcal{X}}}(0{\vert\,}x)-\mu_{\mathcal{X}}(0{\vert\,}x))(\pi_{\mathcal{X}}(1{\vert\,}x)-\hat{\pi}_{\mathcal{X}}(1{\vert\,}x))\), and \({\hat{\mu}_{\mathcal{X}}}\) be consistent because \(\hat{p}_1-p_1=O_{{P_{VX}}}\left(n^{-1/2}\right)\), provided \(1-\hat{\pi}_{\mathcal{X}}(1{\vert\,}x)\) is bounded away from zero.

Parameters with a more geometric interpretation are also included in our class.

Example 4 (Average Approximate Derivative). Let \(\mathfrak{X}=\mathbb{R}\), and \(V\mathrel{\vcenter{:}}= Y\) for a random variable \(Y\). Set \(m(V,X)\mathrel{\vcenter{:}}= Y\), and assume that \(\mu_{\mathcal{X}}(v_\mathit{1},x)=\mathbb{E}\left[\left. Y\,\right\vert\, V_\mathit{1}=v_\mathit{1},X=x\right]\) does not depend on \(v_\mathit{1}\), so we use \(\mu_{\mathcal{X}}(x)\) to refer to its value. Assume that the marginal distribution \(P_X\) of \(X\) admits a Lebesque density \(p_X\).

Fix a strictly positive constant \(\epsilon\in\mathbb{R}\). Setting \(f(v,x,\mu,\gamma)\mathrel{\vcenter{:}}=\frac{\mu(x+\epsilon)-\mu(x)}{\epsilon}\) and \(J_{P_{VX},\epsilon}(\mu)\mathrel{\vcenter{:}}=\mathbb{E}\left[\frac{\mu(X+\epsilon)-\mu(X)}{\epsilon}\right]\), \(\mu\in L_2(P_X)\), yields the parameter \[\begin{align} \chi(P_{VX}) = J_{P_{VX},\epsilon}(\mu_{\mathcal{X}}) = \mathbb{E}\left[\frac{\mu_{\mathcal{X}}(X+\epsilon)-\mu_{\mathcal{X}}(X)}{\epsilon}\right]=\int_{-\infty}^\infty \frac{\mu_{\mathcal{X}}(x+\epsilon)-\mu_{\mathcal{X}}(x)}{\epsilon}p_X(x)\mathrm{d}x. \end{align}\] For small \(\epsilon\), the integrand represents an approximate derivative of \(\mu_{\mathcal{X}}\), although we do not actually require that \(\mu_{\mathcal{X}}\) be differentiable, unlike [4].

Assume that \((P_{VX},\epsilon)\) satisfies \[\begin{align} b(P_{VX},\epsilon)\mathrel{\vcenter{:}}=\int_{-\infty}^\infty \left(\frac{p_X(x-\epsilon)}{p_X(x)}\right)^2 p_X(x)\mathrm{d}x<\infty. \label{priv:eq:b95bound} \end{align}\qquad{(13)}\] Then \(J_{P_{VX},\epsilon}\) has Riesz representer \(r(x)=\frac{p_X(x-\epsilon)-p_X(x)}{\epsilon p_X(x)}\), whose dependence on \((P_{VX},\epsilon)\) is suppressed in notation. The last display ensures that \(J_{P_{VX},\epsilon}\) is continuous so that the Riesz representation theorem applies. Indeed, by a change of variables, and the Cauchy–Schwarz inequality, for \(\mu,\tilde{\mu}\in L_2(P_X)\), \[\begin{align} |J_{P_{VX},\epsilon}(\mu)-J_{P_{VX},\epsilon}(\tilde{\mu})|\leq \mathbb{E}\left[|\mu(X)-\tilde{\mu}(X)|\frac{p_X(X-\epsilon)}{p_X(X)}\right]+\left\lVert{\mu-\tilde{\mu}}\right\rVert_{L_2(P_X)} \\ \leq \left(\sqrt{b(P_{VX},\epsilon)}+1\right)\left\lVert{\mu-\tilde{\mu}}\right\rVert_{L_2(P_X)}. \end{align}\] Hence, 4 , 5 , and 6 hold. In addition, if \((P_{VX},\epsilon)\) satisfies \(\left\lVert{\frac{p_X(\cdot-\epsilon)}{p_X(\cdot)}}\right\rVert_{\infty}<\infty\) — which is stronger than ?? — , then 56 holds by similar arguments.

Example 5 (Ray-Type Integral). Let \(\mathfrak{X}=\mathbb{R}^d\), and \(V\mathrel{\vcenter{:}}=(Y,W)\), for a random variable \(Y\), and a random vector \(W\) with values \(\mathbb{R}^d\) and Lebesgue density \(p_W\). Set \(m(V,X)\mathrel{\vcenter{:}}= Y\), and assume that \(\mu_{\mathcal{X}}(v_\mathit{1},x)\equiv\mu_{\mathcal{X}}(x)\) does not depend on \(v_\mathit{1}\). Suppose that \(P_X\) has a Lebesgue density \(p_X\).

Fix constants \(0<t_0<t_1\). Setting \(f(v,x,\mu,\gamma)\mathrel{\vcenter{:}}=\int_{t_{0}}^{t_{1}} \mu(tw)\left\lVert{w}\right\rVert_{2}\,\mathrm{d}t\) and \(J_{P_{VX},t_{0},t_{1}}(\mu)\mathrel{\vcenter{:}}=\mathbb{E}\int_{t_0}^{t_1} \mu(tW)\left\lVert{W}\right\rVert_{2}\,\mathrm{d}t\), \(\mu\in L_2(P_X)\), where \(\left\lVert{w}\right\rVert_{2}=\sqrt{\sum_{j=1}^d w_j^2}\) is the Euclidean norm on \(\mathbb{R}^d\), yields the parameter \[\begin{align} \chi(P_{VX})\mathrel{\vcenter{:}}= J_{P_{VX},t_0,t_1}(\mu_{\mathcal{X}})= \mathbb{E}\int_{t_0}^{t_1} \mu_{\mathcal{X}}(tW)\left\lVert{W}\right\rVert_{2}\,\mathrm{d}t=\int_{\mathbb{R}^d}\int_{t_0}^{t_1} \mu_{\mathcal{X}}(tw)\left\lVert{w}\right\rVert_{2}\,\mathrm{d}t p_W(w)\,\mathrm{d}w. \end{align}\] The quantity \(f(v,x,\mu,\gamma)\) is the integral of \(\mu\) along the curve \(\zeta_{w}: [t_0,t_1]\to\mathbb{R}^d\), \(\zeta_{w}(t)\mathrel{\vcenter{:}}= t w\), the line segment between \(t_0w\) and \(t_1w\), and \(J_{P_{VX},t_0,t_1}(\mu)\) is its average across draws from \(p_W\). With \(t_0\) close to zero and \(t_1\) to infinity, \(f(v,x,\mu,\gamma)\) approximates the ray emanating from the origin in the direction of \(w\), but care must be taken to ensure integrability: if \((P_{VX},t_0,t_1)\) is such that \[\begin{align} \mathbb{R}^d \ni x\mapsto \frac{\left\lVert{x}\right\rVert_{2}}{p_X(x)}\int_{t_0}^{t_1}p_W\left(\frac{x}{t}\right)\frac{1}{t^{d+1}}\,\mathrm{d}t \in L_2(P_X), \end{align}\] then the map in the display is the Riesz representer of \(J_{P_{VX},t_0,t_1}\), with 4 and 5 holding. (Condition 6 is trivially satisfied.) This follows from \[\begin{align} J_{P_{VX},t_0,t_1}(\mu)=\int_{\mathbb{R}^d}\int_{t_0}^{t_1} \mu(tw)\left\lVert{w}\right\rVert_{2}\,\mathrm{d}t\, p_W(w)\,\mathrm{d}w = \int_{t_0}^{t_1}\int_{\mathbb{R}^d}\mu(u)\left\lVert{u/t}\right\rVert_{2}p_W(u/t)\frac{1}{t^d}\,\mathrm{d}u \,\mathrm{d}t \\ = \int_{\mathbb{R}^d}\mu(u)\frac{\left\lVert{u}\right\rVert_{2}}{p_X(u)}\int_{t_0}^{t_1}p_W(u/t)\frac{1}{t^{d+1}}\,\mathrm{d}t p_X(u) \,\mathrm{d}u = \mathbb{E}\mu(X)\frac{\left\lVert{X}\right\rVert_{2}}{p_X(X)}\int_{t_0}^{t_1}p_W(X/t)\frac{1}{t^{d+1}}\,\mathrm{d}t, \end{align}\] where the second equality is by a change of variables \((wt, t)\mapsto (u, t)\) with a Jacobian whose determinant is \(t^{-d}\).

Example 6 (Line Integral). Instead of taking an integral along a ray as in 5, we can integrate along the curve \(\zeta_{w_{\mathrm{b}},w_\mathrm{e}}:[0,1]\to\mathbb{R}^d\), \(\zeta_{w_{\mathrm{b}},w_{\mathrm{e}}}(t)\mathrel{\vcenter{:}}=(1-t) w_{\mathrm{b}}+tw_{\mathrm{e}}\), which is the straight line connecting \(w_{\mathrm{b}}\) and \(w_{\mathrm{e}}\) in \(\mathbb{R}^d\).

Let \(\mathfrak{X}=\mathbb{R}^d\), and \(V\mathrel{\vcenter{:}}=(Y, W)\), for a random variable \(Y\), and a random vector \(W=(W_{\mathrm{b}}, W_{\mathrm{e}})\) with values \(\mathbb{R}^d\times\mathbb{R}^d\) and Lebesgue density \(p_W\). Set \(m(V,X)\mathrel{\vcenter{:}}= Y\), and assume that \(\mu_{\mathcal{X}}(v_\mathit{1},x)\equiv\mu_{\mathcal{X}}(x)\) does not depend on \(v_\mathit{1}\). Suppose that \(P_X\) has a Lebesgue density \(p_X\). Setting \(f(v,x,\mu,\gamma)\mathrel{\vcenter{:}}=\int_{0}^1 \mu((1-t) w_{\mathrm{b}}+tw_{\mathrm{e}})\left\lVert{w_\mathrm{e}-w_\mathrm{b}}\right\rVert_{2}\,\mathrm{d}t\) and \(J_{P_{VX}}(\mu)=\mathbb{E}\int_{0}^1 \mu((1-t) W_{\mathrm{b}}+tW_{\mathrm{e}})\left\lVert{W_\mathrm{e}-W_\mathrm{b}}\right\rVert_{2}\,\mathrm{d}t\) yields the parameter \[\begin{align} \chi(P_{VX})= J_{P_{VX}}(\mu_{\mathcal{X}})= \mathbb{E}\int_{0}^1 \mu_{\mathcal{X}}((1-t) W_{\mathrm{b}}+tW_{\mathrm{e}})\left\lVert{W_\mathrm{e}-W_\mathrm{b}}\right\rVert_{2}\,\mathrm{d}t \\ = \int_{\mathbb{R}^d\times\mathbb{R}^d} \int_{0}^1 \mu_{\mathcal{X}}((1-t) w_{\mathrm{b}}+tw_{\mathrm{e}})\left\lVert{w_\mathrm{e}-w_\mathrm{b}}\right\rVert_{2}\,\mathrm{d}t\, p_W(w_\mathrm{b}, w_\mathrm{e}) \,\mathrm{d}w_\mathrm{b} \,\mathrm{d}w_\mathrm{e}. \end{align}\] This is the average line integral of \(\mu_{\mathcal{X}}\) across draws of random endpoints \((W_{\mathrm{b}}, W_{\mathrm{e}})\) from \(p_W\). If \(P_{VX}\) is such that \[\begin{align} \mathbb{R}^d\ni x\mapsto \frac{1}{p_X(x)}\int_{0}^1 \int_{\mathbb{R}^d}\left\lVert{\Delta}\right\rVert_{2}p_W(x-t\Delta, x+(1-t)\Delta)\,\mathrm{d}\Delta \,\mathrm{d}t \in L_2(P_{X}), \end{align}\] then the map in the display is the Riesz representer of \(J_{P_{VX}}\). This follows from \[\begin{align} \int_{\mathbb{R}^d\times\mathbb{R}^d} \int_{0}^1 \mu((1-t) w_{\mathrm{b}}+tw_{\mathrm{e}})\left\lVert{w_\mathrm{e}-w_\mathrm{b}}\right\rVert_{2}\,\mathrm{d}t p_W(w_\mathrm{b}, w_\mathrm{e}) \,\mathrm{d}w_\mathrm{b} \,\mathrm{d}w_\mathrm{e} \\ =\int_{\mathbb{R}^d\times\mathbb{R}^d} \int_{0}^1 \mu(u)\left\lVert{\Delta}\right\rVert_{2}p_W(u-t\Delta, u+(1-t)\Delta)|\mathrm{det}J(u,\Delta,t)|\,\mathrm{d}t \,\mathrm{d}\Delta \,\mathrm{d}u, \end{align}\] by a change of variables \[(w_\mathrm{b},w_\mathrm{e},t)\mapsto ((1-t) w_{\mathrm{b}}+tw_{\mathrm{e}}, w_\mathrm{e}-w_\mathrm{b},t)\eqqcolon(u,\Delta, t)\in\mathbb{R}^d\times\mathbb{R}^d\times[0,1].\] Above, \(\mathrm{det}J(u,\Delta,t)\) is the determinant of the Jacobian matrix, in \(\mathbb{R}^{(2d+1)\times(2d+1)}\), of \((u,\Delta,t)\mapsto (g_1, g_2,g_3)(u,\Delta,t)\) \(\mathrel{\vcenter{:}}=(u-t\Delta, u+(1-t)\Delta,t)= (w_\mathrm{b},w_\mathrm{e},t)\), which is \[\begin{align} J(u,\Delta,t)=\begin{bmatrix} \mathrm{D}_ u g_1 & \mathrm{D}_\Delta g_1 & \mathrm{D}_t g_1 \\ \mathrm{D}_ u g_2 & \mathrm{D}_\Delta g_2 & \mathrm{D}_t g_2 \\ \mathrm{D}_ u g_3 & \mathrm{D}_\Delta g_3 & \mathrm{D}_t g_3 \end{bmatrix}(u,\Delta,t) = \begin{bmatrix} \mathrm{I}_d & -t \mathrm{I}_d & -\Delta \\ \mathrm{I}_d & (1-t) \mathrm{I}_d & -\Delta \\ 0_d^\intercal & 0_d^\intercal & 1 \\ \end{bmatrix}, \end{align}\] where \(\mathrm{I}_d\) is the identity matrix in \(\mathbb{R}^d\), and \(0_d^\intercal \in\mathbb{R}^{1\times d}\) is the transposed zero vector. Hence, \(|\mathrm{det}J(u,\Delta,t)|=|\mathrm{det}\begin{bmatrix} \mathrm{I}_d & -t \mathrm{I}_d \\ \mathrm{I}_d & (1-t) \mathrm{I}_d \end{bmatrix}|=1\), because of how the determinant of block matrices is computed. While not pursued here, these arguments could be extended to higher order Bézier curves with random control points \(W_1,\ldots, W_k\), or other curves indeed.

A special subset of our class is a set of parameters with known Riesz representers. Their estimators admit favourable efficiency properties.

Example 7 (Known Riesz Representer). In the general context of 3, suppose that \(f(v,x,\mu,\gamma)=\mu(v_\mathit{1},x)b(v_\mathit{1},x)\), independent of \(\gamma\), for a known* function \(b:\mathfrak{V}_\mathit{1}\times\mathfrak{X}\to\mathbb{R}\). If \(b\in L_2(P_{V_{\mathit{1}}X})\), then, trivially, \(b\) is the Riesz representer of \(L_2(P_{V_{\mathit{1}}X})\ni\mu\mapsto \mathbb{E}\mu(V_\mathit{1},X)b(V_\mathit{1},X)\).*

In this case, 1 is automatically satisfied, and asymptotic efficiency as per 1 follows merely under the consistency requirements of 2 for an estimator of \(\mu_{\mathcal{X}}\).

Situations with a known Riesz representer may exist:

Example 8 (Time Discounting). Consider a simple economics model with two time periods. Furthest in the future, in Period 2, a stock value \(Y\geq 0\) is going to be revealed. In Period 1, a random element \(X\) is realised, which potentially helps predict the stock value through \(\mu_{\mathcal{X}}(x)\mathrel{\vcenter{:}}=\mathbb{E}\left[\left. Y\,\right\vert\, X=x\right]\). In Period 1, an agent owns a stock, and signs a binding contract — after \(X=x\) is realised — that she would sell the stock at price \(\mu_{\mathcal{X}}(x)\) in Period 2. (This is a “fair bet,” as her a priori* expected profit given \(X=x\), \(\mathbb{E}\left[\left. \mu_{\mathcal{X}}(x)-Y\,\right\vert\, X=x\right]\), is zero.)*

Viewed from Period 1, the present value of the agent’s Period-2 income is \(\frac{1}{1+\kappa(x)}\mu_{\mathcal{X}}(x)\), where \(\kappa:\mathfrak{X}\to[0,\infty)\) is a discount rate, a nonrandom, known* function; the larger the \(\kappa\), the more impatient she is. The element \(X\) can be thought of as all relevant information available in Period 1.*

The parameter \[\begin{align} \chi(P_{VX})\mathrel{\vcenter{:}}=\mathbb{E}\frac{1}{1+\kappa(X)}\mu_{\mathcal{X}}(X) \end{align}\] is the expected (before \(X\) is observed) present value of the agent’s Period-2 income* viewed from Period 1. (The expected present value of her profit, \(\mathbb{E}\frac{1}{1+\kappa(X)}(\mu_{\mathcal{X}}(X)-Y)\), is zero.) It is of the form 3 for \(V\mathrel{\vcenter{:}}= Y\) and \(m(V,X)=Y\).*

As \(\kappa\geq 0\), \(x\mapsto \frac{1}{1+\kappa(x)}\in L_2(P_X)\) is the Riesz-representer of \(\mu\mapsto \mathbb{E}\frac{1}{1+\kappa(X)}\mu(X)\), and 4 and 5 both hold. This is an instance of a known Riesz representer in 7. Consequently, a merely consistent estimator of \(\mu_{\mathcal{X}}\) suffices for asymptotically efficient estimation of \(\chi(P_{VX})\). This is attractive as it affords “complicated” predictors \(X\), at essentially no cost (although the asymptotic variance \(P_{VZ}\tilde{\psi}^2\) may change with different choices of \(X\)).

We may extend our class to be indexed by multiple regressions.

Example 9 (Multiple Regressions). The parameter class 3 in 3 can easily be extended to feature linear dependence on multiple infinite-dimensional regressions \[\begin{align} \label{priv:eq:regression95generalisation} \begin{aligned} \mu_{\mathcal{X},1}(v_{\mathit{1}},x)&\mathrel{\vcenter{:}}=\mathbb{E}\left[\left. m_1(V,X)\,\right\vert\, V_{\mathit{1}}=v_{\mathit{1}},X=x\right], \\ \vdots & \\ \mu_{\mathcal{X},K}(v_{\mathit{1}},x)&\mathrel{\vcenter{:}}=\mathbb{E}\left[\left. m_K(V,X)\,\right\vert\, V_{\mathit{1}}=v_{\mathit{1}},X=x\right], \end{aligned} \end{align}\qquad{(14)}\] where the \(m_k: \mathfrak{V}\times \mathfrak{X}\to\mathbb{R}\). Indeed, \(H_K\mathrel{\vcenter{:}}=(L_2(P_{V_\mathit{1}X}))^K\) is a Hilbert space with inner product \(\langle\mu,\nu\rangle_{H_K}\mathrel{\vcenter{:}}=\sum_{k=1}^K\langle\mu_k,\nu_k\rangle_{L_2(P_{V_\mathit{1}X})}=\sum_{k=1}^K \int \mu_k\nu_k\,\mathrm{d}P_{V_\mathit{1}X}\). Then we define \(f:\mathfrak{V}\times\mathfrak{X}\times H_K \times\Gamma\to\mathbb{R}\) requiring the linearity of \(H_K \ni \mu \mapsto f(V,X,\mu,\gamma)\) \(P_{VX}\)-almost surely, for all \(\gamma\in\Gamma\) (meaning that for all \(\beta\in\mathbb{R}\), \(f(V,X,(\beta\mu_1,\ldots,\beta\mu_K),\gamma)=\beta f(V,X,(\mu_1,\ldots,\mu_K),\gamma)\), \(P_{VX}\)-almost surely, for all \(\gamma\in\Gamma\)), and the continuity of \(H_K\ni \mu \mapsto P_{VX}f(V,X,\mu,\gamma)\) for all \(\gamma\in\Gamma\) for the distance \(\left\lVert{\mu-\nu}\right\rVert_{H_K}\mathrel{\vcenter{:}}=\sqrt{\langle\mu-\nu,\mu-\nu\rangle_{H_K}}=\sqrt{\sum_{k=1}^K \int (\mu_k-\nu_k)^2\,\mathrm{d}P_{V_\mathit{1}X}}\). Then the Riesz representation theorem applies equally: for all \((P_{VX},\gamma)\) there exists a unique element \(r_{\gamma}=(r_{\gamma,1},\ldots,r_{\gamma,K})\in H_K\) such that \[\begin{align} P_{VX}f(V,X,\mu,\gamma)=\sum_{k=1}^K P_{V_\mathit{1}X}(r_{\gamma,k}\mu_k) \quad\text{ for all } \mu \in H_K. \end{align}\] This yields results very similar to the single-infinite-dimensional-regression case.

Indeed, let \(\mu_{\mathcal{X}}\mathrel{\vcenter{:}}=(\mu_{\mathcal{X},1}, \ldots, \mu_{\mathcal{X},K})\in H_K\). Then efficient influence function of \(\chi(P_{VX})\mathrel{\vcenter{:}}= P_{VX}f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c))\), in 1 becomes \[\begin{align} \tilde{\chi}(v,x) =&\, \sum_{k=1}^K r_{k}(v_{\mathit{1}},x)(m_k(v,x)-\mu_{\mathcal{X},k}(v_{\mathit{1}},x))\\ &+\frac{\mathbb{1}_{v_{\mathit{2}}=c}}{p_{V_{\mathit{2}}}(c)}(g(v,x)-\gamma_{\mathcal{V}}(c))\mathbb{E}\partial_\gamma f(V,X, \mu_{\mathcal{X}}, \gamma_{\mathcal{V}}(c)) \\ &+ f(v,x, \mu_{\mathcal{X}}, \gamma_{\mathcal{V}}(c))-\chi(P_{VX}), \end{align}\] where \(r=(r_{1},\ldots,r_{K})\in H_K\) is the Riesz representer of \(H_K\ni \mu \mapsto P_{VX}f(V,X,\mu,\gamma_{\mathcal{V}}(c))\).

Correspondingly, the rate-double-robustness property of 1 becomes \[\begin{align} \chi'-\chi(P_{VX})+P_{VX}\tilde{\chi}' =&\, \sum_{k=1}^{K}P_{VX}(r_{k}'-r_{k})(\mu_{\mathcal{X},k}-\mu_{\mathcal{X},k}') \\ &+(\gamma_{\mathcal{V}}(c)-\gamma_{\mathcal{V}}'(c))\left(\frac{p_{V_{\mathit{2}}}(c)}{p_{V_{\mathit{2}}}'(c)}e''-e'\right) \\ &-(\gamma_{\mathcal{V}}(c)-\gamma_{\mathcal{V}}'(c))^2\frac{P_{VX}\partial_\gamma^2 f(V,X,\mu_{\mathcal{X}}',\widetilde{\gamma_{\mathcal{V}}(c)})}{2}, \end{align}\] with \[\begin{align} \tilde{\chi}'(v,x)\mathrel{\vcenter{:}}=&\, \sum_{k=1}^Kr_{k}'(v_{\mathit{1}},x)(m_k(v,x)-\mu_{\mathcal{X},k}'(v_{\mathit{1}},x))+\frac{\mathbb{1}_{v_{\mathit{2}}=c}}{p_{V_{\mathit{2}}}'(c)}(g(v,x)-\gamma_{\mathcal{V}}'(c))e'' \\ &+ f(v,x, \mu_{\mathcal{X}}', \gamma_{\mathcal{V}}'(c))-\chi', \end{align}\] where the \(r_{k}',\mu_{\mathcal{X},k}'\) are both arbitrary elements of \(L_2(P_{V_\mathit{1}X})\) with \(\mu_{\mathcal{X}}'\mathrel{\vcenter{:}}=(\mu_{\mathcal{X},1}',\ldots,\mu_{\mathcal{X},K}')\); \(p_{V_{\mathit{2}}}'(c),\gamma_{\mathcal{V}}'(c),\chi',e''\in\mathbb{R}\) are arbitrary; \(e'\mathrel{\vcenter{:}}= P_{VX}\partial_\gamma f(V,X,\mu_{\mathcal{X}}',\gamma_{\mathcal{V}}'(c));\) and \(\widetilde{\gamma_{\mathcal{V}}(c)}\) is some value between \(\gamma_{\mathcal{V}}(c)\) and \(\gamma_{\mathcal{V}}'(c)\). That is, the rate-double-robustness property holds, with pairwise rate-tradeoffs between the \(r_{k}'\) and \(\mu_{\mathcal{X},k}'\).

Moreover, if \(\mathfrak{V}_{1}=\mathfrak{V}_{\mathit{1},1}\times\ldots \times\mathfrak{V}_{\mathit{1}, L}\) and \(\mathfrak{X}=\mathfrak{X}_1\times\ldots \mathfrak{X}_K\), then we could generalise ?? further by defining \(L\times K\) regressions \[\begin{align} \mu_{\mathcal{X},l,k}(v_{\mathit{1},l},x_k)\mathrel{\vcenter{:}}=\mathbb{E}\left[\left. m_{lk}(V,X)\,\right\vert\, V_{\mathit{1},l}=v_{\mathit{1},l}, X_{k}=x_k\right],\, (v_{\mathit{1},l},x_k) \in \mathfrak{V}_{\mathit{1},l}\times \mathfrak{X}_k, \end{align}\] with \(m_{lk}:\mathfrak{V}\times\mathfrak{X}\to\mathbb{R}\), for \((l,k)\in [L]\times [K]\), and by considering the Hilbert space \(H\mathrel{\vcenter{:}}=\otimes_{l=1}^L\otimes_{k=1}^K L_2(P_{V_{\mathit{1},l}X_k})\) with the inner product \(\langle\mu,\nu\rangle_{H}\mathrel{\vcenter{:}}=\sum_{l=1}^L\sum_{k=1}^K\langle\mu_{l,k},\nu_{l,k}\rangle_{L_2(P_{V_{\mathit{1},l}X_k})}=\sum_{k=1}^K \int \mu_{l,k}\nu_{l,k}\,\mathrm{d}P_{V_{\mathit{1},l}X_k}\). The efficient influence function and the rate-double-robustness would be preserved with the obvious modifications.

Finally, we remark that for \(\mathfrak{V}_{\mathit{2}}=\mathfrak{V}_{\mathit{2},1}\times\ldots\times \mathfrak{V}_{\mathit{2},K}\), dependence on multiple low-dimensional regressions \[\begin{align} \gamma_{\mathcal{V},k}(v_{\mathit{2}})&\mathrel{\vcenter{:}}=\mathbb{E}\left[\left. g_k(V,X)\,\right\vert\, V_{\mathit{2},k}=v_{\mathit{2},k}\right],\quad v_{\mathit{2},k}\in\mathfrak{V}_{\mathit{2},k},\,\, g_k:\mathfrak{V}\times\mathfrak{X}\to\mathbb{R},\,\, k\in[K], \end{align}\] can be accommodated as well. This entails a vector-valued derivative \(\partial_\gamma f\) and corresponding vector \(e\), with number of entries determined by the \(|\mathfrak{V}_{\mathit{2},k}|\).

9.2 Proofs↩︎

This section proves [priv:eqprop:eif,priv:thm:dr] in 3.

We first note that 4 , 5 imply that for all \(\gamma\in\Gamma\), there exists a unique function \(r_{P_{VX},\gamma}:\mathfrak{V}_{\mathit{1}}\times\mathfrak{X}\to\mathbb{R}\), \(r_{P_{VX},\gamma}\in L_2(P_{V_{\mathit{1}}X})\), such that \[\begin{align} \mathbb{E}f(V,X, \mu, \gamma)=\mathbb{E}r_{P_{VX},\gamma}(V_{\mathit{1}},X)\mu(V_{\mathit{1}},X)\quad \text{ for all } \mu\in L_2(P_{V_{\mathit{1}}X}), \end{align}\] where the expectations are taken with respect to \(P_{VX}\). This is the consequence of the Riesz representation theorem, and \(r_{\gamma,P_{VX}}\) is called the Riesz representer of \(\mu\mapsto\mathbb{E}f(V,X, \mu, \gamma)\).

Proof of 1. First, note that by the definition of \(\mu_{\mathcal{X}},\gamma_{\mathcal{V}}\), \[\begin{align} \mathbb{E}r(V_{\mathit{1}},X)(m(V,X)-\mu_{\mathcal{X}}(V_{\mathit{1}},X)) \\ = \mathbb{E}\mathbb{E}\left[\left. r(V_{\mathit{1}},X)(m(V,X)-\mu_{\mathcal{X}}(V_{\mathit{1}},X))\,\right\vert\, V_{\mathit{1}},X\right] \\ =\mathbb{E}r(V_{\mathit{1}},X)\mu_{\mathcal{X}}(V_{\mathit{1}},X)-\mathbb{E}r(V_{\mathit{1}},X)\mu_{\mathcal{X}}(V_{\mathit{1}},X)=0, \\ \mathbb{E}\mathbb{1}_{V_{\mathit{2}}=c}(g(V,X)-\gamma_{\mathcal{V}}(c))= \mathbb{E}\left[\mathbb{1}_{V_{\mathit{2}}=c}g(V,X)\right]-p_{V_{\mathit{2}}}(c)\gamma_{\mathcal{V}}(c) \\ =p_{V_{\mathit{2}}}(c)\mathbb{E}\left[\left. g(V,X)\,\right\vert\, V_{\mathit{2}}=c\right]-p_{V_{\mathit{2}}}(c)\gamma_{\mathcal{V}}(c) = 0, \end{align}\] because \(V_{\mathit{2}}\) is distributed on a finite set. Hence, by the definition of \(\chi(P_{VX})\) in 3 , \(\mathbb{E}\tilde{\chi}=0\). Because \(r\in L_2(P_{V_{\mathit{1}}X})\), 7 implies that \(\tilde{\chi}\) is in the tangent set \(L_2^0(P_{VX})\) of the nonparametric model.

Consider now a regular submodel \(t\mapsto P_{VX,t}\) for \(t\in\mathbb{R}\) in the neighbourhood of zero with \(P_{VX,0}=P_{VX}\) and the property that \(\frac{\mathrm{d}}{\mathrm{d}t}\big\vert_{t=0} P_{VX,t}(\mathrm{d}v,\mathrm{d}x)=s(v,x)P_{VX}(\mathrm{d}v,\mathrm{d}x)\) for the score function \(s\) running through the tangent set \(L_2^0(P_{VX})\); for example \(\mathrm{d}P_{VX}=\exp(ts)(P_{VX}\exp(ts))^{-1}\mathrm{d}P_{VX}\) as in [45]. See [38] for regular submodels. For ?? , it suffices to show \(\frac{\mathrm{d}}{\mathrm{d}t}\big\vert_{t=0}\chi(P_{VX,t})=\mathbb{E}\tilde{\chi }s\). Assuming that differentiation and expectation commutes, we have \[\begin{align} \frac{\mathrm{d}}{\mathrm{d}t}\big\vert_{t=0}\chi(P_{VX,t})=&\,\frac{\mathrm{d}}{\mathrm{d}t}\big\vert_{t=0}\int_{\mathfrak{V}\times\mathfrak{X}} f(v,x,\mu_{\mathcal{X},t},\gamma_{\mathcal{V},t}(c))\,\mathrm{d}P_{VX,t}(v,x) \\ =&\, \int_{\mathfrak{V}\times\mathfrak{X}} \frac{\mathrm{d}}{\mathrm{d}t}\big\vert_{t=0}f(v,x,\mu_{\mathcal{X},t},\gamma_{\mathcal{V},t}(c))\,\mathrm{d}P_{VX}(v,x) \\ &+ \mathbb{E}f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c))s(V,X). \end{align}\] Here, \[\begin{align} \int_{\mathfrak{V}\times\mathfrak{X}} \frac{\mathrm{d}}{\mathrm{d}t}\big\vert_{t=0}f(v,x,\mu_{\mathcal{X},t},\gamma_{\mathcal{V},t}(c))\,\mathrm{d}P_{VX}(v,x) \\ = \frac{\mathrm{d}}{\mathrm{d}t}\big\vert_{t=0}\mathbb{E}f(V,X,\mu_{\mathcal{X},t},\gamma_{\mathcal{V}}(c)) +\frac{\mathrm{d}}{\mathrm{d}t}\big\vert_{t=0}\mathbb{E}f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V},t}(c)). \end{align}\] By 10 , \(\mathbb{E}f(V,X,\mu_{\mathcal{X},t},\gamma_{\mathcal{V}}(c))=\mathbb{E}r(V_{\mathit{1}},X)\mu_{\mathcal{X},t}(V_{\mathit{1}},X)\). Taking derivatives, the properties of (conditional) expectation give \[\begin{align} \label{priv:eq:mux95t95derivative} \begin{aligned} \frac{\mathrm{d}}{\mathrm{d}t}\big\vert_{t=0}\mu_{\mathcal{X},t}(v_{\mathit{1}},x)&=\mathbb{E}\left[\left. (m(V,X)-\mu_{\mathcal{X}}(v_{\mathit{1}},x))s(V,X)\,\right\vert\, V_{\mathit{1}}=v_{\mathit{1}},X=x\right], \\ \frac{\mathrm{d}}{\mathrm{d}t}\big\vert_{t=0}\gamma_{\mathcal{V},t}(c)&=\mathbb{E}\left[\left. (g(V,X)-\gamma_{\mathcal{V}}(c))s(V,X)\,\right\vert\, V_{\mathit{2}}=c\right] \\ &= \mathbb{E}\frac{\mathbb{1}_{V_{\mathit{2}}=c}}{p_{V_{\mathit{2}}}(c)}(g(V,X)-\gamma_{\mathcal{V}}(c))s(V,X). \end{aligned} \end{align}\tag{35}\] This can be seen in a few steps. First, given a collection of coordinates \(V_\mathit{j}\) of \(V\), we can assume without loss of generality that \(V=(V_\mathit{j},V_{-\mathit{j}})\) for another collection of coordinates \(V_{-\mathit{j}}\) of \(V\). Second, \(P_{V_{-\mathit{j}}{\vert\,}V_{\mathit{j}}X}(B {\vert\,}v_{\mathit{j}},x)=\frac{\mathrm{d}P_{V_{-\mathit{j}}V_{\mathit{j}}X}(B,\cdot,\cdot)}{\mathrm{d}P_{V_{\mathit{j}}X}}(v_{\mathit{j}},x)\) can be verified to be the conditional distribution of \(V_{-\mathit{j}}\) given \((V_{\mathit{j}},X)=(v_{\mathit{j}},x)\) as in the proof of 4, noting that \(P_{V_{-\mathit{j}}V_{\mathit{j}}X}(B,\cdot,\cdot)\ll P_{V_{\mathit{j}}X}\) for all \(B\in\mathscr{F}_{{\mathfrak{V}}_{\mathit{j}}}\). Third, if \(\pi_t\mathrel{\vcenter{:}}=\frac{\mathrm{d}\lambda_t}{\mathrm{d}\nu_t}\) for valid submodels \(t\mapsto(\lambda_t,\nu_t)\) of measures, then \(\partial_t \pi_t=\frac{\mathrm{d}\partial_t \lambda_t}{\mathrm{d}\nu_t}-(\frac{\mathrm{d}\lambda_t}{\mathrm{d}\nu_t})(\frac{\mathrm{d}\partial_t \nu_t}{\mathrm{d}\nu_t})\) provided the densities on the right are well defined. Fourth, the given submodel \(t\mapsto P_{VX,t}\) induces the marginal \(\mathrm{d}P_{V_{\mathit{j}}X,t}(v_{\mathit{j}},x)=\int_{\mathfrak{V}_{\mathit{j}}} s(v_{-\mathit{j}},v_{\mathit{j}},x) P_{V_{-\mathit{j}}V_{\mathit{j}}X}(\mathrm{d}v_{-\mathit{j}}, \mathrm{d}v_{\mathit{j}}, \mathrm{d}x)\). Finally, apply these steps to obtain \[\begin{align} \frac{\mathrm{d}\partial_t P_{V_{-\mathit{j}},V_{\mathit{j}},X,t}(\mathrm{d}v_{-\mathit{j}},\cdot,\cdot)\big\vert_{t=0}}{\mathrm{d}P_{V_{\mathit{j}}X}}(v_{\mathit{j}},x)=s(v_{-\mathit{j}},v_{\mathit{j}}, x)P_{V_{-\mathit{j}}{\vert\,}V_{\mathit{j}}X}(\mathrm{d}v_{-\mathit{j}} {\vert\,}v_{\mathit{j}},x), \\ \frac{\mathrm{d}\partial_t P_{V_{\mathit{j}}X,t}\big\vert_{t=0}}{\mathrm{d}P_{V_{\mathit{j}X}}}(v_{\mathit{j}},x)=\int s(v_{-\mathit{j}},v_{\mathit{j}}, x)P_{V_{-\mathit{j}}{\vert\,}V_{\mathit{j}}X}(\mathrm{d}v_{-\mathit{j}} {\vert\,}v_{\mathit{j}},x). \end{align}\]

The display 35 then implies \[\begin{align} \frac{\mathrm{d}}{\mathrm{d}t}\big\vert_{t=0}\mathbb{E}r(V_{\mathit{1}},X)\mu_{\mathcal{X},t}(V_{\mathit{1}},X)\\ = \mathbb{E}\left[r(V_{\mathit{1}},X)\mathbb{E}\left[\left. \left(m(V,X)-\mu_{\mathcal{X}}(V_{\mathit{1}},X)\right)s(V,X)\,\right\vert\, V_{\mathit{1}},X\right]\right] \\ = \mathbb{E}r(V_{\mathit{1}},X)\left(m(V,X)-\mu_{\mathcal{X}}(V_{\mathit{1}},X)\right)s(V,X),\\ \frac{\mathrm{d}}{\mathrm{d}t}\big\vert_{t=0}\mathbb{E}f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V},t}(c)) =\mathbb{E}\left[\partial_\gamma f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c))\frac{\mathrm{d}}{\mathrm{d}t}\big\vert_{t=0}\gamma_{\mathcal{V},t}(c)\right] \\ =\mathbb{E}\left[\partial_\gamma f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c))\mathbb{E}\left[\frac{\mathbb{1}_{V_{\mathit{2}}=c}}{p_{V_{\mathit{2}}}(c)}(g(V,X)-\gamma_{\mathcal{V}}(c))s(V,X)\right]\right] \\ = \mathbb{E}\left[\mathbb{E}\left[\partial_\gamma f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c))\right] \frac{\mathbb{1}_{V_{\mathit{2}}=c}}{p_{V_{\mathit{2}}}(c)}(g(V,X)-\gamma_{\mathcal{V}}(c))s(V,X)\right]. \end{align}\] Hence, \[\begin{align} \frac{\mathrm{d}}{\mathrm{d}t}\big\vert_{t=0}\chi(P_{VX,t}) = \frac{\mathrm{d}}{\mathrm{d}t}\big\vert_{t=0}\mathbb{E}f(V,X,\mu_{\mathcal{X},t},\gamma_{\mathcal{V}}(c)) +\frac{\mathrm{d}}{\mathrm{d}t}\big\vert_{t=0}\mathbb{E}f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V},t}(c)) \\ +\mathbb{E}\left[f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c))s(V,X)\right] \\ = \mathbb{E}r(V_{\mathit{1}},X)\left(m(V,X)-\mu_{\mathcal{X}}(V_{\mathit{1}},X)\right)s(V,X) \\ + \mathbb{E}\left[\mathbb{E}\left[\partial_\gamma f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c))\right] \frac{\mathbb{1}_{V_{\mathit{2}}=c}}{p_{V_{\mathit{2}}}(c)}(g(V,X)-\gamma_{\mathcal{V}}(c))s(V,X)\right] \\ +\mathbb{E}\left[f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c))s(V,X)\right] \\ = \mathbb{E}\tilde{\chi}(V,X)s(V,X), \end{align}\] where the last equality follows from \(\mathbb{E}s(V,X)=0\). ◻

Proof of 1. First, \[\begin{align} P_{VX}\big[ -r(V_{\mathit{1}},X)h(V_{\mathit{1}},X)+f(V,X,h,\gamma_{\mathcal{V}}(c))\big]&=0, \tag{36} \\ P_{VX}\big[ (m(V,X)-\mu_{\mathcal{X}}(V_{\mathit{1}},X))h(V_{\mathit{1}},X)\big]&=0 \tag{37} \end{align}\] for all \(h\in L_2(P_{V_{\mathit{1}}X})\), where the first equality is by 10 and the second is by the definition of \(\mu_{\mathcal{X}}\) and the tower property of expectation. Then, because \(\tilde{\chi}\) is an influence function satisfying \(P_{VX}\tilde{\chi}=0\), and \(\tilde{\chi},\chi_0\) — not depending on \((v,x)\) — are constants with respect to \(P_{VX}\)-integration, \[\begin{align} \chi'-\chi_0+P_{VX}\tilde{\chi}' = P_{VX}\left[\chi'+\tilde{\chi}'\right] - P_{VX}\left[\chi_0+\tilde{\chi}\right] \\ = P_{VX}\left\{r'(V_{\mathit{1}},X)(m(V,X)-\mu_{\mathcal{X}}'(V_{\mathit{1}},X))+\frac{\mathbb{1}_{V_{\mathit{2}}=c}}{p_{V_{\mathit{2}}}'(c)}(g(V,X)-\gamma_{\mathcal{V}}'(c))e'' \right. \\ +\left. \vphantom{\frac{\mathbb{1}_{V_{\mathit{2}}=c}}{p_{V_{\mathit{2}}}'(c)}} f(V,X, \mu_{\mathcal{X}}', \gamma_{\mathcal{V}}'(c)) \right\} \\ -P_{VX}\left\{r(V_{\mathit{1}},X)(m(V,X)-\mu_{\mathcal{X}}(V_{\mathit{1}},X))\vphantom{\frac{\mathbb{1}_{V_{\mathit{2}}=c}}{p_{V_{\mathit{2}}}(c)}} \right. \\ +\left. \frac{\mathbb{1}_{V_{\mathit{2}}=c}}{p_{V_{\mathit{2}}}(c)}(g(V,X)-\gamma_{\mathcal{V}}(c))P_{VX} \partial_\gamma f(V,X, \mu_{\mathcal{X}}, \gamma_{\mathcal{V}}(c))+\vphantom{\frac{\mathbb{1}_{V_{\mathit{2}}=c}}{p_{V_{\mathit{2}}}(c)}} f(V,X, \mu_{\mathcal{X}}, \gamma_{\mathcal{V}}(c)) \right\} \\ -P_{VX}\left\{-r(V_{\mathit{1}},X)(\mu_{\mathcal{X}}'(V_{\mathit{1}},X)-\mu_{\mathcal{X}}(V_{\mathit{1}},X)) \right. \\ \left. +f(V,X,\mu_{\mathcal{X}}',\gamma_{\mathcal{V}}(c))-f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c))\right\} \\ -P_{VX}\left\{\left[m(V,X)-\mu_{\mathcal{X}}(V_{\mathit{1}},X)\right]\left[r'(V_{\mathit{1}},X)-r(V_{\mathit{1}},X)\right]\right\}, \end{align}\] where the last two expectations are zero: we apply 36 and 37 choosing \(h\) to be \(\mu_{\mathcal{X}}'-\mu_{\mathcal{X}}\) and \(r'-r\), respectively, and use the linearity of \(f\) in 5 . As \(P_{VX}\frac{\mathbb{1}_{V_{\mathit{2}}=c}}{p_{V_{\mathit{2}}}(c)}(g(V,X)-\gamma_{\mathcal{V}}(c))=0\) by the tower property and the definition of \(\gamma_{\mathcal{V}}\), \[\begin{align} \chi'-\chi_0+P_{VX}\tilde{\chi}' = -P_{VX}(r-r')(\mu_{\mathcal{X}}-\mu_{\mathcal{X}}') \\ +P_{VX}\left\{\frac{\mathbb{1}_{V_{\mathit{2}}=c}}{p_{V_{\mathit{2}}}'(c)}(g(V,X)-\gamma_{\mathcal{V}}'(c)) e'' \right. \\ \left. \phantom{\frac{\mathbb{1}_{V_{\mathit{2}}=c}}{p_{V_{\mathit{2}}}'(c)}} +f(V,X,\mu_{\mathcal{X}}',\gamma_{\mathcal{V}}'(c))-f(V,X,\mu_{\mathcal{X}}',\gamma_{\mathcal{V}}(c))\right\}. \end{align}\] Next, a Taylor-approximation of \(\gamma\mapsto f(V,X,\mu_{\mathcal{X}}',\gamma)\) by 6 with a mean-value representation of the the remainder term gives \[\begin{align} f(V,X,\mu_{\mathcal{X}}',\gamma_{\mathcal{V}}(c))=&\,f(V,X,\mu_{\mathcal{X}}',\gamma_{\mathcal{V}}'(c))+\partial_\gamma f(V,X,\mu_{\mathcal{X}}',\gamma_{\mathcal{V}}'(c))\big[\gamma_{\mathcal{V}}(c)-\gamma_{\mathcal{V}}'(c)\big] \\ &+\frac{1}{2} \partial_\gamma^2 f(V,X,\mu_{\mathcal{X}}',\widetilde{\gamma_{\mathcal{V}}(c)})\big[\gamma_{\mathcal{V}}(c)-\gamma_{\mathcal{V}}'(c)\big]^2. \end{align}\] As \(P_{VX}\mathbb{1}_{V_{\mathit{2}}=c}(g(V,X)-\gamma_{\mathcal{V}}'(c))=p_{V_{\mathit{2}}}(c)(\gamma_{\mathcal{V}}(c)-\gamma_{\mathcal{V}}'(c))\), \[\begin{align} P_{VX}\left\{\frac{\mathbb{1}_{V_{\mathit{2}}=c}}{p_{V_{\mathit{2}}}'(c)}(g(V,X)-\gamma_{\mathcal{V}}'(c)) e'' +f(V,X,\mu_{\mathcal{X}}',\gamma_{\mathcal{V}}'(c))-f(V,X,\mu_{\mathcal{X}}',\gamma_{\mathcal{V}}(c))\right\} \\ = \frac{p_{V_{\mathit{2}}}(c)}{p_{V_{\mathit{2}}}'(c)}(\gamma_{\mathcal{V}}(c)-\gamma_{\mathcal{V}}'(c))e''- P_{VX}\partial_\gamma f(V,X,\mu_{\mathcal{X}}',\gamma_{\mathcal{V}}'(c))\big[\gamma_{\mathcal{V}}(c)-\gamma_{\mathcal{V}}'(c)\big] \\ -\frac{1}{2}P_{VX} \partial_\gamma^2 f(V,X,\mu_{\mathcal{X}}',\widetilde{\gamma_{\mathcal{V}}(c)})\big[\gamma_{\mathcal{V}}(c)-\gamma_{\mathcal{V}}'(c)\big]^2, \end{align}\] which yields the assertion by the definition of \(e'\). ◻

10 Privacy↩︎

This section contains complementary results to and proofs of the claims in 4 and 5. 10.1 complements 4 by deriving the tangent set of the private model, and proves 2. 10.2 proves 5’s 1. Our proofs rest on the auxiliary results [priv:lem:linopqx_properties,priv:lem:qtc_implications,priv:lem:norms] in 10.3, proven there.

Finally, in 10.4, the relation between total-variation and differential privacy is quantified, while 10.5 discusses alternative choices for invertible privacy mechanisms in the sense of 13 .

10.1 Private Inferential Properties↩︎

This section complements 4 by deriving the tangent set of the private model, and proves 2.

1 shows that the tangent set of the model \(\mathcal{P}_{VZ}(Q,\mathcal{P}_{VX})\) is \[\mathcal{T}_{VZ}(Q,\mathcal{P}_{VX})=Q_{\mathcal{X}}^{*}\mathcal{T}_{VX}=\left\{(v,z)\mapsto\mathbb{E}\left[\left. s(V,X)\,\right\vert\, V=v,Z=z\right]: s\in\mathcal{T}_{VX}\right\},\] which is typical of mixture models such as \(P_{VZ}\) in 11 (e.g.[46]), and is also in agreement with [26]. 1 also shows that \(\mathcal{P}_{VZ}(Q,\mathcal{P}_{\mathfrak{V}\mathfrak{X}})\) remains nonparametric if \(Q_\mathcal{X}\) is invertible, and when \(X\) is discrete with \(|\mathfrak{X}|=|\mathfrak{Z}|\), this holds with “if and only if.”

Lemma 1 (Private Semiparametric Properties). Let \(\mathcal{Q}_J,\mathcal{Q}^{\mathrm{I}}_J,\mathcal{Q}_{\delta}\) be the sets of mechanism defined in ?? , ?? , 15 , respectively. Then the tangent set \(\mathcal{T}_{VZ}(Q, \mathcal{P}_{VX})\) at \(P_{VZ}\) in the model 12 is as follows.

  1. Let \(Q\in\mathcal{Q}(\mathfrak{X}\to\mathfrak{Z})\) be arbitrary. Then \(\mathcal{T}_{VZ}(Q,\mathcal{P}_{VX})=\left\{Q_{\mathcal{X}}^{*}s: s\in\mathcal{T}_{VX}\right\}\) for the tangent set \(\mathcal{T}_{VX}\subset L_2^0(P_{VX})\) at \(P_{VX}\) in any model \(\mathcal{P}_{VX}\), and \(Q_{\mathcal{X}}^{*}\) in 3.

  2. Suppose that \(\mathcal{T}_{VX}= L_2^0(P_{VX})\) and \(Q\in\mathcal{Q}_J\). Then \(\mathcal{T}_{VZ}(Q, \mathcal{P}_{\mathfrak{V}\mathfrak{X}})=L_2^0(P_{VZ})\) if and only if \(Q\in\mathcal{Q}^{\mathrm{I}}_J\).

  3. Suppose that \(\mathcal{T}_{VX}= L_2^0(P_{VX})\) and \(Q\in\mathcal{Q}_{\delta}\). Then the closure of \(\mathcal{T}_{VZ}(Q, \mathcal{P}_{\mathfrak{V}\mathfrak{X}})\) in \(L_2(P_{VZ})\) is \(L_2^0(P_{VZ})\).

Proof of 1. Assertion [priv:lem:semiparametric95vz95transform]. Consider the submodel \(t\mapsto \mathrm{d}P_{VX,t}=e^{ts}(P_{VX}e^{ts})^{-1}\mathrm{d}P_{VX}\) for \(s\in\mathcal{T}_{VX}\subset L_2^0(P_{VX})\) (e.g.[45]). Clearly, \(P_{VX,t}\ll P_{VX}\) for all \(t\). This submodel is differentiable in the quadratic mean with score \(s\) ([38]): under \(P_{VX,t}\ll P_{VX}\), \[\begin{align} \int \left[\frac{1}{t}\left(\sqrt{\frac{\mathrm{d}P_{VX,t}}{{\mathrm{d}P_{VX}}}}-1\right)-\frac{1}{2}s\right]^2\,\mathrm{d}P_{VX}\to 0\,\text{ as } t\to0. \label{priv:eq:dqm} \end{align}\tag{38}\] The submodel \(P_{VX,t}\) induces the submodel \(P_{VZ,t}\mathrel{\vcenter{:}}=\int_{(\cdot)\times\mathfrak{X}} Q(\cdot{\vert\,}x)\,\mathrm{d}P_{VX,t}(v,x)\) for \(P_{VZ}=\int_{(\cdot)\times\mathfrak{X}} Q(\cdot{\vert\,}x)\,\mathrm{d}P_{VX}(v,x)\) because \(Q\) is known, following the construction 11 . Thus, \(P_{VZ,t}\ll P_{VZ}\) for all \(t\). For the assertion, it is then necessary and sufficient that \(P_{VZ,t}\) be differentiable in quadratic mean with score \(Q_{\mathcal{X}}^{*}s\): \[\begin{align} \int \left[\frac{1}{t}\left(\sqrt{\frac{\mathrm{d}P_{VZ,t}}{{\mathrm{d}P_{VZ}}}}-1\right)-\frac{1}{2}Q_{\mathcal{X}}^{*}s\right]^2\,\mathrm{d}P_{VZ}\to 0\,\text{ as } t\to0. \label{priv:eq:dqm95vz} \end{align}\tag{39}\] To show this, we follow [47]. Note that \(\frac{\mathrm{d}P_{VZ,t}}{{\mathrm{d}P_{VZ}}}=Q_{\mathcal{X}}^{*}\frac{\mathrm{d}P_{VX,t}}{{\mathrm{d}P_{VX}}}\), because one can verify (as in the proof of 4) that the conditional distribution on the right in \(Q_{\mathcal{X}}^{*}\) is \(P_{X{\vert\,}VZ}(B{\vert\,}v,z)=\frac{\mathrm{d}\int_{(\cdot)\times B}Q(\cdot{\vert\,}x)\mathrm{d}P_{VX}(\tilde{v}, x)}{\mathrm{d}P_{VZ}}(v,z)\). Further, \[\begin{align} \frac{\mathrm{d}P_{VZ,t}}{{\mathrm{d}P_{VZ}}}=Q_{\mathcal{X}}^{*}\frac{\mathrm{d}P_{VX,t}}{{\mathrm{d}P_{VX}}} = Q_{\mathcal{X}}^{*}\left(\sqrt{\frac{\mathrm{d}P_{VX,t}}{{\mathrm{d}P_{VX}}}} -Q_{\mathcal{X}}^{*}\sqrt{\frac{\mathrm{d}P_{VX,t}}{{\mathrm{d}P_{VX}}}} + Q_{\mathcal{X}}^{*}\sqrt{\frac{\mathrm{d}P_{VX,t}}{{\mathrm{d}P_{VX}}}} \right)^2 \nonumber \\ =Q_{\mathcal{X}}^{*}\left(\sqrt{\frac{\mathrm{d}P_{VX,t}}{{\mathrm{d}P_{VX}}}} -Q_{\mathcal{X}}^{*}\sqrt{\frac{\mathrm{d}P_{VX,t}}{{\mathrm{d}P_{VX}}}}\right)^2+ \left(Q_{\mathcal{X}}^{*}\sqrt{\frac{\mathrm{d}P_{VX,t}}{{\mathrm{d}P_{VX}}}} \right)^2 \eqqcolon\sigma_t^2 + \left(Q_{\mathcal{X}}^{*}\sqrt{\frac{\mathrm{d}P_{VX,t}}{{\mathrm{d}P_{VX}}}} \right)^2, \label{priv:eq:sigmat} \end{align}\tag{40}\] where the last equality follows from \(Q_{\mathcal{X}}^{*}\) being the conditional expectation given \((V,Z)\), and we introduced \(\sigma_t^2\in L_2(P_{VZ})\). It follows that \[\begin{align} L_2(P_{VZ})\ni\delta_t\mathrel{\vcenter{:}}=\sqrt{\frac{\mathrm{d}P_{VZ,t}}{{\mathrm{d}P_{VZ}}}}- Q_{\mathcal{X}}^{*}\sqrt{\frac{\mathrm{d}P_{VX,t}}{{\mathrm{d}P_{VX}}}}\geq 0. \label{priv:eq:positivedelta} \end{align}\tag{41}\]

Write \[\begin{align} r_t\mathrel{\vcenter{:}}=\sqrt{\frac{\mathrm{d}P_{VX,t}}{{\mathrm{d}P_{VX}}}} - 1 - \frac{1}{2}ts,\quad \bar{r}_t \mathrel{\vcenter{:}}=\sqrt{\frac{\mathrm{d}P_{VZ,t}}{{\mathrm{d}P_{VZ}}}} - 1 - \frac{1}{2}t Q_{\mathcal{X}}^{*}s \label{priv:eq:residuals95dqm} \end{align}\tag{42}\] for the residuals \(r_t\in L_2(P_{VX})\) and \(\bar{r}_t \in L_2(P_{VZ})\). By 38 , \(P_{VX}r_t^2=o\left(t^2\right)\) as \(t\to0\). To show 39 , we show \(P_{VZ}\bar{r}_t^2=o\left(t^2\right)\) as \(t\to0\). Using that \((x+y)^2\leq 2(x^2+y^2)\) for all \(x,y\in\mathbb{R}\), expand \[\begin{align} P_{VX}\bar{r}_t^2 \leq 2 P_{VZ}(Q_{\mathcal{X}}^{*}r_t)^2 + 2 P_{VZ}(\bar{r}_t - Q_{\mathcal{X}}^{*}r_t)^2. \end{align}\] Here, the first term \(P_{VZ}(Q_{\mathcal{X}}^{*}r_t)^2\leq\) \(P_{VZ}Q_{\mathcal{X}}^{*}r_t^2=P_{VX}r_t^2=o\left(t^2\right)\), where the first inequality is Jensen’s, conditional on \((V,Z)\); the second term \[\begin{align} P_{VZ}(\bar{r}_t - Q_{\mathcal{X}}^{*}r_t)^2 = P_{VZ}\left(\sqrt{\frac{\mathrm{d}P_{VZ,t}}{{\mathrm{d}P_{VZ}}}} - Q_{\mathcal{X}}^{*}\sqrt{\frac{\mathrm{d}P_{VX,t}}{{\mathrm{d}P_{VX}}}} \right)^2 = P_{VZ}\delta_t^2. \end{align}\] By definitions 40 and 41 , \(\left(\delta_t+Q_{\mathcal{X}}^{*}\sqrt{\frac{\mathrm{d}P_{VX,t}}{{\mathrm{d}P_{VX}}}}\right)^2=\frac{\mathrm{d}P_{VZ,t}}{{\mathrm{d}P_{VZ}}}=\sigma_t^2+\left(Q_{\mathcal{X}}^{*}\sqrt{\frac{\mathrm{d}P_{VX,t}}{{\mathrm{d}P_{VX}}}}\right)^2\), yielding the algebraic identity \[\begin{align} \delta_t^2=\sigma_t^2-2\delta_t Q_{\mathcal{X}}^{*}\sqrt{\frac{\mathrm{d}P_{VX,t}}{{\mathrm{d}P_{VX}}}}.\label{priv:eq:delta95identity} \end{align}\tag{43}\] But \(\delta_t\geq 0\) by 41 , and so is \(Q_{\mathcal{X}}^{*}\sqrt{\frac{\mathrm{d}P_{VX,t}}{{\mathrm{d}P_{VX}}}}\geq 0\) as a density is nonnegative; hence \(\delta_t^2\leq \sigma_t^2\). Using 42 , \[\begin{align} \begin{aligned}\label{priv:eq:sigmabound} \sigma_t^2= Q_{\mathcal{X}}^{*}\left(\frac{1}{2}t(s-Q_{\mathcal{X}}^{*}s)+r_t-Q_{\mathcal{X}}^{*}r_t\right)^2 \leq \frac{1}{2} t^2 Q_{\mathcal{X}}^{*}(s-Q_{\mathcal{X}}^{*}s)^2 + 2 Q_{\mathcal{X}}^{*}(r_t-Q_{\mathcal{X}}^{*}r_t)^2 \\ \leq \frac{t^2}{2} Q_{\mathcal{X}}^{*}s^2 + 2 Q_{\mathcal{X}}^{*}r_t, \end{aligned} \end{align}\tag{44}\] because the projection \(Q_{\mathcal{X}}^{*}\) decreases the conditional variance. Here, by the tower property of expectations, \(P_{VZ}Q_{\mathcal{X}}^{*}s^2=P_{VX}s^2=O\left(1\right)\) as \(s\in L_2(P_{VX})\) since it is a score, and \(P_{VZ}Q_{\mathcal{X}}^{*}r_t^2=P_{VX}r_t^2=o\left(t^2\right)\) by 38 . Hence, \(P_{VZ}\sigma_t^2\to0\) as \(t\to 0\).

For a fixed, strictly positive \(\beta\in\mathbb{R}\), define the set of \((V,Z)\), \[A_{t,\beta}\mathrel{\vcenter{:}}=\left\{Q_{\mathcal{X}}^{*}\sqrt{\frac{\mathrm{d}P_{VX,t}}{{\mathrm{d}P_{VX}}}} \geq \frac{1}{2}, \sigma_t^2\leq \beta\right\}.\] On \(A_{t,\beta}\), squaring 43 gives \[\delta_t^2=\frac{\sigma_t^4-\delta_t^4-2\delta_t^3Q_{\mathcal{X}}^{*}\sqrt{\frac{\mathrm{d}P_{VX,t}}{{\mathrm{d}P_{VX}}}}}{4\left(Q_{\mathcal{X}}^{*}\sqrt{\frac{\mathrm{d}P_{VX,t}}{{\mathrm{d}P_{VX}}}}\right)^2}\leq \sigma_t^4\leq \beta \sigma_t^2\] since \(\delta_t\geq 0\); while on its complement \(A_{t,\beta}^\mathrm{c}\), like \(P_{VZ}\)-everywhere, \(\delta_t^2\leq\sigma_t^2\) as seen above. From 44 and 38 , \[\begin{align} P_{VZ}\delta_t^2 = P_{VZ}\mathbb{1}_{A_{t,\beta}}\delta_t^2+P_{VZ}\mathbb{1}_{A_{t\beta}^{\mathrm{c}}}\delta_t^2 \leq \beta P_{VZ} \sigma_t^2 + P_{VZ}\mathbb{1}_{A_{t,\beta}^{\mathrm{c}}}\sigma_t^2 \\ \leq \frac{\beta t^2}{2} P_{VZ}Q_{\mathcal{X}}^{*}s^2 + o\left(t^2\right)+\frac{t^2}{2}P_{VZ}\mathbb{1}_{A_{t,\beta}^{\mathrm{c}}}Q_{\mathcal{X}}^{*}s^2 + o\left(t^2\right). \end{align}\] Since \(P_{VZ}(A_{t,\beta}^{\mathrm{c}})\to 0\) as \(t\to 0\) by the previous paragraph, a small enough choice of \(\beta\) shows the right side \(o\left(t^2\right)\) Conclude that \(P_{VZ}\bar{r}_t^2\leq P_{VZ}\delta_t^2=o\left(t^2\right)\), whereby 39 holds.

Assertions [priv:lem:semiparametric95vz95nonpara95discr] and [priv:lem:semiparametric95vz95nonpara95gen]. We build on [46]. By [priv:lem:semiparametric95vz95transform], \(\mathcal{T}_{VZ}=\left\{Q_{\mathcal{X}}^{*}s: s\in L_2^0(P_{VX})\right\}\), which is the range \(R(Q_{\mathcal{X}}^{*})\) of \(Q_{\mathcal{X}}^{*}\). It follows from the defining relation between \(Q_\mathcal{X}\) and \(Q_{\mathcal{X}}^{*}\) that \(R(Q_{\mathcal{X}}^{*})^\perp=N((Q_{\mathcal{X}}^{*})^{*})\), where \((Q_{\mathcal{X}}^{*})^*=Q_\mathcal{X}\) in Hilbert spaces \(L_2(P_{VX})\) and \(L_2(P_{VZ})\), and \[R(Q_{\mathcal{X}}^{*})^\perp\mathrel{\vcenter{:}}=\left\{k\in L_2(P_{VZ}): P_{VZ}k\kappa =0 \text{ holds for all } \kappa\in R(Q_{\mathcal{X}}^{*})\right\}\] is the orthocomplement of \(R(Q_{\mathcal{X}}^{*})\), and \(N(Q_\mathcal{X})\mathrel{\vcenter{:}}=\left\{k\in L_2(P_{VZ}): Q_\mathcal{X}k = 0\right\}\) is the kernel of \(Q_\mathcal{X}\). By properties of Hilbert spaces, it follows that \(\overline{R(Q_{\mathcal{X}}^{*})}=N(Q_\mathcal{X})^\perp\), where \(\overline{R(Q_{\mathcal{X}}^{*})}\) is the closure of \(R(Q_{\mathcal{X}}^{*})\) in \(L_2(P_{VZ})\). Studying the kernel \(N(Q_\mathcal{X})\), the relation \[\begin{align} 0=(Q_\mathcal{X}k)(v,x)\, \text{ for all } (v,x)\in\mathfrak{V}\times\mathfrak{X}, \end{align}\] for [priv:lem:semiparametric95vz95nonpara95discr] is equivalent to \[\begin{align} 0_{J} = Q^\intercal \begin{bmatrix} k(v,z_1) \\ k(v,z_2) \\ \vdots \\ k(v,z_J) \end{bmatrix}\, \text{ for all } v\in\mathfrak{V}\quad \\ \iff (Q^\intercal)^{-1}0_J=0_J= \begin{bmatrix} k(v,z_1) \\ k(v,z_2) \\ \vdots \\ k(v,z_J) \end{bmatrix}\, \text{ for all } v\in\mathfrak{V}, \end{align}\] by invertibility of \(Q\). Hence, \(k=0\), and thus \(\overline{R(Q_{\mathcal{X}}^{*})}=N(Q_\mathcal{X})^\perp=L_2(P_{VZ})\) so that the model remains nonparametric. For [priv:lem:semiparametric95vz95nonpara95gen], \(Q_\mathcal{X}^{-1}0=0\) with \(Q_\mathcal{X}^{-1}\) in 3 [priv:lem:linopqx95properties95inv95qtc], yields the same conclusion. ◻

Proof of 2. We follow [46]. First, we address the nonparametric model \(\mathcal{P}_{VX}=\mathcal{P}_{\mathfrak{V}\mathfrak{X}}\) with tangent set \(\mathcal{T}_{VX}=L_2^0(P_{VX})\) and efficient influence function \(\tilde{\chi}\) in ?? . Consider the submodel \(t\mapsto P_{VX,t}\) of 1 with scores \(s\in\mathcal{T}_{VX}=L_2^0(P_{VX})\). Because \(Q\in\mathcal{Q}_{\psi}\), these submodels induce submodel \(t\mapsto P_{VZ,t}\) and the tangent set \(\mathcal{T}_{VZ}(Q,\bar{P}_{VX})=\left\{Q_{\mathcal{X}}^{*}s: s\in L_2^0(P_{VX})\right\}\) by 1. By construction, since \(Q\in\mathcal{Q}_{\psi}\), \(\psi(P_{VZ,t})=\chi(P_{VX,t})\) for all \(t\in\mathbb{R}\).

The efficient influence function \(\tilde{\psi}\) of \(\psi(P_{VZ})\) exists if and only if \[\frac{\mathrm{d}}{\mathrm{d}t}\big\vert_{t=0}\psi(P_{VZ,t})=P_{VZ}[\tilde{\psi }(Q_{\mathcal{X}}^{*}s)]\] for all regular submodels \(P_{VZ,t}\) with score \(Q_{\mathcal{X}}^{*}s\in \left\{Q_{\mathcal{X}}^{*}s: s\in L_2^0(P_{VX})\right\}\). But \(\frac{\mathrm{d}}{\mathrm{d}t}\big\vert_{t=0}\psi(P_{VZ,t})=\frac{\mathrm{d}}{\mathrm{d}t}\big\vert_{t=0}\chi(P_{VX,t})\) by the previous paragraph. Since \(\tilde{\chi}\) is the efficient influence function of \(\chi(P_{VX})\) by assumption, we must also have that \[\frac{\mathrm{d}}{\mathrm{d}t}\big\vert_{t=0}\chi(P_{VX,t})=P_{VX}[s\tilde{\chi}].\] Thus, for all \(s\in L_2^0(P_{VX})\), \[\begin{align} P_{VZ}[\tilde{\psi }(Q_{\mathcal{X}}^{*}s)]=\frac{\mathrm{d}}{\mathrm{d}t}\big\vert_{t=0}\psi(P_{VZ,t})=\frac{\mathrm{d}}{\mathrm{d}t}\big\vert_{t=0}\chi(P_{VX,t})=P_{VX}[s\tilde{\chi}]. \end{align}\] By the definition of the adjoint \((Q_{\mathcal{X}}^{*})^{*}=Q_\mathcal{X}\), the inner product \(P_{VZ}[\tilde{\psi }(Q_{\mathcal{X}}^{*}s)]\) is equal to \(P_{VZ}[\tilde{\psi }(Q_{\mathcal{X}}^{*}s)]=P_{VX}[(((Q_{\mathcal{X}}^{*})^{*})\tilde{\psi})s]=P_{VX}[(Q_\mathcal{X}\tilde{\psi})s]\), which by the last display is equal to \(P_{VX}[s\tilde{\chi}]\) for all \(s\in L_2^0(P_{VX})\). Hence \(P_{VX}[(Q_\mathcal{X}\tilde{\psi})s]=P_{VX}[s\tilde{\chi}]\) for all \(s\in L_2^0(P_{VX})\), or, by rearrangement, \(P_{VX}[(Q_\mathcal{X}\tilde{\psi}-\tilde{\chi})s]=0\) for all \(s\in L_2^0(P_{VX})\). Equivalently, \(Q_\mathcal{X}\tilde{\psi}-\tilde{\chi}\) must be in the orthocomplement of \(L_2^0(P_{VX})\subset L_2(P_{VX})\), which, as we show below, is \[L_2^0(P_{VX})^\perp=\left\{f\in L_2(P_{VX}):f-P_{VX}f=0\, P_{VX}\text{-a.s.}\right\}.\] Now, \(\tilde{\chi}\) is the efficient influence function for \(\chi(P_{VX})\), so \(P_{VX}\tilde{\chi}=0\). For \(\tilde{\psi}\) to be an influence function for \(\psi(P_{VZ})\), we must have \(P_{VZ}\tilde{\psi}=0\), but by 3[priv:lem:linopqx95properties95change], \(P_{VZ}\tilde{\psi}=P_{VX}Q_\mathcal{X}\tilde{\psi}\). Hence, \(\tilde{\chi},Q_\mathcal{X}\tilde{\psi }\in L_2^0(P_{VX})\) and \(Q_\mathcal{X}\tilde{\psi }- \tilde{\chi }\in L_2^0(P_{VX})\). But \(Q_\mathcal{X}\tilde{\psi }- \tilde{\chi }\in L_2^0(P_{VX})^\perp\) too as we showed above. Since \(L_2^0(P_{VX})\cap L_2^0(P_{VX})^\perp=\left\{f:f=0\, P_{VX}\text{-a.s.}\right\}\), we must have \(Q_\mathcal{X}\tilde{\psi}=\tilde{\chi}\) \(P_{VX}\)-a.s.. Because \(Q\in\mathcal{Q}_{\psi}\), \(Q_\mathcal{X}^{-1}\) exists, giving \(\tilde{\psi}=Q_\mathcal{X}^{-1}\tilde{\chi}\).

To see that the orthocomplement \(L_2^0(P_{VX})^\perp\) of \(L_2^0(P_{VX})\) in \(L_2(P_{VX})\) is \[\left\{f\in L_2(P_{VX}):f-P_{VX}f=0\, P_{VX}\text{-a.s.}\right\},\] take some \(f\in L_2^0(P_{VX})^{\perp}\). Because \(f\in L_2^0(P_{VX})^{\perp}\) and \(f-P_{VX}f \in L_2^0(P_{VX})\), we must have \(P_{VX}[f(f-P_{VX}f)]=0\). Because \(P_{VX}[f(f-P_{VX}f)]=P_{VX}[(f-P_{VX}f)(f-P_{VX}f)]\), we must have \(f-P_{VX}f=0\) \(P_{VX}\)-a.s..Because \(f\in L_2^0(P_{VX})^{\perp}\) was chosen arbitrarily, the assertion follows.

Second, we address an arbitrary model \(\mathcal{P}_{VX}\subset\mathcal{P}_{\mathfrak{V}\mathfrak{X}}\) with \(\mathcal{T}_{VX}\subset L_2^0(P_{VX})\) and efficient influence function \(\varphi\in Q_\mathcal{X}Q_{\mathcal{X}}^{*}\mathcal{T}_{VX}\). Let \(\tilde{\Psi}\) be the efficient influence function of \(\psi(P_{VZ})\) in the model \(\mathcal{P}_{VZ}(Q, \mathcal{P}_{VX})\) of 12 . We have \(P_{VX}[(Q_\mathcal{X}\tilde{\Psi}-\varphi)s]=0\) for all \(s\in \mathcal{T}_{VX}\) by the arguments above, that is, \[Q_\mathcal{X}\tilde{\Psi}-\varphi\in\mathcal{T}_{VX}^\perp\mathrel{\vcenter{:}}=\left\{f\in L_2(P_{VX}): P_{VX}fs=0\text{ for all } s\in\mathcal{T}_{VX}\right\}.\] If \(Q_\mathcal{X}\tilde{\Psi}-\varphi \in\mathcal{T}_{VX}\) also, then \(Q_\mathcal{X}\tilde{\Psi}-\varphi=0\) a.s.must be, from where the assertion follows by the invertibility of \(Q_\mathcal{X}\). From the previous display, \(\tilde{\Psi}=Q_\mathcal{X}^{-1}(f_0+\varphi)\) for some \(f_0\in\mathcal{T}_{VX}^\perp.\) Since \(\tilde{\Psi}\in\mathcal{T}_{VZ}(Q, \mathcal{P}_{VX})\), which set is \(\left\{Q_{\mathcal{X}}^{*}s: s\in \mathcal{T}_{VX}\right\}\) by 1, we also have that \((Q_{\mathcal{X}}^{*})^{-1}\tilde{\Psi}\in\mathcal{T}_{VX}\). But then \(P_{VX}f_0(Q_{\mathcal{X}}^{*})^{-1}\tilde{\Psi}=0\) because \(f_0\in\mathcal{T}_{VX}^\perp\). Whence, \(P_{VX}f_0(Q_{\mathcal{X}}^{*})^{-1}Q_\mathcal{X}^{-1}(f_0+\varphi)=0\), or, equivalently, \[P_{VX}f_0(Q_{\mathcal{X}}^{*})^{-1}Q_\mathcal{X}^{-1}f_0=-P_{VX}f_0(Q_{\mathcal{X}}^{*})^{-1}Q_\mathcal{X}^{-1}\varphi.\] By the definition of \(Q_\mathcal{X}\) and \(Q_{\mathcal{X}}^{*}\) and by the invertibility of \(Q_\mathcal{X}\), one can verify with the tower property of expectations that \((Q_{\mathcal{X}}^{*})^{-1}Q_\mathcal{X}^{-1}=(Q_\mathcal{X}Q_{\mathcal{X}}^{*})^{-1}\) is a positive definite operator, so that the left side of the previous display is zero if and only if \(f_0=0\) a.s..But the right side is zero because \(f_0\in\mathcal{T}_{VX}^\perp\), and \((Q_\mathcal{X}Q_{\mathcal{X}}^{*})^{-1}\varphi\in\mathcal{T}_{VX}\) by the assumption \(\varphi\in Q_\mathcal{X}Q_{\mathcal{X}}^{*}\mathcal{T}_{VX}\). Hence, \(f_0=0\) a.s., so \(\tilde{\Psi}=Q_\mathcal{X}^{-1}\varphi\) a.s..  ◻

10.2 Private Estimation↩︎

This section proves the main result 1 of 5 via 2.

5 in 10.3 establishes some technical results concerning norms under \(P_{VX}\) and \(P_{VZ}\) and the continuity of the operators \(Q_\mathcal{X},Q_\mathcal{X}^{-1}\). In the light of these technical results, under 2, 2 shows that the empirical process term in 26 vanishes as \(o_{P_{VZ}}\left(1\right)\). It rests on the same arguments as its nonprivate counterpart, 9, but it relies on the continuity of the operator \(Q_\mathcal{X}^{-1}\). It is solely because of this that 2 is more involved than its nonprivate counterpart, 4, and that it requires that the whole \((V,X)\) be distributed on a finite set, as opposed to only \(X\) be finitely distributed. For the proof, note that \[\begin{align} \partial_\gamma \bar{f}(v,z,\mu,\bar{\gamma})\mathrel{\vcenter{:}}=\frac{\partial \bar{f}}{\partial \gamma}(v,z,\mu,\bar{\gamma})= \frac{\partial }{\partial \gamma}(Q_\mathcal{X}^{-1}(v,x)\mapsto f(v,x,\mu,\bar{\gamma}))(v,z) \\ =(Q_\mathcal{X}^{-1}(v,x)\mapsto \partial_\gamma f(v,x,\mu,\bar{\gamma}))(v,z),\quad (v,z,\mu,\bar{\gamma})\in \mathfrak{V}\times\mathfrak{Z}\times L_2(P_{V_{\mathit{1}}X})\times\Gamma. \end{align}\]

Lemma 2 (Vanishing Empirical Process Term — Private Estimators). Assume that \(\eta\) is estimated by 21 and 2 holds. Then \(\check{e}-e=o_{P_{VZ}}\left(1\right)\) and \((\bar{\mathbb{P}}_n-P_{VZ})(\hat{\tilde{\psi}}-\tilde{\psi})=o_{P_{VZ}}\left(n^{-1/2}\right)\).

Proof of 2. As in the proof of 9, we apply that if a random function \(\hat{q} \in L_2(P_{VZ})\) is independent of the random sample generating the process \(\bar{\mathbb{P}}_n\), then \[\begin{align} \label{priv:eq:meansq95empproc95vz} \begin{aligned} \int (\hat{q}(v,z)-q(v,z))^2 \,\mathrm{d}P_{VZ}(v,z)=o_{P_{VZ}}\left(1\right) \text{ implies } \\ \sqrt{n}(\bar{\mathbb{P}}_n-P_{VZ})(\hat{q} -q)=o_{P_{VZ}}\left(1\right). \end{aligned} \end{align}\tag{45}\] In particular, we shall combine this with 5, establishing the boundedness of \(Q_\mathcal{X}^{-1}\) for \(\left\lVert{\cdot}\right\rVert_{\mathrm{L}_2}\), to show the convergence of \[\begin{align} \left\lVert{Q_\mathcal{X}^{-1}T}\right\rVert_{L_2(P_{VZ})}\lesssim \left\lVert{T}\right\rVert_{\mathrm{L}_2} \label{priv:eq:meansq95empproc95vz95qinv} \end{align}\tag{46}\] to zero in \(P_{VZ}\)-probability for some \(T \in \mathrm{L}_2\).

By the linearity of \(Q_\mathcal{X}^{-1}\), \(\hat{\tilde{\psi}}-\tilde{\psi}=Q_\mathcal{X}^{-1}(\check{\tilde{\chi}}-\tilde{\chi})\). By the definitions ?? and 20 , \[\begin{align} \check{\tilde{\chi}}(v,x)-\tilde{\chi}(v,x) =&\, \bar{T}_1(v,x)+\bar{T}_2(v,x)+\bar{T}_3(v,x)+\bar{T}_4, \label{priv:eq:emproc95decomp95vz} \\ \bar{T}_1(v,x)\mathrel{\vcenter{:}}=&\, \check{r}(v_{\mathit{1}},x)(m(v,x)-{\check{\mu}_{\mathcal{X}}}(v_{\mathit{1}},x)) \nonumber \\ &-r(v_{\mathit{1}},x)(m(v,x)-\mu_{\mathcal{X}}(v_{\mathit{1}},x)), \nonumber \\ \bar{T}_2(v,x)\mathrel{\vcenter{:}}=&\, \frac{\mathbb{1}_{v_{\mathit{2}}=c}}{\check{p}_{V_{\mathit{2}}}(c)}(g(v,x)-\check{\gamma}_{\mathcal{V}}(c))\check{e}-\frac{\mathbb{1}_{v_{\mathit{2}}=c}}{p_{V_{\mathit{2}}}(c)}(g(v,x)-\gamma_{\mathcal{V}}(c))e, \nonumber \\ \bar{T}_3(v,x)\mathrel{\vcenter{:}}=&\, f(v,x, {\check{\mu}_{\mathcal{X}}}, \check{\gamma}_{\mathcal{V}}(c)) -f(v,x, \mu_{\mathcal{X}}, \gamma_{\mathcal{V}}(c)), \nonumber \\ \bar{T}_4\mathrel{\vcenter{:}}=&\, -\psi(\hat{P}_{VZ})+\psi(P_{VZ}). \nonumber \end{align}\tag{47}\] As \(\bar{T}_4\) is constant, not depending on \((v,x)\), \(Q_\mathcal{X}^{-1}\bar{T}_4=\bar{T}_4\) and \((\bar{\mathbb{P}}_n-P_{VZ}) Q_\mathcal{X}^{-1}\bar{T}_4=0\). It remains to show \((\bar{\mathbb{P}}_n-P_{VZ})Q_\mathcal{X}^{-1}\bar{T}_j=o_{P_{VZ}}\left(n^{-1/2}\right)\) for \(j=1,2,3\) by the linearity of the process \(\bar{\mathbb{P}}_n-P_{VZ}\).

Term \(\bar{T}_1.\,\,\) In the light of 46 , \((\bar{\mathbb{P}}_n-P_{VZ})Q_\mathcal{X}^{-1}\bar{T}_1=o_{P_{VZ}}\left(n^{-1/2}\right)\) can be established along the same steps as that of \(T_1\) in the proof of 9. In particular, \(\left\lVert{Q_\mathcal{X}^{-1}\bar{T}_1}\right\rVert_{L_2(P_{VZ})}\lesssim \left\lVert{\bar{T}_1}\right\rVert_{\mathrm{L}_2}\). Namely, suppressing the arguments, write \[\begin{align} \bar{T}_1=\check{r}(m-{\check{\mu}_{\mathcal{X}}})-r(m-\mu_{\mathcal{X}}) &= (\check{r}-r+r)(m-{\check{\mu}_{\mathcal{X}}})-r(m-\mu_{\mathcal{X}}) \\ &= (\check{r}-r)(m-{\check{\mu}_{\mathcal{X}}}) + r(\mu_{\mathcal{X}}-{\check{\mu}_{\mathcal{X}}}). \end{align}\] By 2, \(\left\lVert{m-{\check{\mu}_{\mathcal{X}}}}\right\rVert_{\infty}=O_{{P_{VZ}}}\left(1\right)\); either directly by ?? , or by ?? and ?? , noting that \(\left\lVert{m-{\check{\mu}_{\mathcal{X}}}}\right\rVert_{\infty}\leq \left\lVert{m-\mu_{\mathcal{X}}}\right\rVert_{\infty}+\left\lVert{\mu_{\mathcal{X}}-{\check{\mu}_{\mathcal{X}}}}\right\rVert_{\infty}=O\left(1\right)+o_{P_{VZ}}\left(1\right)=O_{{P_{VZ}}}\left(1\right)\). Then the convergence ?? of \(r\) implies that \((\bar{\mathbb{P}}_n-P_{VZ}) Q_\mathcal{X}^{-1}((\check{r}-r)(m-{\check{\mu}_{\mathcal{X}}}))=o_{P_{VZ}}\left(n^{-1/2}\right)\) by 46 as \[\begin{align} \left\lVert{(\check{r}-r)(m-{\check{\mu}_{\mathcal{X}}})}\right\rVert_{\mathrm{L}_2}\leq\left\lVert{m-{\check{\mu}_{\mathcal{X}}}}\right\rVert_{\infty}\left\lVert{\check{r}-r}\right\rVert_{\mathrm{L}_2}=O_{{P_{VZ}}}\left(1\right)o_{P_{VZ}}\left(1\right)=o_{P_{VZ}}\left(1\right) \end{align}\] since \(\left\lVert{|q|^2}\right\rVert_{\infty}=\left\lVert{q}\right\rVert_{\infty}^2\).

By 2, either ?? , or ?? and ?? . In the former case, \[\begin{align} \left\lVert{r(\mu_{\mathcal{X}}-{\check{\mu}_{\mathcal{X}}})}\right\rVert_{\mathrm{L}_2}\leq\left\lVert{\mu_{\mathcal{X}}-{\check{\mu}_{\mathcal{X}}}}\right\rVert_{\infty}\left\lVert{r}\right\rVert_{\mathrm{L}_2}=o_{P_{VZ}}\left(1\right), \end{align}\] because \(r\in \mathrm{L}_2\). In the latter case, \[\begin{align} \left\lVert{r(\mu_{\mathcal{X}}-{\check{\mu}_{\mathcal{X}}})}\right\rVert_{\mathrm{L}_2}\leq \bar{R} \left\lVert{\mu_{\mathcal{X}}-{\check{\mu}_{\mathcal{X}}}}\right\rVert_{\mathrm{L}_2}=o_{P_{VZ}}\left(1\right), \end{align}\] since ?? bounds \(r\) and \({\check{\mu}_{\mathcal{X}}}\) is convergent by ?? . Thus, \((\bar{\mathbb{P}}_n-P_{VZ})Q_\mathcal{X}^{-1}\) \((r(\mu_{\mathcal{X}}-{\check{\mu}_{\mathcal{X}}}))=o_{P_{VZ}}\left(n^{-1/2}\right)\) by 45 . Conclude that \((\bar{\mathbb{P}}_n-P_{VZ})Q_\mathcal{X}^{-1}\bar{T}_1=o_{P_{VZ}}\left(n^{-1/2}\right)\).

Term \(\bar{T}_2.\,\,\) By the mean-value theorem there exists \((\tilde{\gamma}_{\mathcal{V}}(c), \tilde{p}_{V_{\mathit{2}}}(c), \tilde{e})\) between \((\gamma_{\mathcal{V}}(c), p_{V_{\mathit{2}}}(c), e)\) and \((\check{\gamma}_{\mathcal{V}}(c), \check{p}_{V_{\mathit{2}}}(c), \check{e})\) such that \[\begin{align} \bar{T}_2(v,x)=&\,\frac{\mathbb{1}_{v_{\mathit{2}}=c}}{\check{p}_{V_{\mathit{2}}}(c)}(g(v,x)-\check{\gamma}_{\mathcal{V}}(c))\check{e}-\frac{\mathbb{1}_{v_{\mathit{2}}=c}}{p_{V_{\mathit{2}}}(c)}(g(v,x)-\gamma_{\mathcal{V}}(c))e\\ =&\,-\frac{\mathbb{1}_{v_{\mathit{2}}=c}}{\tilde{p}_{V_{\mathit{2}}}(c)}\tilde{e}(\check{\gamma}_{\mathcal{V}}(c)-\gamma_{\mathcal{V}}(c)) \\ &-\frac{\mathbb{1}_{v_{\mathit{2}}=c}}{\tilde{p}_{V_{\mathit{2}}}(c)^2}(g(v,x)-\tilde{\gamma}_{\mathcal{V}}(c))\tilde{e}(\check{p}_{V_{\mathit{2}}}(c)-p_{V_{\mathit{2}}}(c)) \\ &+\frac{\mathbb{1}_{v_{\mathit{2}}=c}}{\tilde{p}_{V_{\mathit{2}}}(c)}(g(v,x)-\tilde{\gamma}_{\mathcal{V}}(c))(\check{e}-e). \end{align}\] The standard central limit theorem applies to the i.i.d.sequence \[\bigg((Q_\mathcal{X}^{-1}(v,x)\mapsto\mathbb{1}_{v_{\mathit{2}}=c})(V_i,Z_i)\bigg)_{i\in[n]},\] hence \(\sqrt{n}(\bar{\mathbb{P}}_n-P_{VZ})((Q_\mathcal{X}^{-1}(v,x)\mapsto\mathbb{1}_{v_{\mathit{2}}=c})(V,Z))=O_{{P_{VZ}}}\left(1\right)\). By the linearity of the process \(\sqrt{n}(\bar{\mathbb{P}}_n-P_{VZ}) Q_\mathcal{X}^{-1}\), \[\begin{align} \sqrt{n}(\bar{\mathbb{P}}_n-P_{VZ})\big\{\big[Q_\mathcal{X}^{-1}(v,x)\mapsto\mathbb{1}_{v_{\mathit{2}}=c}g(v,x)-\mathbb{1}_{v_{\mathit{2}}=c}\tilde{\gamma}_{\mathcal{V}}(c)\big](V,Z)\big\} \\ = \sqrt{n}(\bar{\mathbb{P}}_n-P_{VZ})\big\{\big[Q_\mathcal{X}^{-1}(v,x)\mapsto \mathbb{1}_{v_{\mathit{2}}=c}g(v,x)\big](V,Z)\big\} \\ - \tilde{\gamma}_{\mathcal{V}}(c)\sqrt{n}(\bar{\mathbb{P}}_n-P_{VZ})\big\{\big[Q_\mathcal{X}^{-1}(v,x)\mapsto\mathbb{1}_{v_{\mathit{2}}=c}\big](V,Z)\big\} \\ =(1-\tilde{\gamma}_{\mathcal{V}}(c))O_{{P_{VZ}}}\left(1\right)=O_{{P_{VZ}}}\left(1\right) \end{align}\] again by the standard central limit theorem and ?? . Suppose that \(\check{e}-e=o_{P_{VZ}}\left(1\right)\), which we show later. Then by ?? and ?? , \((\bar{\mathbb{P}}_n-P_{VZ})Q_\mathcal{X}^{-1}\bar{T}_2=o_{P_{VZ}}\left(n^{-1/2}\right)\).

Term \(\bar{T}_3.\,\,\) Recall that \(\bar{T}_3(v,x)=f(v,x, {\check{\mu}_{\mathcal{X}}}, \check{\gamma}_{\mathcal{V}}(c))-f(v,x, \mu_{\mathcal{X}}, \gamma_{\mathcal{V}}(c))\). By the consistency of \(\check{\gamma}_{\mathcal{V}}\) and \({\check{\mu}_{\mathcal{X}}}\) (?? and ?? or ?? ), we have \(\left\lVert{\bar{T}_3}\right\rVert_{\mathrm{L}_2}=o_{P_{VZ}}\left(1\right)\) by the continuous mapping theorem and ?? . Conclude by 45 and 46 that \((\bar{\mathbb{P}}_n-P_{VZ})Q_\mathcal{X}^{-1}\bar{T}_3=o_{P_{VZ}}\left(n^{-1/2}\right)\).

Consistency of \(\check{e}.\,\,\) By the definition of \(e,\check{e}\), \[\begin{align} \check{e}-e=&\, \bar{\mathbb{P}}_n''\partial_\gamma \bar{f}(V,Z,{\check{\mu}_{\mathcal{X}}},\check{\gamma}_{\mathcal{V}}(c)) - P_{VZ}\partial_\gamma \bar{f}(V,Z,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c)) \\ =&\, \bar{\mathbb{P}}_n''\partial_\gamma \bar{f}(V,Z,{\check{\mu}_{\mathcal{X}}},\check{\gamma}_{\mathcal{V}}(c))-P_{VZ}\partial_\gamma \bar{f}(V,Z,{\check{\mu}_{\mathcal{X}}},\check{\gamma}_{\mathcal{V}}(c)) \\ &+ P_{VZ}\partial_\gamma \bar{f}(V,Z,{\check{\mu}_{\mathcal{X}}},\check{\gamma}_{\mathcal{V}}(c))- P_{VZ}\partial_\gamma f(V,Z,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c)) \\ =&\, (\bar{\mathbb{P}}_n''-P_{VZ})\partial_\gamma \bar{f}(V,Z,{\check{\mu}_{\mathcal{X}}},\check{\gamma}_{\mathcal{V}}(c)) \\ &+ P_{VZ}\big[\partial_\gamma \bar{f}(V,Z,{\check{\mu}_{\mathcal{X}}},\check{\gamma}_{\mathcal{V}}(c))-\partial_\gamma \bar{f}(V,Z,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c)) \big] \\ =&\, (\bar{\mathbb{P}}_n''-P_{VZ})\partial_\gamma \bar{f}(V,Z,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c)) \\ &+(\bar{\mathbb{P}}_n''-P_{VZ})\big[\partial_\gamma \bar{f}(V,Z,{\check{\mu}_{\mathcal{X}}},\check{\gamma}_{\mathcal{V}}(c))-\partial_\gamma \bar{f}(V,Z,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c))\big] \\ &+ P_{VZ}\big[\partial_\gamma \bar{f}(V,Z,{\check{\mu}_{\mathcal{X}}},\check{\gamma}_{\mathcal{V}}(c))-\partial_\gamma \bar{f}(V,Z,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c)) \big] \end{align}\] Here, the first term is \(O_{{P_{VZ}}}\left(n^{-1/2}\right)=o_{P_{VZ}}\left(1\right)\) by the standard central limit theorem, and the second and third term are \(o_{P_{VZ}}\left(1\right)\) by the continuity ?? of \(\partial_\gamma f\) using that \(\partial_\gamma \bar{f} =Q_\mathcal{X}^{-1}\partial_\gamma f\) along the same arguments concerning \(\bar{T}_3\) above. ◻

Proof of 1. Follows from [priv:lem:chi_eff_empprocess_priv,priv:thm:dr], noting that, for \(\bar{e}'\) in 29 , \[\check{e}-\bar{e}'=(\bar{\mathbb{P}}_n''-P_{VZ})\partial_\gamma \bar{f}(V,Z,{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c))\] is \(o_{P_{VZ}}\left(1\right)\) by the consistency proof of \(\check{e}\) in 2. ◻

10.3 Auxiliary Results↩︎

In this section, auxiliary results underpinning the proofs in 10 are derived: in 3, the properties of the linear operator \(Q_\mathcal{X}\); in 4, the distributions of the (partly) unobserved data \((V,X,Z)\) under mechanism 15 ; in 5, norms under \(P_{VX}\) and \(P_{VZ}\), and consequent continuity of \(Q_\mathcal{X},Q_\mathcal{X}^{-1}\).

The following properties of \(Q_\mathcal{X}\) play an essential role in the derivation of the tangent set and the efficient influence function.

Lemma 3 (Properties of \(Q_\mathcal{X}\) in 16 ).

  1. The operator \(Q_\mathcal{X}\) is the conditional expectation operator \((Q_\mathcal{X}k)(v,x)=\mathbb{E}\left[\left. k(V,Z)\,\right\vert\, V=v,X=x\right]\), \((v,x)\in\mathfrak{V}\times\mathfrak{X}\).

  2. The operator \(Q_\mathcal{X}\) has adjoint \(Q_{\mathcal{X}}^{*}:L_2(P_{VX})\to L_2(P_{VZ})\), \((Q_{\mathcal{X}}^{*}h)(v,z)=\mathbb{E}\left[\left. h(V,X)\,\right\vert\, V=v,Z=z\right]\), \((v,z)\in\mathfrak{V}\times\mathfrak{Z}\).

  3. Change of measure: \(P_{VZ}k = P_{VX}Q_\mathcal{X}k\) for all \(k\in L_2(P_{VZ})\). In particular, if \(Q_\mathcal{X}: L_2(P_{VZ})\to S \subset L_2(P_{VX})\) has an inverse \(Q_\mathcal{X}^{-1}: S\to L_2(P_{VZ})\) so that \(Q_\mathcal{X}Q_\mathcal{X}^{-1}h=h\) for all \(h\in S\), then \(k\mathrel{\vcenter{:}}= Q_\mathcal{X}^{-1}h\) yields \(P_{VZ}Q_\mathcal{X}^{-1}h=P_{VX}Q_\mathcal{X}Q_\mathcal{X}^{-1}h=P_{VX}h.\)

  4. If \(X\) is distributed on a finite set with \(|\mathfrak{Z}|=|\mathfrak{X}|=J\), then the operator \(Q_\mathcal{X}: L_2(P_{VZ})\to L_2(P_{VX})\) of 16 can be represented in the matrix notation ?? as, for all \(v\in\mathfrak{V}\), \[\begin{align} \begin{bmatrix} (Q_\mathcal{X}k)(v,x_1) \\ (Q_\mathcal{X}k)(v,x_2) \\ \vdots \\ (Q_\mathcal{X}k)(v,x_J) \end{bmatrix} =Q^\intercal \begin{bmatrix} k(v,z_1) \\ k(v,z_2) \\ \vdots \\ k(v,z_J) \end{bmatrix}, \end{align}\] and it has inverse \(Q_\mathcal{X}^{-1}:L_2(P_{VX})\to L_2(P_{VZ})\) if and only if \(Q\) is invertible, given by, for all \(v\in\mathfrak{V}\), \[\begin{align} \begin{bmatrix} k(v,z_1) \\ k(v,z_2) \\ \vdots \\ k(v,z_J) \end{bmatrix} = (Q^\intercal)^{-1} \begin{bmatrix} h(v,x_1) \\ h(v,x_2) \\ \vdots \\ h(v,x_J) \end{bmatrix} \end{align}\] with \((Q^\intercal)^{-1}=(Q^{-1})^\intercal\).

  5. Let \(Q\in\mathcal{Q}_{\delta}\) in 15 . Then \((Q_\mathcal{X}k)(v,x)=\alpha k(v,x)+(1-\alpha)\int_\mathcal{X}k(v,z)\bar{Q}(\mathrm{d}z)\); moreover, \(Q_\mathcal{X}:L_2(P_{VZ})\to L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})\) is a bounded, hence continuous, linear operator for the norm \[\begin{align} \left\lVert{h}\right\rVert_{L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})}\mathrel{\vcenter{:}}=\left\lVert{h}\right\rVert_{L_2(P_{VX})}+\left\lVert{h}\right\rVert_{L_2(P_V\otimes \bar{Q})}\label{priv:eq:normqbar} \end{align}\qquad{(15)}\] on \(L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})\), where \(P_V\otimes \bar{Q}\) is the distribution of a random element \((V,\bar{Z})\) with independent coordinates \(V\sim P_V\) and \(\bar{Z}\sim\bar{Q}\).

  6. Let \(Q\in\mathcal{Q}_{\delta}\) in 15 . The inverse of \(Q_\mathcal{X}: L_2(P_{VZ})\to L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})\) exists, and is, as in [48], \[(Q_\mathcal{X}^{-1}h)(v,z)=\frac{1}{\alpha}h(v,z)-\frac{1-\alpha}{\alpha}\int_\mathfrak{X}h(v,x)\bar{Q}(\mathrm{d}x)\] for all \(h \in L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q}).\) That is, \(Q_\mathcal{X}Q_\mathcal{X}^{-1}h=h\) and \(Q_\mathcal{X}^{-1}Q_\mathcal{X}k\) \(=k\) for all \(h\in L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})\) and all \(k\in L_2(P_{VZ})\). Moreover \(Q_\mathcal{X}^{-1}:L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})\to L_2(P_{VZ})\) is a bounded, hence continuous, linear operator for the norm ?? .

  7. Let \(Q\in\mathcal{Q}_{\delta}\) in 15 . Then for all \(h\in L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})\), \[\begin{align} P_{VX}h^2+\frac{1-\alpha}{\alpha}\left((P_V\otimes\bar{Q})h^2- P_V\left(\int h(V,x)\bar{Q}(\mathrm{d}x)\right)^2\right) \leq P_{VZ}(Q_\mathcal{X}^{-1}h)^2 \\ \leq \frac{2-\alpha}{\alpha} P_{VX}\tilde{\chi}^2+2\frac{(1-\alpha)(2-\alpha)}{\alpha^2}(P_V\otimes\bar{Q})h^2. \end{align}\] The second term in the lower bound is nonnegative due to Jensen’s inequality.

Proof of 3. Assertion [priv:lem:linopqx95properties95exp]. Follows directly from \(Z\mid(V,X)\sim Q(\cdot{\vert\,}X)\).

Assertion [priv:lem:linopqx95properties95adj]. By definition, the adjoint \(Q_{\mathcal{X}}^{*}\) of \(Q_\mathcal{X}\) satisfies \(P_{VX}[(Q_\mathcal{X}k) h]=P_{VZ}[kQ_{\mathcal{X}}^{*}h]\). By the tower property of expectation we can verify that, for \(Q_{\mathcal{X}}^{*}\) given in [priv:lem:linopqx95properties95adj], \[\begin{align} P_{VX}[(Q_\mathcal{X}k) h]=\mathbb{E}\left[\mathbb{E}\left[\left. k(V,Z)\,\right\vert\, V,X\right]h(V,X)\right]=\mathbb{E}\left[\mathbb{E}\left[\left. k(V,Z)h(V,X)\,\right\vert\, V,X\right]\right] \\ =\mathbb{E}\left[\mathbb{E}\left[\left. k(V,Z)h(V,X)\,\right\vert\, V,Z\right]\right]=\mathbb{E}\left[k(V,Z)\mathbb{E}\left[\left. h(V,X)\,\right\vert\, V,Z\right]\right]=P_{VZ}[kQ_{\mathcal{X}}^{*}h]. \end{align}\]

Assertion [priv:lem:linopqx95properties95change]. By the tower property of expectation, \[P_{VZ}k=\mathbb{E}k(V,Z) =\mathbb{E}\mathbb{E}\left[\left. k(V,Z)\,\right\vert\, V,X\right] = P_{VX}Q_\mathcal{X}k.\]

Assertion [priv:lem:linopqx95properties95exp95inv95disc]. Under the discrete model for \(X\), \[\int_\mathfrak{X}k(v,z)Q(\mathrm{d}z{\vert\,}x)=\sum_{z\in\mathfrak{X}}k(v,z)Q(\left\{z\right\}{\vert\,}x),\] from which the assertion directly follows.

Assertion [priv:lem:linopqx95properties95exp95qtc]. Under the mechanism \(Q\) in 15 , we have \[\begin{align} (Q_\mathcal{X}k)(v,x)= \int_\mathfrak{X}k(v,z)Q(\mathrm{d}z{\vert\,}x) = \int_\mathfrak{X}k(v,z)\left[\alpha\delta_x(\mathrm{d}z)+(1-\alpha)\bar{Q}(\mathrm{d}z)\right] \\ = \alpha k(v,x)+(1-\alpha)\int_\mathfrak{X}k(v,z) \bar{Q}(\mathrm{d}z). \end{align}\] Next we show that \(Q_\mathcal{X}: L_2(P_{VZ})\to L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})\) is bounded. Take some \(k\in L_2(P_{VZ})\), and let \((Q_\mathcal{X}k)(v,x)= \alpha k(v,x)+(1-\alpha)\int_\mathfrak{X}k(v,z) \bar{Q}(\mathrm{d}z)\eqqcolon\phi_1(v,x)+\phi_2(v)\). By Minkowski’s inequality, \[\left\lVert{Q_\mathcal{X}k}\right\rVert_{L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})}\leq \left\lVert{\phi_1}\right\rVert_{L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})}+\left\lVert{\phi_2}\right\rVert_{L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})},\] where, by definition ?? , \[\left\lVert{\phi_1}\right\rVert_{L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})}=\left\lVert{\phi_1}\right\rVert_{L_2(P_{VX})}+\left\lVert{\phi_1}\right\rVert_{L_2(P_V\otimes \bar{Q})}\] with \(\left\lVert{\phi_1}\right\rVert_{L_2(P_{VX})}=\alpha\left\lVert{k}\right\rVert_{L_2(P_{VX})}\leq \sqrt{\alpha}\left\lVert{k}\right\rVert_{L_2(P_{VZ})}\) by 4[priv:lem:qtc95implications95norms], and \[\left\lVert{\phi_2}\right\rVert_{L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})}=2\left\lVert{\phi_2}\right\rVert_{L_2(P_{V})}.\] By Jensen’s inequality, \[\left(\int_\mathfrak{X}k(v,z) \bar{Q}(\mathrm{d}z)\right)^2\leq \int_\mathfrak{X}k(v,z)^2 \bar{Q}(\mathrm{d}z),\] thus \[\begin{align} \left\lVert{\phi_2}\right\rVert_{L_2(P_{V})}\leq (1-\alpha)\sqrt{\int_\mathfrak{V}\int_\mathfrak{X}k(v,z)^2 \bar{Q}(\mathrm{d}z)P_V(\mathrm{d}v)}\\ =(1-\alpha)\sqrt{\int_\mathfrak{V}\int_\mathfrak{X}k(v,z)^2 \bar{q}(z) \nu_X(\mathrm{d}z) p_V(v)\nu_V(\mathrm{d}v)} \\ \leq \sqrt{(1-\alpha)} \sqrt{\int_\mathfrak{V}\int_\mathfrak{X}k(v,z)^2 p_{VZ}(v,z) \nu_X(\mathrm{d}z)\nu_V(\mathrm{d}v)} = \sqrt{(1-\alpha)} \left\lVert{k}\right\rVert_{L_2(P_{VZ})}, \end{align}\] where we used that by 4 [priv:lem:qtc95implications95dens], \[\bar{q}(z)p_V(v)=\frac{1}{1-\alpha}p_{VZ}(v,z)-\frac{\alpha}{1-\alpha}p_{VX}(v,z)\leq \frac{1}{1-\alpha}p_{VZ}(v,z)\] for \(\alpha \in(0,1)\). Finally, by the previous display, \(\left\lVert{\phi_1}\right\rVert_{L_2(P_V\otimes \bar{Q})}\leq \sqrt{\frac{1}{1-\alpha}}\left\lVert{\phi_1}\right\rVert_{L_2(P_{VZ})}\), yielding \[\begin{align} \left\lVert{Q_\mathcal{X}k}\right\rVert_{L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})} \\ \leq \sqrt{\alpha}\left\lVert{k}\right\rVert_{L_2(P_{VZ})} + \frac{\alpha}{\sqrt{1-\alpha}}\left\lVert{k}\right\rVert_{L_2(P_{VZ})}+2 \sqrt{(1-\alpha)} \left\lVert{k}\right\rVert_{L_2(P_{VZ})}. \end{align}\]

Assertion [priv:lem:linopqx95properties95inv95qtc]. It suffices to show \(Q_\mathcal{X}(Q_\mathcal{X}^{-1}h)=h\) for all \(h\in L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})\). Under 15 , \[\begin{align} (Q_\mathcal{X}(Q_\mathcal{X}^{-1}h))(v,x) = \alpha (Q_\mathcal{X}^{-1}h)(v,x)+(1-\alpha) \int_{\mathfrak{X}} (Q_\mathcal{X}^{-1}h)(v,z)\bar{Q}(\mathrm{d}z). \end{align}\] Here, the first term is \[\begin{align} \alpha (Q_\mathcal{X}^{-1}h)(v,x) &= \alpha \left\{ \frac{1}{\alpha}h(v,x)-\frac{1-\alpha}{\alpha}\int_\mathfrak{X}h(v,t)\bar{Q}(\mathrm{d}t)\right\} \\ &=h(v,x) -(1-\alpha)\int_\mathfrak{X}h(v,t)\bar{Q}(\mathrm{d}t), \end{align}\] while the second term is \[\begin{align} (1-\alpha) \int_{\mathfrak{X}} (Q_\mathcal{X}^{-1}h)(v,z)\bar{Q}(\mathrm{d}z)\\ = (1-\alpha) \int_{\mathfrak{X}} \left\{\frac{1}{\alpha}h(v,z)-\frac{1-\alpha}{\alpha}\int_\mathfrak{X}h(v,t)\bar{Q}(\mathrm{d}t)\right\} \bar{Q}(\mathrm{d}z) \\ =(1-\alpha)\left\{ \frac{1}{\alpha}\int_{\mathfrak{X}} h(v,z)\bar{Q}(\mathrm{d}z) - \frac{1-\alpha}{\alpha}\int_\mathfrak{X}h(v,t)\bar{Q}(\mathrm{d}t) \right\} \\ = (1-\alpha)\int_\mathfrak{X}h(v,t)\bar{Q}(\mathrm{d}t), \end{align}\] where we used that \(\bar{Q}\) is a probability measure with \(\int_{\mathfrak{X}}\bar{Q}(\mathrm{d}z)=1\). Collecting terms, conclude that \(Q_\mathcal{X}(Q_\mathcal{X}^{-1}h)=h\). We showed that \(Q_\mathcal{X}^{-1}\) is the inverse of \(Q_\mathcal{X}\), which is by [priv:lem:linopqx95properties95exp95qtc] a bounded operator.

To see that \(Q_\mathcal{X}^{-1}: L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})\to L_2(P_{VZ})\) is bounded, take some \(h\in L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})\). As \((Q_\mathcal{X}^{-1}h)(v,z)=\frac{1}{\alpha}h(v,z)-\frac{1-\alpha}{\alpha}\int_\mathfrak{X}h(v,x)\bar{Q}(\mathrm{d}x)\), we have \[\begin{align} \left\lVert{Q_\mathcal{X}^{-1}h}\right\rVert_{L_2(P_{VZ})}\leq \frac{1}{\alpha}\left\lVert{h}\right\rVert_{L_2(P_{VZ})}+\frac{1-\alpha}{\alpha}\left\lVert{\int_\mathfrak{X}h(\cdot,x)\bar{Q}(\mathrm{d}x)}\right\rVert_{L_2(P_V)}, \end{align}\] where, by 4 [priv:lem:qtc95implications95jointobs], \[\begin{align} \left\lVert{h}\right\rVert_{L_2(P_{VZ})}=\sqrt{\int h^2\,\mathrm{d}[\alpha P_{VX}+(1-\alpha) P_V \otimes \bar{Q}]} \\ \leq \sqrt{\left\lVert{h}\right\rVert_{L_2(P_{VX})}^2 + \left\lVert{h}\right\rVert_{L_2(P_V \otimes \bar{Q})}^2} \leq \sqrt{(\left\lVert{h}\right\rVert_{L_2(P_{VX})} + \left\lVert{h}\right\rVert_{L_2(P_V \otimes \bar{Q})})^2} \\ \leq \left\lVert{h}\right\rVert_{L_2(P_{VX})} + \left\lVert{h}\right\rVert_{L_2(P_V \otimes \bar{Q})} = \left\lVert{h}\right\rVert_{L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})}. \end{align}\] By Jensen’s inequality, \[\begin{align} \left\lVert{\int_\mathfrak{X}h(\cdot,x)\bar{Q}(\mathrm{d}x)}\right\rVert_{L_2(P_V)} \\ \leq \sqrt{\int h^2 \,\mathrm{d}P_V\otimes\bar{Q}}=\left\lVert{h}\right\rVert_{L_2(P_V \otimes \bar{Q})}\leq \left\lVert{h}\right\rVert_{L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})}, \end{align}\] whereby \[\begin{align} \left\lVert{Q_\mathcal{X}^{-1}h}\right\rVert_{L_2(P_{VZ})}\leq \frac{1}{\alpha}\left\lVert{h}\right\rVert_{L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})} +\frac{1-\alpha}{\alpha}\left\lVert{h}\right\rVert_{L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})}. \end{align}\]

Assertion [priv:lem:linopqx95properties95bounds]. By [priv:lem:linopqx95properties95inv95qtc], \[\begin{align} (Q_\mathcal{X}^{-1}h)^2(v,z)= \frac{1}{\alpha^2}h^2(v,z)-2\frac{1-\alpha}{\alpha^2} h(v,z)\int h(v,x)\bar{Q}(\mathrm{d}x)+\left(\frac{1-\alpha}{\alpha}\int h(v,x)\bar{Q}(\mathrm{d}x)\right)^2. \label{priv:eq:qinv95sqr} \end{align}\tag{48}\] Here, because \(xy\leq (x^2+y^2)/2\) for all \(x,y\in\mathbb{R}\), \[\begin{align} \bigg\vert h(v,z)\int h(v,x)\bar{Q}(\mathrm{d}x) \bigg\vert \leq \frac{1}{2}\left( h^2(v,z)+\left(\int h(v,x)\bar{Q}(\mathrm{d}x)\right)^2\right). \label{priv:eq:hproduct95bound} \end{align}\tag{49}\] Then 48 yields the lower bound \[\begin{align} P_{VZ}(Q_\mathcal{X}^{-1}h)^2\geq \frac{1}{\alpha^2}P_{VZ}h^2- \frac{1-\alpha}{\alpha^2}\left(P_{VZ}h^2+P_V \left(\int h(V,x)\bar{Q}(\mathrm{d}x)\right)^2 \right) \\ +P_V\left(\frac{1-\alpha}{\alpha}\int h(V,x)\bar{Q}(\mathrm{d}x)\right)^2 = \frac{1}{\alpha}P_{VZ}h^2-\frac{1-\alpha}{\alpha}P_V \left(\int h(V,x)\bar{Q}(\mathrm{d}x)\right)^2 \end{align}\] Here, \(P_{VZ}=\alpha P_{VX}+(1-\alpha)P_V\otimes\bar{Q}\) by 4, which proves the lower bound. For the upper bound, 48 and 49 give \[\begin{align} P_{VZ}(Q_\mathcal{X}^{-1}h)^2\leq \frac{1}{\alpha^2}P_{VZ}h^2+ \frac{1-\alpha}{\alpha^2}\left(P_{VZ}h^2+P_V \left(\int h(V,x)\bar{Q}(\mathrm{d}x)\right)^2 \right) \\ +P_V\left(\frac{1-\alpha}{\alpha}\int h(V,x)\bar{Q}(\mathrm{d}x)\right)^2 = \frac{2-\alpha}{\alpha^2} P_{VZ}h^2 + \frac{(2-\alpha)(1-\alpha)}{\alpha^2} P_V \left(\int h(V,x)\bar{Q}(\mathrm{d}x)\right)^2. \end{align}\] Then \(P_{VZ}=\alpha P_{VX}+(1-\alpha)P_V\otimes\bar{Q}\) and Jensen’s inequality \(P_V \left(\int h(v,x)\bar{Q}(\mathrm{d}x)\right)^2\leq (P_V\otimes\bar{Q})h^2\) prove the upper bound. ◻

By the construction of \(Z_i\) in 4.1, the sequence \(((V_i,X_i,Z_i))_{i\in[n]}\) is an i.i.d.sample from the distribution of the partly unobserved data \((V,X,Z)\), \[\begin{align} P_{VXZ}(B_{\mathsf{v}}, B_{\mathsf{x}}, B_{\mathsf{z}}) &= \int \mathbb{1}_{(v,x)\in B_{\mathsf{v}}\times B_{\mathsf{x}}}Q(B_{\mathsf{z}}{\vert\,}x) \,\mathrm{d}P_{VX}(v,x)\label{priv:eq:distribution95vxz} \end{align}\tag{50}\] for \(B_{\mathsf{v}}\in\mathscr{F}_{\mathfrak{V}}, B_{\mathsf{x}}\in\mathscr{F}_{\mathfrak{X}}, B_{\mathsf{z}}\in\mathscr{F}_{\mathfrak{Z}}\). When \(Q\in\mathcal{Q}_{\delta}\) in 15 , this distribution is as follows.

Lemma 4 (Distributions under 15 ). If the mechanism \(Q\in\mathcal{Q}_{\delta}\) in 15 , then:

  1. The joint distribution of \((V,X,Z)\) is \(P_{VXZ}(B_{\mathsf{v}},B_{\mathsf{x}},B_{\mathsf{z}})=\alpha P_{VX}(B_{\mathsf{v}},B_{\mathsf{x}}\cap B_{\mathsf{z}})+(1-\alpha)P_{VX}(B_{\mathsf{v}},B_{\mathsf{x}})\bar{Q}(B_{\mathsf{z}})\) for \(B_{\mathsf{v}}\in\mathscr{F}_{\mathfrak{V}}, B_{\mathsf{x}}\in\mathscr{F}_{\mathfrak{X}}, B_{\mathsf{z}}\in\mathscr{F}_{\mathfrak{Z}}\).

  2. The joint distribution of \((V,Z)\) is \(P_{VZ}(B_{\mathsf{v}},B_{\mathsf{z}})=\alpha P_{VX}(B_{\mathsf{v}},B_{\mathsf{z}})+(1-\alpha)P_{V}(B_{\mathsf{v}})\bar{Q}(B_{\mathsf{z}})\) for \(B_{\mathsf{v}}\in\mathscr{F}_{\mathfrak{V}}, B_{\mathsf{z}}\in\mathscr{F}_{\mathfrak{Z}}\). Hence, we have the absolute-continuity relations \(P_{VX}\ll P_{VZ}\), so \(P_X\ll P_Z\), and \(\bar{Q}\ll P_Z\).

  3. If \(P_{VX}\) has density \(p_{VX}\) with respect to some measure \(\nu_V\times\nu_X\) and \(\bar{Q}\) has density \(\bar{q}\) with respect to \(\nu_X\), then \(P_{VZ}\) has density \(p_{VZ}(v,z)\mathrel{\vcenter{:}}=\alpha p_{VX}(v,z)+(1-\alpha)p_{V}(v) \bar{q}(z)\) for \((v,z)\in\mathfrak{V}\times\mathfrak{X}\) with respect to \(\nu_V\times\nu_X\).

  4. For all \(h:\mathfrak{V}\times\mathfrak{X}\to\mathbb{R}\) and all \(p\in[1,\infty]\), \(\left\lVert{h}\right\rVert_{L_p(P_{VX})}\leq \alpha^{-1/p}\left\lVert{h}\right\rVert_{L_p(P_{VZ})}\).

  5. The Markov kernel \[\begin{align} P_{X\mid Z}(B{\vert\,}z) &\mathrel{\vcenter{:}}=\alpha\frac{\mathrm{d}P_X}{\mathrm{d}P_{Z}}(z)\delta_z(B)+(1-\alpha)P_X(B)\frac{\mathrm{d}\bar{Q}}{\mathrm{d}P_{Z}}(z),\quad B\in\mathscr{F}_{\mathfrak{X}}, z\in\mathfrak{X}, \end{align}\] is the conditional distribution of \(X\) given \(Z=z\), where the Radom-Nykodým derivatives \(\frac{\mathrm{d}P_X}{\mathrm{d}P_{Z}}\), \(\frac{\mathrm{d}\bar{Q}}{\mathrm{d}P_{Z}}\) exist by [priv:lem:qtc95implications95jointobs]. If \(P_{VX}\) and \(\bar{Q}\) have densities \(p_{VX}\) and \(\bar{q}\) with respect to \(\nu_V\times\nu_X\) and \(\nu_X\), respectively, then for \(p_Z\) induced by [priv:lem:qtc95implications95jointobs], the last display is equal to \[\begin{align} P_{X\mid Z}(B{\vert\,}z)=\alpha\frac{p_X(z)}{p_Z(z)}\delta_z(B)+(1-\alpha)\frac{\bar{q}(z)}{p_Z(z)}P_X(B), \quad B\in\mathscr{F}_{\mathfrak{X}}, z\in\mathfrak{X}. \end{align}\]

  6. Suppose that \(X\) given \(V=v\), \(v\in\mathfrak{V}\), admits a conditional distribution \(P_{X{\vert\,}V}(\cdot{\vert\,}v)\). Then the Markov kernel \[\begin{align} P_{Z{\vert\,}V}(B{\vert\,}v)\mathrel{\vcenter{:}}=\alpha P_{X{\vert\,}V}(B{\vert\,}v)+(1-\alpha)\bar{Q}(B),\quad B\in\mathscr{F}_{\mathfrak{X}}, v\in\mathfrak{V}, \end{align}\] is the conditional distribution of \(Z\) given \(V=v\). Hence, \(\bar{Q}\ll P_{Z{\vert\,}V}(\cdot{\vert\,}v)\) for any \(v\in\mathfrak{V}\).

  7. Suppose that \(X\) given \(V=v\), \(v\in\mathfrak{V}\), admits a conditional distribution \(P_{X{\vert\,}V}(\cdot{\vert\,}v)\). For each \(v\in\mathfrak{V}\), let \(\bar{q}(z{\vert\,}v)\mathrel{\vcenter{:}}=\frac{\mathrm{d}\bar{Q}}{\mathrm{d}P_{Z{\vert\,}V}(\cdot{\vert\,}v)}(z)\) be the Radom-Nykodým derivative of \(\bar{Q}\) with respect to \(P_{Z{\vert\,}V}(\cdot{\vert\,}v)\), which exists by [priv:lem:qtc95implications95cond95v]. Then the Markov kernel \[\begin{align} P_{X{\vert\,}VZ}(B{\vert\,}v,z)\mathrel{\vcenter{:}}=\alpha \frac{\mathrm{d}P_{VX}}{\mathrm{d}P_{VZ}}(v,z)\delta_z(B) + (1-\alpha) \bar{q}(z{\vert\,}v) P_{X{\vert\,}V}(B{\vert\,}v), \end{align}\] for \(B\in\mathscr{F}_{\mathfrak{X}}, (v,z)\in\mathfrak{V}\times\mathfrak{X}\), is the conditional distribution of \(X\) given \((V,Z)=(v,z)\).

Proof of 4. Assertion [priv:lem:qtc95implications95jointall]. Plug 15 into 50 , using the definition of the Dirac measure \(\delta_x(B)=\mathbb{1}_{x\in B}\), to find \[\begin{align} P_{VXZ}(B_{\mathsf{v}},B_{\mathsf{x}},B_{\mathsf{z}})=&\;\alpha \int \mathbb{1}_{(v,x)\in B_{\mathsf{v}}\times B_{\mathsf{x}}} \mathbb{1}_{x\in B_{\mathsf{z}}} \,\mathrm{d}P_{VX}(v,x) \\ &+(1-\alpha) \bar{Q}(B_{\mathsf{z}})\int\mathbb{1}_{(v,x)\in B_{\mathsf{v}}\times B_{\mathsf{x}}} \,\mathrm{d}P_{VX}(v,x) \\ =&\, \alpha \int\mathbb{1}_{(v,x)\in B_{\mathsf{v}}\times (B_{\mathsf{x}}\cap B_{\mathsf{z}})}\,\mathrm{d}P_{VX}(v,x)\\ & +(1-\alpha) \bar{Q}(B_{\mathsf{z}})\int \mathbb{1}_{(v,x)\in B_{\mathsf{v}}\times B_{\mathsf{x}}} \,\mathrm{d}P_{VX}(v,x) \\ =&\, \alpha P_{VX}(B_{\mathsf{v}},B_{\mathsf{x}}\cap B_{\mathsf{z}})+(1-\alpha)\bar{Q}(B_{\mathsf{z}})P_{VX}(B_{\mathsf{v}},B_{\mathsf{x}}). \end{align}\]

Assertion [priv:lem:qtc95implications95jointobs]. Follows from [priv:lem:qtc95implications95jointall] as the marginal distribution by setting \(B_{\mathsf{x}}\mathrel{\vcenter{:}}=\mathfrak{X}\).

Assertion [priv:lem:qtc95implications95dens]. A convex combination of two densities, \(p_{VZ}\) is nonnegative. From [priv:lem:qtc95implications95jointobs], write \(P_{VZ}(B_{\mathsf{v}},B_{\mathsf{z}})\) as \[\begin{align} \alpha \int_{B_{\mathsf{v}}}\int_{B_{\mathsf{z}}} p_{VX}(v,x)\,\mathrm{d}\nu_X(x) \,\mathrm{d}\nu_V(v)+(1-\alpha)\int_{B_{\mathsf{z}}}\bar{q}(z)\,\mathrm{d}\nu_X(x)\int_{B_{\mathsf{v}}} p_V(v)\,\mathrm{d}\nu_V(v) \\ = \int_{B_{\mathsf{v}}}\int_{B_{\mathsf{z}}} \big\{ \alpha p_{VX}(v,x)+(1-\alpha)\bar{q}(x) p_V(v) \big\}\,\mathrm{d}\nu_X(x)\,\mathrm{d}\nu_V(v). \end{align}\]

Assertion [priv:lem:qtc95implications95norms]. By [priv:lem:qtc95implications95jointobs], \(P_{VX}\leq \alpha^{-1}P_{VZ}\). Hence, \[\begin{align} \left\lVert{h}\right\rVert_{L_p(P_{VX})}= \left(\int_{\mathfrak{X}} \int_{\mathfrak{V}} |h(v,x)|^p \,\mathrm{d}P_{VX}(v,x) \right)^{1/p}\leq \alpha^{-1/p} \left\lVert{h}\right\rVert_{L_p(P_{VZ})}. \end{align}\]

Assertion [priv:lem:qtc95implications95cond95z]. It is sufficient and necessary to verify that for the \(P_{X|Z}\) given in [priv:lem:qtc95implications95cond95z], \(P_{XZ}(B_{\mathsf{x}},B_{\mathsf{z}})=\int_{B_{\mathsf{z}}} P_{X|Z}(B_{\mathsf{x}}{\vert\,}z)\,\mathrm{d}P_Z(z)\). From the right, \[\begin{align} \int_{B_{\mathsf{z}}} P_{X|Z}(B_{\mathsf{x}}{\vert\,}z) \,\mathrm{d}P_Z(z) \\ = \int_{B_{\mathsf{z}}} \bigg\{ \alpha\frac{\mathrm{d}P_X}{\mathrm{d}P_{Z}}(z)\delta_z(B_{\mathsf{x}})+(1-\alpha)P_X(B_{\mathsf{x}})\frac{\mathrm{d}\bar{Q}}{\mathrm{d}P_{Z}}(z) \bigg\}\,\mathrm{d}P_Z(z) \\ = \alpha \int_{B_{\mathsf{z}}\cap B_{\mathsf{x}}}\frac{\mathrm{d}P_X}{\mathrm{d}P_{Z}}(z) \,\mathrm{d}P_Z(z) +(1-\alpha)P_X(B_{\mathsf{x}})\int_{B_{\mathsf{z}}}\frac{\mathrm{d}\bar{Q}}{\mathrm{d}P_{Z}}(z)\,\mathrm{d}P_Z(z) \\ = \alpha P_X(B_{\mathsf{x}}\cap B_{\mathsf{z}})+(1-\alpha)P_X(B_{\mathsf{x}}) \bar{Q}(B_{\mathsf{z}}), \end{align}\] in which we recognise \(P_{XZ}(B_{\mathsf{x}},B_{\mathsf{z}})\) by [priv:lem:qtc95implications95jointall].

Assertion [priv:lem:qtc95implications95cond95vz]. Follows from the same arguments as [priv:lem:qtc95implications95cond95z].

Assertion [priv:lem:qtc95implications95cond95vz]. Follows from the same arguments as [priv:lem:qtc95implications95cond95z], by verifying that for the \(P_{X{\vert\,}VZ}\) given in the assertion, we have \(P_{VXZ}(B_{\mathsf{v}},B,B_{\mathsf{z}})=\int_{B_{\mathsf{v}}\times B_{\mathsf{z}}}P_{X{\vert\,}VZ}(B{\vert\,}v,z)\,\mathrm{d}P_{VZ}(v,z)\) for \(P_{VXZ}\) in [priv:lem:qtc95implications95jointall] and for all \(B_{\mathsf{v}}\in\mathscr{F}_{\mathfrak{V}}\), \(B\in\mathscr{F}_{\mathfrak{X}}\), \(B_{\mathsf{z}}\in\mathscr{F}_{\mathfrak{X}}\). Specifically, noting that \(\mathrm{d}P_{VZ}(v,z)=P_{Z{\vert\,}V}(\mathrm{d}z{\vert\,}v)\mathrm{d}P_{V}(v)\) by definition of the conditional distribution \(P_{Z{\vert\,}V}\), we have from the right, \[\begin{align} \int_{B_{\mathsf{v}}\times B_{\mathsf{z}}} \left\{ \alpha \frac{\mathrm{d}P_{VX}}{\mathrm{d}P_{VZ}}(v,z)\delta_z(B) + (1-\alpha) \bar{q}(z{\vert\,}v) P_{X{\vert\,}V}(B{\vert\,}v) \right\} \,\mathrm{d}P_{VZ}(v,z) \\ = \alpha \int_{B_{\mathsf{v}}\times (B_{\mathsf{z}}\cap B)}\frac{\mathrm{d}P_{VX}}{\mathrm{d}P_{VZ}}(v,z) \,\mathrm{d}P_{VZ}(v,z)\\ +(1-\alpha)\int_{B_{\mathsf{v}}}P_{X{\vert\,}V}(B{\vert\,}v)\int_{B_{\mathsf{z}}}\frac{\mathrm{d}\bar{Q}}{\mathrm{d}P_{Z{\vert\,}V}(\cdot{\vert\,}v)}(z)P_{Z{\vert\,}V}(\mathrm{d}z{\vert\,}v)\mathrm{d}P_{V}(v)\\ =\alpha P_{VX}(B_{\mathsf{v}},B_{\mathsf{z}}\cap B)+(1-\alpha)\bar{Q}(B_{\mathsf{z}})\int_{B_{\mathsf{v}}}P_{X{\vert\,}V}(B{\vert\,}v)\mathrm{d}P_V(v) \\ =\alpha P_{VX}(B_{\mathsf{v}},B_{\mathsf{z}}\cap B)+(1-\alpha)\bar{Q}(B_{\mathsf{z}}) P_{VX}(B_{\mathsf{v}},B), \end{align}\] in which we recognise \(P_{VXZ}(B_{\mathsf{v}},B,B_{\mathsf{z}})\) in [priv:lem:qtc95implications95jointall]. ◻

5 is key in proving 2.

Lemma 5 (Norms and continuity of \(Q_\mathcal{X},Q_\mathcal{X}^{-1}\)). With \(c_Q\), we denote strictly positive constants which depend only on \(Q\) and whose value may differ in every display in this lemma.

If \(Q\in\mathcal{Q}_{\delta}\) in 15 , or \(Q\in \mathcal{Q}_J\) in ?? with \(\min_{(z,x)\in\mathfrak{X}^2}Q(\left\{z\right\}{\vert\,}x)>0\), then for all \(h:\mathfrak{V}\times\mathfrak{X}\to\mathbb{R}\) and all \(p\in[1,\infty]\), \[\begin{align} \left\lVert{h}\right\rVert_{L_p(P_{VX})}\leq c_Q\left\lVert{h}\right\rVert_{L_p(P_{VZ})}. \label{priv:lem:norms:eq:lp} \end{align}\qquad{(16)}\]

If \(Q\in\mathcal{Q}^{\mathrm{I}}_J\) in ?? and \(0<\inf_{(v,x)\in\mathfrak{V}\times\mathfrak{X}}p_{VX}(v,x)\leq\sup_{(v,x)\in\mathfrak{V}\times\mathfrak{X}}p_{VX}(v,x)<\infty\), then for all \(h\in L_2(P_{VX})\), \[\begin{align} \left\lVert{Q_\mathcal{X}^{-1}h}\right\rVert_{L_2(P_{VZ})}&\leq c_Q\left\lVert{h}\right\rVert_{L_2(P_{VX})}. \label{priv:lem:norms:eq:q295295disc} \end{align}\qquad{(17)}\]

If \(Q\in\mathcal{Q}_{\delta}\) in 15 , we have for all \(h\in L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})\), \[\begin{align} \left\lVert{Q_\mathcal{X}^{-1}h}\right\rVert_{L_2(P_{VZ})}&\leq c_Q\left\lVert{h}\right\rVert_{L_2(P_{VX})\cap L_2(P_V\otimes\bar{Q})}. \label{priv:lem:norms:eq:q295295gen} \end{align}\qquad{(18)}\]

For any \(Q\in\mathcal{Q}(\mathfrak{X}\to\mathfrak{Z})\), we have for all \(k\in L_2(P_{VZ})\), \[\begin{align} \left\lVert{Q_\mathcal{X}k}\right\rVert_{L_2(P_{VX})}&\leq \left\lVert{k}\right\rVert_{L_2(P_{VZ})}. \label{priv:lem:norms:eq:q2951} \end{align}\qquad{(19)}\]

Proof of 5. Suppose that \(Q\in\mathcal{Q}_{\delta}\). Then ?? is 4 [priv:lem:qtc95implications95norms]; ?? is 3 [priv:lem:linopqx95properties95inv95qtc].

Suppose now that \(X\) is distributed on a finite set \(\mathfrak{X}=\mathfrak{Z}\) and \(P_{VX}\) has density \(p_{VX}\). First, we show ?? . We have, for any \(c>0\), \[\begin{align} \left\lVert{h}\right\rVert_{L_p(P_{VZ})}^p - \frac{1}{c}\left\lVert{h}\right\rVert_{L_p(P_{VX})}^p \\ = \int_\mathfrak{V}\sum_{x\in\mathfrak{X}}|h(v,x)|^p \left(p_{VZ}(v,x)-\frac{1}{c}p_{VX}(v,x)\right) \,\mathrm{d}\nu_V(v). \end{align}\] By assumption, \[\underline{q}\mathrel{\vcenter{:}}=\min_{(x,\bar{x})\in\mathcal{X}^2}Q(\left\{x\right\}{\vert\,}\bar{x})>0.\] Since \(p_{VZ}(v,x)=\sum_{\bar{x} \in \mathfrak{X}} Q(\left\{x\right\}{\vert\,}\bar{x})p_{VX}(v,\bar{x})\), we have \[\begin{align} p_{VZ}(v,x)-\frac{1}{c}p_{VX}(v,x) \geq \underline{q}\sum_{\bar{x}\in\mathfrak{X}}p_{VX}(v,\bar{x}) - \frac{1}{c} p_{VX}(v,x) \\ = \underline{q}\sum_{\bar{x}\in\mathfrak{X}:\bar{x}\neq x}p_{VX}(v,\bar{x})+\left(\underline{q}-\frac{1}{c}\right)p_{VX}(v,x). \end{align}\] Hence, setting \(c\mathrel{\vcenter{:}}= 1/\underline{q}\) implies \(p_{VZ}(v,x)-\frac{1}{c}p_{VX}(v,x)\geq 0\). Thus \[\left\lVert{h}\right\rVert_{L_p(P_{VZ})}^p - \frac{1}{c}\left\lVert{h}\right\rVert_{L_p(P_{VX})}^p=\left\lVert{h}\right\rVert_{L_p(P_{VZ})}^p - \underline{q}\left\lVert{h}\right\rVert_{L_p(P_{VX})}^p\geq 0.\]

Second, we show ?? . Consider the matrix representation of \(Q_\mathcal{X}^{-1}\) in 3 [priv:lem:linopqx95properties95exp95inv95disc]. The linear operator (matrix) \(Q:\mathbb{R}^{J\times J}\to \mathbb{R}^{J\times 1}\) has inverse \(Q^{-1}\) because \(Q\in\mathcal{Q}_{\psi}\). As \((Q^{-1})^\intercal:\mathbb{R}^{J\times J}\to \mathbb{R}^{J\times 1}\) is a linear operator on a finite-dimensional space, it is a continuous and bounded linear operator (e.g.[49]). Whence, \(\left\lVert{(Q^{-1})^\intercal t}\right\rVert_{2}\lesssim \left\lVert{t}\right\rVert_{2}\) for all \(t\in\mathbb{R}^{J\times 1}\), where \(\left\lVert{.}\right\rVert_{2}\) is the Euclidean norm on \(\mathbb{R}^{J\times 1}\). The assertion follows by \[\begin{align} \left\lVert{Q_\mathcal{X}^{-1}h}\right\rVert_{L_2(P_{VZ})}^2= \int_\mathfrak{V}\sum_{z\in\mathfrak{X}} \left[(Q_\mathcal{X}^{-1}h)(v,z) \right]^2 p_{VZ}(v,z)\,\mathrm{d}\nu_V(v) \\ \leq \left\lVert{p_{VZ}}\right\rVert_{\infty}\int_\mathfrak{V}\sum_{z\in\mathfrak{X}} \left[(Q_\mathcal{X}^{-1}h)(v,z) \right]^2 \,\mathrm{d}\nu_V(v) \lesssim \int_\mathfrak{V}\sum_{z\in\mathfrak{X}} \left[ \sum_{x\in\mathfrak{X}}((Q^{-1})^\intercal)_{z,x} h(v,x) \right]^2\,\mathrm{d}\nu_V(v) \\ = \int_\mathfrak{V}\left\lVert{(Q^{-1})^\intercal (h(v,x))_{x\in\mathcal{X}}}\right\rVert_{2}^2\,\mathrm{d}\nu_V(v) \lesssim \int_\mathfrak{V}\left\lVert{(h(v,x))_{x\in\mathcal{X}}}\right\rVert_{2}^2\,\mathrm{d}\nu_V(v) \\ = \int_\mathfrak{V}\sum_{x\in\mathfrak{X}} h(v,x)^2\,\mathrm{d}\nu_V(v) \simeq \int_\mathfrak{V}\sum_{x\in\mathfrak{X}} h(v,x)^2 p_{VX}(v,x)\,\mathrm{d}\nu_V(v) = \left\lVert{h}\right\rVert_{L_2(P_{VX})}, \end{align}\] since \(\left\lVert{p_{VZ}}\right\rVert_{\infty}<\infty\) by the assumption \(\left\lVert{p_{VX}}\right\rVert_{\infty}<\infty\), and \(\inf_{(v,x)\in\mathfrak{V}\times\mathfrak{X}} p_{VX}(v,x)>0\) by assumption.

Finally, we show ?? . By 3 [priv:lem:linopqx95properties95exp] and Jensen’s inequality, \[((Q_\mathcal{X}k)(v,x))^2\leq \mathbb{E}\left[\left. k(V,Z)^2\,\right\vert\, V=v,X=x\right].\] But then \[\begin{align} \left\lVert{Q_\mathcal{X}k}\right\rVert_{L_2(P_{VX})}^2= \mathbb{E}\left[((Q_\mathcal{X}k)(V,X))^2\right] \\ \leq \mathbb{E}\left[\mathbb{E}\left[\left. k(V,Z)^2\,\right\vert\, V,X\right]\right]=\mathbb{E}k(V,Z)^2 = \left\lVert{k}\right\rVert_{L_2(P_{VZ})}^2. \end{align}\] ◻

10.4 Total-Variational and Differential Privacy↩︎

Recall the definition of differential privacy, which is one of the most common privacy guarantees.

Definition 2 (Local \((\alpha,\delta)\)-Differential Privacy: \((\alpha,\delta)\)-LDP [2]). For \(\alpha,\delta\geq 0\), the privacy mechanism \(Q\in\mathcal{Q}(\mathfrak{X}\to\mathfrak{Z})\) is locally \((\alpha,\delta)\)-differentially private if \(Q(B{\vert\,}x)\leq e^\alpha Q(B{\vert\,}x')+\delta\) for all \(B\in\mathscr{F}_{\mathfrak{Z}}\) and for all \(x,x'\in\mathfrak{X}\).

6 shows how \(\alpha\)-LTVP in 1 and \((\alpha,\delta)\)-LDP are related. An \(\alpha\)-LTVP is always an \((.,\delta=\alpha)\)-LDP, which provides weaker privacy guarantee than an \((\alpha,\delta=0)\)-LDP. Conversely, when \((\alpha,\delta)\) is small enough — so that privacy is strict —, \((\alpha,\delta)\)-LDP is an \((e^\alpha-1+\delta)\)-LTVP. For example, an \((\alpha,\delta=0)\)-LDP is also an \((e^\alpha-1)\)-LTVP for \(\alpha\leq \log(2)\approx 0.69\).

Lemma 6 (\((\alpha,\delta)\)-LDP and \(\alpha\)-LTVP). For all \(0\leq \alpha\leq 1\), every \(\alpha\)-LTVP mechanism is \((\tilde{\alpha},\alpha)\)-LDP for any \(\tilde{\alpha}\geq 0\). For all \(\alpha,\delta\geq 0\) such that \(e^\alpha-1+\delta\leq 1\), every \((\alpha,\delta)\)-LDP mechanism is \((e^\alpha-1+\delta)\)-LTVP.

Proof. Let \(Q\) be \(\alpha\)-LTVP. As \(|Q|=Q\), we have for all \(B\in\mathscr{F}_{\mathfrak{Z}}\), \[\begin{align} Q(B{\vert\,}x)=|Q(B{\vert\,}x)|=|Q(B{\vert\,}x)-Q(B{\vert\,}x')+Q(B{\vert\,}x')| \\ \leq |Q(B{\vert\,}x')|+|Q(B{\vert\,}x)-Q(B{\vert\,}x')| \leq Q(B{\vert\,}x')+\alpha\leq e^{\tilde{\alpha}}Q(B{\vert\,}x')+\alpha \end{align}\] for all \(x,x'\in\mathfrak{X}\) and for any \(\tilde{\alpha}\geq 0\). Hence, \(Q\) is \((\tilde{\alpha},\alpha)\)-LDP.

Now let \(Q\) be an \((\alpha,\delta)\)-LDP mechanism. Then for all \(x,x'\in\mathfrak{X}\) and \(B\in\mathscr{F}_{\mathfrak{Z}}\), \[\begin{align} |Q(B {\vert\,}x)-Q(B{\vert\,}x')|= \begin{cases} Q(B{\vert\,}x)-Q(B{\vert\,}x')\quad&\text{ if }\, Q(B {\vert\,}x)-Q(B{\vert\,}x')\geq 0 \\ Q(B{\vert\,}x')-Q(B{\vert\,}x)\quad&\text{ if }\, Q(B {\vert\,}x)-Q(B{\vert\,}x')< 0 \end{cases} \\ \leq \begin{cases} e^\alpha Q(B{\vert\,}x')+\delta-Q(B{\vert\,}x')\quad&\text{ if }\, Q(B {\vert\,}x)-Q(B{\vert\,}x')\geq 0 \\ e^\alpha Q(B{\vert\,}x)+\delta-Q(B{\vert\,}x)\quad&\text{ if }\, Q(B {\vert\,}x)-Q(B{\vert\,}x')<0 \end{cases} \\ = \begin{cases} (e^\alpha-1) Q(B{\vert\,}x')+\delta\quad&\text{ if }\, Q(B {\vert\,}x)-Q(B{\vert\,}x')\geq 0 \\ (e^\alpha-1) Q(B{\vert\,}x)+\delta\quad&\text{ if }\, Q(B {\vert\,}x)-Q(B{\vert\,}x')< 0 \end{cases} \\ \leq e^\alpha-1+\delta, \end{align}\] because \(Q\) maps to \([0,1]\) and \(e^\alpha-1\geq 0\) for all \(\alpha\geq 0\). Hence, if \(e^\alpha-1+\delta \leq 1\), \(Q\) is \((e^\alpha-1+\delta)\)-LTVP. ◻

10.5 Invertible Privacy Mechanisms↩︎

In 4.1, we saw that for a generic \(X\), a mechanism \(Q\in\mathcal{Q}_{\delta}\) implies the existence of \(L_Q\). In some special cases, choices different from \(\mathcal{Q}_{\delta}\) also imply the existence of \(L_Q\) in 13 .

The special cases arise when \(Q(\cdot{\vert\,}x)=\bar{Q}_{\mathrm{c}}(\cdot{\vert\,}x)\), where \(\bar{Q}_{\mathrm{c}}(\cdot{\vert\,}x)\) admits a \(\nu_X\)-density \(\bar{q}_{\mathrm{c}}(\cdot{\vert\,}x)\) (as is also the case for the discretely distributed covariates). Then inverting 11 is equivalent to solving a Fredholm integral equation of the first kind (see e.g.[48]) where \(\bar{q}_{\mathrm{c}}\) induces a linear integral operator. Such equations are usually ill-posed [48], but in a few favourable cases they do admit a unique solution which also retains the nonparametric model class.

One such favourable case occurs when \(\mathfrak{Z}=\mathfrak{X}=\mathbb{R}^K\) and \(\bar{q}_{\mathrm{c}}\) corresponds to, for example, the Laplace mechanism satisfying \(\alpha\)-LDP, adding Laplace noise from the Laplace density \(p_\varepsilon\). Then for the Fourier transform \(\mathcal{F}\) and its inverse \(\mathcal{F}^{-1}\), we have \(p_{VX}(v,x)=(\mathcal{F}^{-1}(w\mapsto \frac{(\mathcal{F}p_{VZ}(v,\cdot))(w)}{(\mathcal{F}p_\varepsilon)(w)}))(x)\) by the convolution theorem. Another favourable case occurs when \(\mathfrak{Z}=\mathfrak{X}\) and \(\bar{q}_{\mathrm{c}}\) induces a compact integral operator. Then the equation admits a unique and smooth solution ([49]). One class of compact linear operators are Hilbert-Schmidt operators ([50], [49]). The \(\bar{q}_{\mathrm{c}}\) induces a Hilbert-Schmidt operator if \[\int_{\mathfrak{X}\times\mathfrak{X}}\bar{q}_{\mathrm{c}}(z{\vert\,}x)^2\,\mathrm{d}\nu_X(z)\,\mathrm{d}\nu_X(x)<\infty\] ([50]), or if \((z,x)\mapsto \bar{q}_{\mathrm{c}}(z{\vert\,}x)\) was continuous and the domain \(\mathfrak{Z}=\mathfrak{X}\subset\mathbb{R}^K\) was compact ([49]).

The first, convolution case is specific to privacy achieved by additive noise and image space \(\mathbb{R}^K\), which is not suitable for a generic covariate under our consideration, for example when \(X\) also contains coordinates distributed on a finite set. The second, compact case also places restrictions on the domain, or require \(\int_{\mathfrak{X}\times\mathfrak{X}}\bar{q}_{\mathrm{c}}(z{\vert\,}x)^2\,\mathrm{d}\nu_X(z)\,\mathrm{d}\nu_X(x)<\infty\), which we do not expect to hold unless \(\mathfrak{X}\) is compact (for example, when \(\mathfrak{X}=\mathbb{R}\) and the mechanism is the additive Laplace noise, then the last integral is infinite). Moreover, these cases assume the existence of a density. In contrast to these, our mechanism 15 allows for more generic covariate types and space \(\mathfrak{Z}=\mathfrak{X}\) handled smoothly by a single mechanism.

11 Estimation of Nuisance Parameters↩︎

This section proves 2 and 3 of 6, and details their application to the estimation of the regression and of the Riesz representer. 11.1 contains results on finite-dimensional models, while 11.2 on infinite-dimensional models.

11.1 Finite-Dimensional Models↩︎

In 11.1.1, we prove 2, and we apply it to the estimation of the regression and the Riesz representer in [priv:app:priv_nuisance:subsec:finite_dim:subsubsec:regression,priv:app:priv_nuisance:subsec:finite_dim:subsubsec:riesz], respectively.

11.1.1 Method-of-Moments↩︎

Before proving 2 on the private method-of-moments, it is helpful to recall the nonprivate version.

Lemma 7 (Nonprivate Generalised method-of-moments ([5] and [6])). Suppose that \(\theta_0\in\Theta\subset\mathbb{R}^K\) is the unique minimiser \[\theta_0=\arg\min_{\theta\in\Theta} P_{VX}\Xi_\theta,\] for a fixed \(\Xi_{\theta}:\mathfrak{V}\times\mathfrak{X}\to\mathbb{R}\), \(\theta\in\Theta\). Further suppose that the derivative \(\mathrm{D}_\theta \Xi_{\tilde{\theta}}(v,x)\) of \(\theta\mapsto \Xi_\theta(v,x)\) exists at all \(\tilde{\theta}\in\mathrm{Nb}(\theta_0)\), where \(\mathrm{Nb}(\theta_0)\) is a neighbourhood of \(\theta_0\), and for all \((v,x)\in\mathfrak{V}\times\mathfrak{X}\). Assume that

  1. The value \(\theta_0\) is in the interior of the compact \(\Theta\).

  2. Let \[\phi_{\tilde{\theta}}(v,x)\mathrel{\vcenter{:}}=\mathrm{D}_\theta \Xi_{\tilde{\theta}}(v,x)^\intercal,\quad (v,x)\in\mathfrak{V}\times\mathfrak{X},\] be the \(\mathbb{R}^{K\times 1}\)-valued derivative of \(\theta\mapsto\Xi_\theta(v,x)\) at \(\tilde{\theta}\). The \(\phi_{\theta}\) satisfies \[\begin{align} \left\lVert{\left\lVert{\phi_{\theta_0}}\right\rVert_{2}^2}\right\rVert_{L_1(P_{VX})}<\infty, \label{priv:lemma:rate95mm95phi95bound} \end{align}\qquad{(20)}\] where, for a fixed \((v,x)\in\mathfrak{V}\times\mathfrak{X}\), \(\left\lVert{\phi_{\theta_0}}\right\rVert_{2}^2(v,x)\) is sum of the \(K\) squared entries of \(\phi_{\theta_0}(v,x)\).

  3. The \(\mathbb{R}^{K\times K}\)-valued derivative \(\dot{\phi}_{\tilde{\theta}}(v,x)\mathrel{\vcenter{:}}=\mathrm{D}_\theta \phi_{\tilde{\theta}}(v,x)\) of \(\theta\mapsto \phi_{\theta}(v,x)\) at \(\tilde{\theta}\) exist at all \(\tilde{\theta}\in\mathrm{Nb}(\theta_0)\) and all \((v,x)\in\mathfrak{V}\times\mathfrak{X}\). The map \(\theta\mapsto \dot{\phi}_{\theta}(v,x)\) is continuous at all \(\theta\in\mathrm{Nb}(\theta_0)\) for all \((v,x)\in\mathfrak{V}\times\mathfrak{X}\). The expectation of \(\dot{\phi}_{\theta_0}\) exists, and the matrix \(P_{VX}\dot{\phi}_{\theta_0}\) is invertible. Furthermore, \[\begin{align} \left\lVert{\sup_{\theta\in\mathrm{Nb}(\theta_0)}\left\lVert{\dot{\phi}_\theta}\right\rVert_{1}}\right\rVert_{L_1(P_{VX})}<\infty,\label{priv:lemma:rate95mm95phideriv95bound} \end{align}\qquad{(21)}\] where, for a fixed \((v,x)\in\mathfrak{V}\times\mathfrak{X}\), \(\left\lVert{\dot{\phi}_\theta}\right\rVert_{1}(v,x)\) is the sum of the absolute values of the \(K^2\) entries of \(\dot{\phi}_\theta(v,x)\).

Let \(A_n\in\mathbb{R}^{K\times K}\) be an arbitrary sequence of (possibly random and then \(\sigma(\mathcal{S}')\)-measurable) matrices with \(A_n\overset{P_{VX}}{\to}A_0\) as \(n\to\infty\) for a symmetric positive definite \(A_0\). Then the solution \(\hat{\theta}\) to \[\tilde{\theta}\mapsto \Lambda_n(\tilde{\theta})\mathrel{\vcenter{:}}=\big(\mathbb{P}_n'\phi_{\tilde{\theta}}^\intercal \big)A_n \big(\mathbb{P}_n'\phi_{\tilde{\theta}}\big)\equiv 0\] up to \(\Lambda_n(\hat{\theta})=o_{P_{VX}}\left(n^{-1/2}\right)\) satisfies \(\sqrt{n}(\hat{\theta}-\theta_0)\overset{P_{VX}}{\rightsquigarrow} \mathcal{N}(0,\Sigma)\) as \(n\to\infty\), where \[\begin{align} \Sigma \mathrel{\vcenter{:}}=({\dot{\Phi}^\intercal}A_0{\dot{\Phi}})^{-1}{\dot{\Phi}^\intercal}A_0\Phi A_0{\dot{\Phi}}({\dot{\Phi}^\intercal}A_0{\dot{\Phi}})^{-1},\quad \dot{\Phi} \mathrel{\vcenter{:}}= P_{VX} \dot{\phi}_{\theta_0}, \quad \Phi \mathrel{\vcenter{:}}= P_{VX}\phi_{\theta_0}\phi_{\theta_0}^\intercal. \end{align}\]

Further, let \(\xi_{\theta}:\mathfrak{V}\times\mathfrak{X}\to\mathbb{R}\), \(\theta\in\Theta\), be (possibly random and then \(\sigma(\mathcal{S}')\)-measurable) functions with the derivative of \(\theta\mapsto \xi_\theta(v,x)\) at \(\tilde{\theta}\), \(\mathrm{D}_\theta\xi_{\tilde{\theta}}(v,x)\), existent at all \(\tilde{\theta}\in\mathrm{Nb}(\theta_0)\) for all \((v,x)\in\mathfrak{V}\times\mathfrak{X}\). If \[\begin{align} \left\lVert{\sup_{\tilde{\theta}\in\mathrm{Nb}(\theta_0)}\left\lVert{\mathrm{D}_\theta \xi_{\tilde{\theta}}}\right\rVert_{2}}\right\rVert_{L_2(P_{VX})}=O_{{P_{VX}}}\left(1\right),\label{priv:lemma:rate95mm95xideriv95bound} \end{align}\qquad{(22)}\] where, for a fixed \((v,x)\in\mathfrak{V}\times\mathfrak{X}\), \(\left\lVert{\mathrm{D}_\theta \xi_{\tilde{\theta}}}\right\rVert_{2}^2(v,x)\) is the sum of squared entries of the vector \(\mathrm{D}_\theta \xi_{\tilde{\theta}}(v,x)\), then \(\left\lVert{\xi_{\hat{\theta}}-\xi_{\theta_0}}\right\rVert_{L_2(P_{VX})}=O_{{P_{VX}}}\left(n^{-1/2}\right)\).

Proof of 7. We have the identification \(P_{VX}\phi_{\theta_0}=0_{K}\). The conditions of the lemma are versions of Assumptions 2.1–2.5 and 3.1–3.6 in [5] adapted to our setting, whereby Theorems 2.1 and 3.2 ibid apply, giving \(\sqrt{n}(\hat{\theta}-\theta_0)\overset{P_{VX}}{\rightsquigarrow}\mathcal{N}(0,\Sigma)\) (see also [6]). Then by the mean-value theorem and ?? , \[\left\lVert{\xi_{\hat{\theta}}-\xi_{\theta_0}}\right\rVert_{L_2(P_{VX})}\leq \left\lVert{\hat{\theta}-\theta_0}\right\rVert_{2}\left\lVert{\sup_{\tilde{\theta}\in\mathrm{Nb}(\theta_0)}\left\lVert{\mathrm{D}_\theta \xi_{\tilde{\theta}}}\right\rVert_{2}}\right\rVert_{L_2(P_{VX})}=O_{{P_{VX}}}\left(n^{-1/2}\right).\] ◻

2 is modelled after 7 but the \(\left\lVert{\cdot}\right\rVert_{L_1(P_{VX})}\) in conditions ?? and ?? replaced with \(\left\lVert{\cdot}\right\rVert_{L_1(P_{VZ})}\) in 2. The restriction of \(\left\lVert{\cdot}\right\rVert_{L_1(P_{VZ})}\)-bounds in ?? and ?? instead of the \(\left\lVert{\cdot}\right\rVert_{L_1(P_{VX})}\)-bounds in the nonprivate case of 7 appears to be unavoidable in our construction: we need to control \(\left\lVert{\cdot}\right\rVert_{L_1(P_{VZ})}\), and while 3 [priv:lem:qtc95implications95norms] shows that \(\left\lVert{h}\right\rVert_{L_p(P_{VX})}\lesssim\left\lVert{h}\right\rVert_{L_p(P_{VZ})}\) when \(Q\in\mathcal{Q}_{\delta}\) in 15 , the converse \(\left\lVert{h}\right\rVert_{L_p(P_{VZ})}\lesssim\left\lVert{h}\right\rVert_{L_p(P_{VX})}\) may fail. One can replace \(\left\lVert{\cdot}\right\rVert_{L_p(P_{VZ})}\) with \(\left\lVert{\cdot}\right\rVert_{\infty}\)-bounds, which could be easier to verify. For example, in estimating \(\mu_{\mathcal{X}}=\mu_{\theta_0}\), if \(m\) is bounded and \(\mu_{\theta}\) is a generalised linear model with continuous second derivative, then the \(\left\lVert{.}\right\rVert_{\infty}\)-bounds hold provided \(\mathfrak{V}_\mathit{1}\times\mathfrak{X}\) is compact.

Proof of 2. Because \(Q_\mathcal{X}^{-1}\) and differentiation with respect to \(\theta\) commutes, \[\begin{align} \bar{\phi}_{\tilde{\theta}}^\intercal=\mathrm{D}_\theta \bar{\Xi}_{\tilde{\theta}}=\mathrm{D}_\theta Q_\mathcal{X}^{-1}\Xi_{\tilde{\theta}}= Q_\mathcal{X}^{-1}\mathrm{D}_\theta \Xi_{\tilde{\theta}}=Q_\mathcal{X}^{-1}\phi_{\tilde{\theta}}^\intercal, \end{align}\] and then also \[\begin{align} \mathrm{D}_\theta \bar{\phi}_{\tilde{\theta}}=\mathrm{D}_\theta Q_\mathcal{X}^{-1}\phi_{\tilde{\theta}}= Q_\mathcal{X}^{-1}\mathrm{D}_\theta\phi_{\tilde{\theta}}=Q_\mathcal{X}^{-1}\dot{\phi}_{\tilde{\theta}}. \end{align}\] (Here, we denote with \(Q_\mathcal{X}^{-1}\phi_{\tilde{\theta}}^\intercal\) the vector and with \(Q_\mathcal{X}^{-1}\dot{\phi}_{\tilde{\theta}}\) the matrix where \(Q_\mathcal{X}^{-1}\) is applied element-wise to the coordinate functions of \(\phi_{\tilde{\theta}}^\intercal\) and \(\dot{\phi}_{\tilde{\theta}}\), respectively.) 3 [priv:lem:linopqx95properties95change] implies the identification \(P_{VZ}\bar{\phi}_{\theta_0}^\intercal=P_{VX}\phi_{\theta_0}^\intercal=0_{K}\), as \(P_{VX}\phi_{\theta_0}^\intercal=0_{K}\) by the identifiability conditions of 7; and also \(P_{VZ}\mathrm{D}_\theta\bar{\phi}_{\tilde{\theta}}=P_{VZ} Q_\mathcal{X}^{-1}\dot{\phi}_{\tilde{\theta}}=P_{VX}\dot{\phi}_{\tilde{\theta}}\). By this, and the strengthening of \(\left\lVert{\cdot}\right\rVert_{L_p(P_{VX})}\) to \(\left\lVert{\cdot}\right\rVert_{L_p(P_{VZ})}\) in ?? ,?? , and ?? , all assumptions in 7 that hold under \((\mathbb{P}_n', P_{VX})\) continue to hold under \((\bar{\mathbb{P}}_n', P_{VZ})\). Thus, the assertion follows with \(\dot{\Phi}\) being the same as in 7. ◻

11.1.2 Regression \(\mu_{\mathcal{X}}\)↩︎

Suppose that \(\mu_{\mathcal{X}}=\mu_{\theta_0}\) belonging to the model \[\begin{align} \mathcal{M}(\Theta)\mathrel{\vcenter{:}}=\left\{\mu_{\theta}:\theta\in\Theta\subset\mathbb{R}^{K}\right\} \label{priv:eq:muXmodel} \end{align}\tag{51}\] for a fixed \(K\geq 1\).

With 7 and 2, we can attain the parametric root-\(n\) rate for the regression in both the nonprivate and the private setting.

Corollary 2 (Regression — Rate of Nonprivate Estimator \(\mu_{\hat{\theta}}\)). Assume that \(\mu_{\mathcal{X}}=\mu_{\theta_0}\) belonging to the model in 51 . For \((v,x,\tilde{\theta})\in\mathfrak{V}\times\mathfrak{X}\times\Theta\), let \(\Xi_{\tilde{\theta}}(v,x)\mathrel{\vcenter{:}}=\Delta_{\mu_{\tilde{\theta}}}^2(v,x)=(m(v,x) - \mu_{\tilde{\theta}}(v_{\mathit{1}},x))^2\), \[\begin{align} \phi_{\tilde{\theta}}(v,x)\mathrel{\vcenter{:}}=\mathrm{D}_{\theta}\Delta_{\mu_{\tilde{\theta}}}^2(v,x)^\intercal=2(m(v,x)-\mu_{\tilde{\theta}}(v_\mathit{1},x))\mathrm{D}_\theta \mu_{\tilde{\theta}}(v_\mathit{1},x)^\intercal, \quad \xi_{\tilde{\theta}}\mathrel{\vcenter{:}}=\mu_{\tilde{\theta}}. \end{align}\]

Take a sequence of (possibly random and then \(\sigma(\mathcal{S}')\)-measurable) matrices \(A_n\in\mathbb{R}^{K\times K}\), and let \(\hat{\theta}\in\Theta\) be the solution to the estimating equation \(\tilde{\theta}\mapsto \Lambda_n(\tilde{\theta})\mathrel{\vcenter{:}}=\big(\mathbb{P}_n' \phi_{\tilde{\theta}}^\intercal\big)A_n\big(\mathbb{P}_n' \phi_{\tilde{\theta}}\big)\equiv 0\) up to \(\Lambda_n(\hat{\theta})=o_{P_{VX}}\left(n^{-1/2}\right)\).

If \((\Xi_\theta, \phi_{\theta},A_n,\xi_\theta)\) satisfy the conditions of 7, then \(\sqrt{n}(\hat{\theta}-\theta_0)\overset{P_{VX}}{\rightsquigarrow} \mathcal{N}(0,\Sigma)\) as \(n\to\infty\) for some \(\Sigma\), and \(\left\lVert{\mu_{\hat{\theta}}-\mu_{\mathcal{X}}}\right\rVert_{L_2(P_{VX})}=O_{{P_{VX}}}\left(n^{-1/2}\right)\).

Corollary 3 (Regression — Rate of Private Estimator \(\mu_{\check{\theta}}\)). Assume that \(P_{VZ}\in\mathcal{P}_{VZ}(Q, \mathcal{P}_{VX})\), for a fixed \(Q\in\mathcal{Q}_{\psi}\) and \(P_{VX}\in\mathcal{P}_{VX}\) subject to the parametric assumption 51 . Let \((\Xi_\theta, \phi_{\theta},\xi_\theta)\), \(\theta\in\Theta\), be as defined in 2.

Take a sequence of (possibly random and then \(\sigma(\bar{\mathcal{S}}')\)-measurable) matrices \(\bar{A}_n\in\mathbb{R}^{K\times K}\), and let \(\check{\theta}\in\Theta\) be the solution to the estimating equation \(\tilde{\theta}\mapsto \bar{\Lambda}_n(\tilde{\theta})\mathrel{\vcenter{:}}=\big(\bar{\mathbb{P}}_n' \bar{\phi}_{\tilde{\theta}}^\intercal\big)\bar{A}_n \big(\bar{\mathbb{P}}_n' \bar{\phi}_{\tilde{\theta}}\big)\equiv 0\) up to \(\bar{\Lambda}_n(\check{\theta})=o_{P_{VZ}}\left(n^{-1/2}\right)\), where \(\bar{\phi}_{\tilde{\theta}}\mathrel{\vcenter{:}}=\mathrm{D}_{\theta}\bar{\Xi}_{\tilde{\theta}}\) with \(\bar{\Xi}_{\theta}\mathrel{\vcenter{:}}= Q_\mathcal{X}^{-1}\Xi_{\theta}\) for the inverse \(Q_\mathcal{X}^{-1}\) of \(Q_\mathcal{X}\) in 16 .

If \((\Xi_\theta, \phi_{\theta}, \bar{A}_n, \xi_\theta)\) satisfy the conditions of 2, then \(\sqrt{n}(\check{\theta}-\theta_0)\overset{P_{VZ}}{\rightsquigarrow} \mathcal{N}(0,\bar{\Sigma})\) as \(n\to\infty\) for some \(\bar{\Sigma}\), and \(\left\lVert{\mu_{\check{\theta}}-\mu_{\mathcal{X}}}\right\rVert_{L_2(P_{VX})}=O_{{P_{VZ}}}\left(n^{-1/2}\right)\).

The proofs of [priv:cor:muXpara_rate,priv:cor:muXpara_rate_priv] are omitted.

11.1.3 Riesz Representer \(r\)↩︎

In this section, we consider the estimation of the Riesz representer \(r\). We begin with an identification result, which is not limited to finite-dimensional models (see also [3]).

Lemma 8 (Identification of \(r\)). For all \(\gamma\in\Gamma\), the Riesz representer \(r_\gamma\) of \(L_2(P_{V_\mathit{1}}X)\ni\mu\mapsto P_{VX} f(V,X,\mu, \gamma)\) satisfies 31 .

Proof of 8. By the definition of \(r_\gamma\), \[\begin{align} P_{VX}f(V,X,h,\gamma)=P_{VX}r_\gamma(V_\mathit{1},X)h(V_\mathit{1},X) \text{ for all } h \in L_2(P_{V_\mathit{1}}X). \end{align}\] Then \[\begin{align} P_{VX}\Upsilon_{\gamma,h}(V,X)=P_{VX}\left[h^2-2f\right]=P_{VX}\left[h^2-2hr_\gamma\right]=P_{VX}\left[(h-r_\gamma)^2 - r_\gamma^2 \right], \end{align}\] whereby \(\arg\min_{h\in L_2(P_{V_\mathit{1}}X)}P_{VX}\Upsilon_{\gamma,h}=\arg\min_{h\in L_2(P_{V_\mathit{1}}X)}P_{VX}\left[(h-r_\gamma)^2\right]=r_\gamma\). ◻

Suppose that the Riesz representer \(r_{\gamma}\) of \(\mu\mapsto P_{VX}f(V,X,\mu,\gamma)\) is \(r_{\gamma}=r_{\gamma,\theta_0}\) for each \(\gamma\in\Gamma\) uniformly, where \(r_{\gamma,\theta_0}\) belongs to the model \[\begin{align} \mathcal{R}_\gamma(\Theta)\mathrel{\vcenter{:}}=\left\{r_{\gamma,\theta}:\theta\in\Theta\subset\mathbb{R}^{K}\right\}, \label{priv:eq:rieszXmodelgamma} \end{align}\tag{52}\] with the map \((\gamma,\theta)\mapsto r_{\gamma,\theta}\) known. A prototypical example is a parametric model for the propensity score in inferring the average treatment effect on the treated.

Example 10 (name=Average Treatment Effect on the Treated,continues=priv:ex:att_dr). Recall that for \(\gamma_{\mathcal{V}}(c)\mathrel{\vcenter{:}}=\bar{\gamma}\mathrel{\vcenter{:}}=\mathbb{E}D=p_1\), the Riesz representer for \(\mathbb{E}\left[\left. Y^0\,\right\vert\, D=1\right]\) is \(r(d,x)=\frac{1-d}{\gamma_{\mathcal{V}}(c)}\frac{\pi_{\mathcal{X}}(1{\vert\,}x)}{1-\pi_{\mathcal{X}}(1{\vert\,}x)}\), and for \(\mathbb{E}\left[\left. Y^1-Y^0\,\right\vert\, D=1\right]\), it is \(r(d,x)=\frac{d}{\gamma_{\mathcal{V}}(c)}-\frac{1-d}{\gamma_{\mathcal{V}}(c)}\frac{\pi_{\mathcal{X}}(1{\vert\,}x)}{1-\pi_{\mathcal{X}}(1{\vert\,}x)}\). Suppose that the propensity score \(\pi_{\mathcal{X}}(1{\vert\,}x)=\pi_{\theta_0}(x)\) for some \(\pi_{\theta_0}\) in the presumed model \(\left\{\pi_{\theta}:\theta\in\Theta\subset\mathbb{R}^{K}\right\}\), for instance, the logistic model \(\pi_\theta(x)=(1+\exp(-\theta^\intercal x))^{-1}\). Then \[\begin{align} r_{\gamma,\theta_0}(d,x) =\frac{1-d}{\gamma}\frac{\pi_{\theta_0}(x)}{1-\pi_{\theta_0}(x)},\quad r_{\gamma,\theta_0}(d,x) =\frac{d}{\gamma}-\frac{1-d}{\gamma}\frac{\pi_{\theta_0}(x)}{1-\pi_{\theta_0}(x)}. \end{align}\] are the Riesz representers for \(\mathbb{E}\left[\left. Y^0\,\right\vert\, D=1\right]\) and \(\mathbb{E}\left[\left. Y^1-Y^0\,\right\vert\, D=1\right]\), respectively, for \(\gamma=\gamma_{\mathcal{V}}(c)\) \(=\bar{\gamma}=\mathbb{E}D\). Conversely, \(r_{\gamma_{\mathcal{V}}(c),\theta}\) is not a Riesz representer unless \(\theta=\theta_0\). Let \(\lambda_{0,\theta}(d,x)\mathrel{\vcenter{:}}=(1-d)\frac{\pi_\theta(x)}{1-\pi_\theta(x)}\) and \(\lambda_{1,\theta}(d,x)\mathrel{\vcenter{:}}= d- (1-d)\frac{\pi_\theta(x)}{1-\pi_\theta(x)}\). Then the models for the Riesz representers are \[\begin{align} \mathcal{R}_\gamma(\Theta)=\left\{\gamma^{-1}\lambda_{0,\theta}:\theta\in\Theta\right\}, \quad \mathcal{R}_\gamma(\Theta)=\left\{\gamma^{-1}\lambda_{1,\theta}:\theta\in\Theta\right\} \end{align}\] for \(\mathbb{E}\left[\left. Y^0\,\right\vert\, D=1\right]\) and \(\mathbb{E}\left[\left. Y^1-Y^0\,\right\vert\, D=1\right]\), respectively.

In the parametric model 52 , the Riesz representer \(r_{\gamma,\theta_0}\) is identified by 8 as \[\begin{align} r_{\gamma,\theta_0}&= \arg\min_{\rho \in\mathcal{R}_\gamma(\Theta)} P_{VX}\Upsilon_{\gamma,\rho}(V,X) \nonumber \\ &= \arg\min_{\rho\in\mathcal{R}_\gamma(\Theta)}P_{VX}\bigg[\rho(v_\mathit{1},x)^2-2f(V,X,\rho,\gamma)\bigg]. \label{priv:eq:rieszgamma95para} \end{align}\tag{53}\] Importantly, for all \(\gamma\in\Gamma\) uniformly, \(\theta=\theta_0\), and only \(\theta=\theta_0\), gives the Riesz representer in the model \(\mathcal{R}_\gamma(\Theta)\). An implication is that for any \(\gamma\in\Gamma\), \[\begin{align} \theta_0 = \arg\min_{\theta\in\Theta}P_{VX}\Upsilon_{\gamma,r_{\gamma,\theta}}(V,X)= \arg\min_{\theta\in\Theta}P_{VX}\bigg[r_{\gamma,\theta}(v_\mathit{1},x)^2-2f(V,X,r_{\gamma,\theta},\gamma)\bigg].\label{priv:eq:rieszgamma95para95theta} \end{align}\tag{54}\] This permits us to infer \(\theta_0\) without suffering any bias from the unknown \(\gamma_{\mathcal{V}}(c)\). Take an arbitrary \(\gamma_0\in\Gamma\). In the nonprivate setting, 54 can directly be used to construct a method-of-moments estimator \(\hat{\theta}_{\gamma_0}\) of \(\theta_0\), and next estimate the representer as \(\hat{r}\mathrel{\vcenter{:}}= r_{\hat{\gamma}_{\mathcal{V}}(c),\hat{\theta}_{\gamma_0}}\) for \(\hat{\gamma}_{\mathcal{V}}(c)\) of 66 . In the private setting, 2 can be used to construct a private estimator \(\check{\theta}_{\gamma_0}\) of \(\theta_0\), and then estimate the representer as \(\check{r}\mathrel{\vcenter{:}}= r_{\check{\gamma}_{\mathcal{V}}(c),\check{\theta}_{\gamma_0}}\) for \(\check{\gamma}_{\mathcal{V}}(c)\) of 24 . If \((\gamma,\theta)\mapsto r_{\gamma,\theta}\) is smooth enough, then analogues to [priv:cor:muXpara_rate,priv:cor:muXpara_rate_priv] hold, giving root-\(n\) rates.

Corollary 4 (Riesz Representer — Rate of Nonprivate Estimator \(r_{\hat{\gamma}_{\mathcal{V}}(c),\hat{\theta}_{\gamma_0}}\)). Suppose that \(r_{\gamma}=r_{\gamma,\theta_0}\) for the model \(\mathcal{R}_\gamma(\Theta)\) of 52 . Assume that

  1. The map \((\gamma,\theta)\mapsto r_{\gamma,\theta}(v_\mathit{1},x)\) is differentiable at all \((\gamma,\theta)\in\mathrm{Nb}(\gamma_{\mathcal{V}}(c))\times\mathrm{Nb}(\theta_0)\), where the \(\mathrm{Nb}(w)\) are neighbourhoods of \(w\), for all \((v_\mathit{1},x)\in\mathfrak{V}_\mathit{1}\times\mathfrak{X}\) with partial derivatives \(\partial_\gamma r_{\tilde{\gamma},\tilde{\theta}}(v_\mathit{1},x)\) with respect to \(\gamma\), and \(\mathrm{D}_\theta r_{\tilde{\gamma},\tilde{\theta}}(v_\mathit{1},x)\) with respect to \(\theta\), at \((\tilde{\gamma},\tilde{\theta})\in\mathrm{Nb}(\gamma_{\mathcal{V}}(c))\times\mathrm{Nb}(\theta_0)\).

  2. The map \(\theta\mapsto \Upsilon_{\gamma,r_{\gamma,\theta}}(v,x)=r_{\gamma,\theta}(v_\mathit{1},x)^2-2f(v,x,r_{\gamma,\theta},\gamma)\) is differentiable at all \(\theta\in\mathrm{Nb}(\theta_0)\) for all \((v,x,\gamma)\in\mathfrak{V}\times\mathfrak{X}\times\Gamma\) with \(\mathbb{R}^{ K\times 1}\)-valued derivative \(\mathrm{D}_\theta \Upsilon_{\gamma,r_{\gamma,\tilde{\theta}}}(v,x)^\intercal\eqqcolon\phi_{\gamma,\tilde{\theta}}(v,x)\) at \(\tilde{\theta}\in\mathrm{Nb}(\theta_0)\).

Fix an arbitrary \(\gamma_0\in\Gamma\), and take a sequence of (possibly random and then \(\sigma(\mathcal{S}')\)-measurable) matrices \(A_n\in\mathbb{R}^{K\times K}\). Let \(\hat{\theta}_{\gamma_0}\in\Theta\) be the solution to the estimating equation \(\tilde{\theta}\mapsto \Lambda_n(\tilde{\theta})\mathrel{\vcenter{:}}=\big(\mathbb{P}_n' \phi_{\gamma_0,\tilde{\theta}}^\intercal\big)A_n\big(\mathbb{P}_n' \phi_{\gamma_0,\tilde{\theta}}\big) \equiv 0\) up to \(\Lambda_n(\hat{\theta}_{\gamma_0})=o_{P_{VX}}\left(n^{-1/2}\right)\). If \(\Xi_\theta\mathrel{\vcenter{:}}=\Upsilon_{\gamma_0,r_{\gamma_0,\theta}}\), \(\phi_{\theta}\mathrel{\vcenter{:}}=\phi_{\gamma_0,\theta}\), and \(A_n\) satisfy the conditions of 7 pertaining to \((\Xi_\theta,\phi_\theta, A_n)\) therein, then \(\sqrt{n}(\hat{\theta}_{\gamma_0}-\theta_0)\overset{P_{VX}}{\rightsquigarrow} \mathcal{N}(0,\Sigma_{\gamma_0})\) as \(n\to\infty\) for some \(\Sigma_{\gamma_0}\). If, in addition, \[\begin{align} \left\lVert{\partial_\gamma r_{\tilde{\gamma}_{\mathcal{V}}(c),\tilde{\theta}_{\gamma_0}}}\right\rVert_{L_2(P_{VX})}&=O_{{P_{VX}}}\left(1\right), \label{priv:cor:rieszXpara95rate:diff95riesz95bound95gamma} \\ \left\lVert{\left\lVert{\mathrm{D}_\theta r_{\tilde{\gamma}_{\mathcal{V}}(c),\tilde{\theta}_{\gamma_0}}}\right\rVert_{2}}\right\rVert_{L_2(P_{VX})}&=O_{{P_{VX}}}\left(1\right) \label{priv:cor:rieszXpara95rate:diff95riesz95bound95theta} \end{align}\] {#eq: sublabel=eq:priv:cor:rieszXpara95rate:diff95riesz95bound95gamma,eq:priv:cor:rieszXpara95rate:diff95riesz95bound95theta} for some \((\tilde{\gamma}_{\mathcal{V}}(c),\tilde{\theta}_{\gamma_0})\) between \((\gamma_{\mathcal{V}}(c),\theta_0)\) and \((\hat{\gamma}_{\mathcal{V}}(c),\hat{\theta}_{\gamma_0})\), then \(\left\lVert{\hat{r}-r}\right\rVert_{L_2(P_{VX})}=O_{{P_{VX}}}\left(n^{-1/2}\right)\), where \(\hat{r}\mathrel{\vcenter{:}}= r_{\hat{\gamma}_{\mathcal{V}}(c),\hat{\theta}_{\gamma_0}}\), \(r=r_{\gamma_{\mathcal{V}}(c),\theta_0}\) for the model in 52 and \(\hat{\gamma}_{\mathcal{V}}(c)\) in 66 .

Proof of 4. The asymptotic normality of \(\hat{\theta}_{\gamma_0}\) follows directly from 7 via the identification 54 , whereby \(P_{VX}\phi_{\gamma,\theta_0}=0_{K }\) for any fixed \(\gamma\in\Gamma\). A mean-value expansion of \(r_{\hat{\gamma}_{\mathcal{V}}(c),\hat{\theta}_{\gamma_0}}\) via [priv:cor:rieszXpara95rate:diff95riesz] in combination with ?? and ?? yields the second assertion as \(\hat{\gamma}_{\mathcal{V}}(c)-\gamma_{\mathcal{V}}(c)=O_{{P_{VX}}}\left(n^{-1/2}\right)\), and \(\hat{\theta}_{\gamma_0}-\theta_0=O_{{P_{VX}}}\left(n^{-1/2}\right)\) by the first assertion. ◻

Corollary 5 (Riesz Representer — Rate of Private Estimator \(r_{\check{\gamma}_{\mathcal{V}}(c),\check{\theta}}\)). Assume that \(P_{VZ}\in\mathcal{P}_{VZ}(Q, \mathcal{P}_{VX})\), for a fixed \(Q\in\mathcal{Q}_{\psi}\) and \(P_{VX}\in\mathcal{P}_{VX}\) subject to the parametric assumption of 52 . Suppose that [priv:cor:rieszXpara95rate:diff95riesz] and [priv:cor:rieszXpara95rate:diff95upsilon] of 4 hold, and let \(\Upsilon_{\gamma,r_{\gamma,\theta}}, \phi_{\gamma,\theta}\), \((\gamma,\theta)\in\Gamma\times\Theta\), be as they are defined therein.

Fix an arbitrary \(\gamma_0\in\Gamma\), and take a (possibly random and then \(\sigma(\bar{\mathcal{S}}')\)-measurable) sequence of matrices \(\bar{A}_n\in\mathbb{R}^{K\times K}\). Let \(\check{\theta}_{\gamma_0}\in\Theta\) be the solution to the estimating equation \(\tilde{\theta}\mapsto \bar{\Lambda}_n(\tilde{\theta})\mathrel{\vcenter{:}}=\big(\bar{\mathbb{P}}_n' \bar{\phi}_{\gamma_0,\tilde{\theta}}^\intercal\big)\bar{A}_n \big(\bar{\mathbb{P}}_n' \bar{\phi}_{\gamma_0,\tilde{\theta}}\big)\equiv 0\), up to \(\bar{\Lambda}_n(\check{\theta}_{\gamma_0})=o_{P_{VZ}}\left(n^{-1/2}\right)\), with \(\bar{\phi}_{\gamma,\tilde{\theta}}\mathrel{\vcenter{:}}=\mathrm{D}_\theta \bar{\Upsilon}_{\gamma,r_{\gamma,\tilde{\theta}}}\), \(\bar{\Upsilon}_{\gamma,r_{\gamma,\tilde{\theta}}} \mathrel{\vcenter{:}}= Q_\mathcal{X}^{-1}\Upsilon_{\gamma,r_{\gamma,\tilde{\theta}}}\), for the inverse \(Q_\mathcal{X}^{-1}\) of \(Q_\mathcal{X}\) in 16 .

If \(\Xi_\theta\mathrel{\vcenter{:}}=\Upsilon_{\gamma_0,r_{\gamma_0,\theta}}\), \(\phi_\theta\mathrel{\vcenter{:}}=\phi_{\gamma_0,\theta}\), and \(\bar{A}_n\) satisfy the conditions of 2 pertaining to \((\Xi_\theta, \phi_\theta, \bar{A}_n)\) therein, then \(\sqrt{n}(\check{\theta}_{\gamma_0}-\theta_0)\overset{P_{VZ}}{\rightsquigarrow} \mathcal{N}(0,\bar{\Sigma}_{\gamma_0})\) as \(n\to\infty\) for some \(\bar{\Sigma}_{\gamma_0}\). If, in addition, \(\left\lVert{\partial_\gamma r_{\tilde{\gamma}_{\mathcal{V}}(c),\tilde{\theta}_{\gamma_0}}}\right\rVert_{L_2(P_{VX})}=O_{{P_{VZ}}}\left(1\right)\) and \(\left\lVert{\left\lVert{\mathrm{D}_\theta r_{\tilde{\gamma}_{\mathcal{V}}(c),\tilde{\theta}_{\gamma_0}}}\right\rVert_{2}}\right\rVert_{L_2(P_{VX})}=O_{{P_{VZ}}}\left(1\right)\) for some \((\tilde{\gamma}_{\mathcal{V}}(c),\tilde{\theta}_{\gamma_0})\) between \((\gamma_{\mathcal{V}}(c),\theta_0)\) and \((\check{\gamma}_{\mathcal{V}}(c),\check{\theta}_{\gamma_0})\), then \(\left\lVert{\check{r}-r}\right\rVert_{L_2(P_{VX})}=O_{{P_{VZ}}}\left(n^{-1/2}\right)\), where \(\check{r}\mathrel{\vcenter{:}}= r_{\check{\gamma}_{\mathcal{V}}(c),\check{\theta}_{\gamma_0}}\), \(r=r_{\gamma_{\mathcal{V}}(c),\theta_0}\) for the model in 52 and \(\check{\gamma}_{\mathcal{V}}(c)\) in 24 .

Proof of 5. Follows from 2 as 4 follows from 7. ◻

11.2 Infinite-Dimensional Models↩︎

Proof of 3. Write, for \(T_n\) in ?? , \[\begin{align} \check\theta(v_\mathit{1},x)-\theta(v_\mathit{1},x) =&\, T_n(v_\mathit{1}, x) + M_{n}(v_\mathit{1}, x) + B_{n}(v_\mathit{1}, x), \\ M_{n}(v_\mathit{1}, x) \mathrel{\vcenter{:}}=&\, \frac{1}{n}\sum_{i\in[n]}\left\{\bar{w}_{n,i}(v_\mathit{1},x,V_i',Z_i',\vartheta)-P_{VZ}[\bar{w}_{n,i}(v_\mathit{1},x,V,Z,\vartheta)]\right\}, \\ B_{n}(v_\mathit{1}, x) \mathrel{\vcenter{:}}=&\, \frac{1}{n}\sum_{i\in[n]}P_{VZ}[\bar{w}_{n,i}(v_\mathit{1},x,V,Z,\vartheta)]-\theta(v_\mathit{1},x), \end{align}\] for \((v_\mathit{1},x)\in \mathfrak{V}_{\mathit{1}}\times\mathfrak{X}\), so that \[\begin{align} P_{VX}(\check\theta-\theta)^2 \leq 2\left\{P_{VX}T_n^2 + 2(P_{VX}M_{n}^2+P_{VX}B_{n}^2)\right\}, \end{align}\] since \((x+y)^2\leq2(x^2+y^2)\) for all \(x,y\in\mathbb{R}\). For \(\bar{w}_{ni}\) in 33 , we have by 3 [priv:lem:linopqx95properties95change] that \(P_{VZ}[\bar{w}_{n,i}(v_\mathit{1},x,V,Z,\vartheta)]=P_{VX}[w_{n,i}(v_\mathit{1},x,V,X,\vartheta)]\). Then, in the light of ?? , the bound ?? follows once we show \[\begin{align} P_{VX}M_n^2=O_{{P_{VZ}}}\left(\frac{1}{n^2}\sum_{i\in[n]}\int \mathop{\mathrm{\mathbb{V}}}\left[\bar{w}_{n,i}(v_\mathit{1},x,V,Z,\vartheta)\right] \,\mathrm{d}P_{V_{\mathit{1}}X}(v_{\mathit{1}},x)\right). \label{priv:eq:varianceorder} \end{align}\tag{55}\] By Markov’s inequality, for all \(K>0\), \[\begin{align} \mathbb{P}\left(P_{VX}M_{n}^2>K\right)\leq \frac{1}{K}\mathbb{E}\left[\int M_{n}(v_\mathit{1},x)^2\,\mathrm{d}P_{V_\mathit{1}X}(v_\mathit{1},x)\right]\\ = \frac{1}{K}\int \mathbb{E}\left[M_{n}(v_\mathit{1},x)^2\right]\,\mathrm{d}P_{V_\mathit{1}X}(v_\mathit{1},x). \end{align}\] For any fixed \((v_\mathit{1},x)\), \(\mathbb{E}\left[M_{n}(v_\mathit{1},x)^2\right]=\frac{1}{n^2}\sum_{i\in[n]}\mathop{\mathrm{\mathbb{V}}}\left[\bar{w}_{n,i}(v_\mathit{1},x,V,Z,\vartheta)\right]\), because the \((V_i',Z_i)\) are i.i.d.. Hence 55 holds.

Bound ?? . If \(Q\in\mathcal{Q}_{\delta}\), then, by 3, we have for \(h\in L_2(P_{VZ})\), \[\begin{align} \mathop{\mathrm{\mathbb{V}}}\left[Q_\mathcal{X}^{-1}h(V,Z)\right]\leq 2 \left\{\frac{1}{\alpha^2}\mathop{\mathrm{\mathbb{V}}}\left[h(V,Z)\right]+\left(\frac{1-\alpha}{\alpha}\right)^2\mathop{\mathrm{\mathbb{V}}}\left[\int h(V,z)\bar{Q}(\mathrm{d}z)\right] \right\}, \end{align}\] since for random variables \(X,Y\), we have \(\mathop{\mathrm{\mathbb{V}}}\left[X+Y\right]\leq 2(\mathop{\mathrm{\mathbb{V}}}\left[X\right]+\mathop{\mathrm{\mathbb{V}}}\left[Y\right])\). Clearly, \(\mathop{\mathrm{\mathbb{V}}}\left[h(V,Z)\right]\leq P_{VZ}h^2\), and for the same reason, \(\mathop{\mathrm{\mathbb{V}}}\left[\int h(V,z)\bar{Q}(\mathrm{d}z)\right]\leq (P_V \otimes \bar{Q})h^2\) by Jensen’s inequality. But \(P_V \otimes \bar{Q}\leq \frac{1}{1-\alpha}P_{VZ}\) by 4, so ?? holds. ◻

In the following, we apply 3 to various estimators. The regression \(\mu_{\mathcal{X}}\) is estimated by kernel and orthogonal series estimators in [priv:ex:regression_nw_kernel,priv:ex:regression_ortho], respectively. The Riesz representer \(r\) is estimated by orthogonal series in 13.

Example 11 (name=Regression \(\mu_{\mathcal{X}}\) — Nadaraya–Watson Kernel Estimator). Suppose that \(\mathfrak{V}_{\mathit{1}}\times\mathfrak{X}\subset \mathbb{R}^{d_{\mathit{1}}}\times\mathbb{R}^{d_\mathcal{X}}\) for fixed positive integers \(d_\mathit{1},d_\mathcal{X}\) and that the mechanism \(Q\in\mathcal{Q}_{\delta}\) in 15 . To estimate \(\mu_{\mathcal{X}}\), consider the nonprivate estimator \[\begin{align} {\hat{\mu}_{\mathcal{X}}}(v_\mathit{1},x)\mathrel{\vcenter{:}}=\frac{1}{n}\sum_{i\in[n]} &\frac{K_h(v_\mathit{1},x,V_{\mathit{1}i}',X_i')m(V_i',X_i')}{\frac{1}{n}\sum_{j\in[n]}K_h(v_\mathit{1},x,V_{\mathit{1}j'},X_j')},\\ K_h(v_\mathit{1},x,V_{\mathit{1}i},X_i)&\mathrel{\vcenter{:}}= K\left(\frac{1}{h}\left((v_\mathit{1},x)-(V_{\mathit{1}i},X_i)\right)\right) \end{align}\] for a bandwidth \(h>0\), and a kernel \(K:\mathbb{R}^d\to\mathbb{R}\) with \(d\mathrel{\vcenter{:}}= d_{\mathit{1}}+d_\mathcal{X}\). For the density estimator \(\hat{p}_{V_\mathit{1}X}(v_\mathit{1},x)\mathrel{\vcenter{:}}=\frac{1}{nh^d}\sum_{i\in[n]}K_h(v_\mathit{1},x,V_{\mathit{1}i}',X_i')\), \({\hat{\mu}_{\mathcal{X}}}\) is of the form 32 for \(w_{n,i}(v_\mathit{1}, x, V',X',\vartheta)=w_{n,i}(v_\mathit{1}, x, V',X', p_{V_\mathit{1}X})\mathrel{\vcenter{:}}=\frac{K_h(v_\mathit{1}, x, V_\mathit{1}',X')m(V',X')}{h^d p_{V_\mathit{1}X}(v_\mathit{1},x)}\).

The private estimator \({\check{\mu}_{\mathcal{X}}}\) is defined as in 33 , with \(\check p_{V_\mathit{1}X}\) an arbitrary estimator of \(p_{V_\mathit{1}X}\) constructed from \(\bar{\mathcal{S}}'=((V_i',Z_i'))_{i\in[n]}\). For instance, if \(\bar{Q}\) of \(Q\) in 15 admits a density \(\bar{q}\) with respect to the same dominating measure as \(P_X\) does, then 4 [priv:lem:qtc95implications95dens] motivates the choice \[\begin{align} \label{priv:eq:estim95pvx95priv} \check p_{V_\mathit{1}X}(v_\mathit{1},x)\mathrel{\vcenter{:}}=\frac{1}{\alpha}\check{p}_{V_\mathit{1}Z}(v_\mathit{1},x)-\frac{1-\alpha}{\alpha}\check p_{V_\mathit{1}}(v)\bar{q}(x), \end{align}\qquad{(23)}\] where \((\check p_{V_\mathit{1}Z},\check p_{V_\mathit{1}})\) is an arbitrary estimator of \((p_{V_\mathit{1}Z}, p_{V_\mathit{1}})\) constructed from \(\bar{\mathcal{S}}'\); for example \[\begin{align} \label{priv:eq:kernel95pvz} \begin{aligned} \check p_{V_\mathit{1}Z}(v_\mathit{1},z)\mathrel{\vcenter{:}}=\frac{1}{nh^d}\sum_{i\in[n]}K_h(v_\mathit{1},x,V_{\mathit{1}i}',Z_i'), \\ \check p_{V_\mathit{1}}(v_\mathit{1})\mathrel{\vcenter{:}}=\frac{1}{nh_\mathit{1}^{d_\mathit{1}}}\sum_{i\in[n]}K_{\mathit{1},h_\mathit{1}}(v_\mathit{1},V_{\mathit{1}i}'), \end{aligned} \end{align}\qquad{(24)}\] where \(K_{\mathit{1},h}(v_\mathit{1},V_\mathit{1})\mathrel{\vcenter{:}}= K_\mathit{1}((v_\mathit{1}-V_\mathit{1})/h_\mathit{1})\) for a kernel \(K_\mathit{1}:\mathbb{R}^{d_\mathit{1}}\to\mathbb{R}\) and bandwidth \(h_\mathit{1}>0\). Next, we study \({\check{\mu}_{\mathcal{X}}}\) in relation to 3.

Term \(T_n\) in ?? . Let \(g_h(v_\mathit{1},x,v',x')\mathrel{\vcenter{:}}= K_h(v_\mathit{1},x,v_\mathit{1}',x')m(v',x')\) and \(\bar{g}_h(v_\mathit{1},x,v',z')\mathrel{\vcenter{:}}=(Q_\mathcal{X}^{-1}(v',x')\mapsto g_h(v_\mathit{1},x,v',x'))(v',z')\). Then \[\begin{align} T_n(v_\mathit{1},x)=\frac{p_{V_\mathit{1}X}(v_\mathit{1},x)-\check p_{V_\mathit{1}X}(v_\mathit{1},x)}{p_{V_\mathit{1}X}(v_\mathit{1},x)}\left(\frac{1}{\check p_{V_\mathit{1}X}(v_\mathit{1},x)}\frac{1}{nh^d}\sum_{i\in[n]}\bar{g}_h(v_\mathit{1},x,V_i',Z_i')\right). \end{align}\] Assume that the \((v_\mathit{1},x)\)-supremum of the bracketed factor is \(O_{{P_{VZ}}}\left(1\right)\) independent of \(h\), and \(\left\lVert{\frac{1}{p_{V_\mathit{1}X}}}\right\rVert_{\infty}<\infty\). Then \(P_{VX}T_n^2\) is dominated by \(P_{VX}(p_{V_\mathit{1}X}-\check p_{V_\mathit{1}X})^2\). Suppose that \(\check p_{V_\mathit{1}X}(v_\mathit{1},x)\) is defined as ?? and ?? , and that \(p_{V_\mathit{1}Z}(v_\mathit{1},z)=\alpha p_{V_\mathit{1}X}(v_\mathit{1},z)+(1-\alpha)p_{V_\mathit{1}}(v_\mathit{1})\bar{q}(z)\) has \(\beta\) continuous derivatives which are uniformly bounded. Further suppose that the kernel \(K\) is chosen suitably, so that \(K\) is of order \(\beta\), satisfying \(\int K(u)\,\mathrm{d}u=1\), \(\int u^j K(u)\,\mathrm{d}u = 0\) if \(|j|<\beta\) and \(\int u^j K(u)\,\mathrm{d}u \neq 0\) if \(|j|=\beta\), where \(u^j\) is the multi-index notation \(u^j\mathrel{\vcenter{:}}= u_1^{j_1}u_2^{j_1}\cdots u_d^{j_d}\) for nonnegative integers \(j_1,j_2,\ldots,j_d\) with \(|j|\mathrel{\vcenter{:}}=\sum_{k=1}^d j_k\). Then we expect \(P_{VX}(p_{V_\mathit{1}Z}-\check p_{V_\mathit{1}Z})^2=O_{{P_{VZ}}}\left(\frac{1}{nh^d}+h^{2\beta}\right)\), and, consequently, \[P_{VX}(p_{V_\mathit{1}X}-\check p_{V_\mathit{1}X})^2=O_{{P_{VZ}}}\left(\frac{1}{\alpha^2}\left(\frac{1}{nh^d}+h^{2\beta}\right)\right)\] under suitable kernel \(K_\mathit{1}\) and bandwidth \(h_\mathit{1}\), because the error of \(\check p_{V_\mathit{1}X}\) is dominated by the error of \(\check p_{V_\mathit{1}Z}\) as opposed to that of \(\check p_{V_\mathit{1}}\), since, clearly, \(V_\mathit{1}\) is lower dimensional than \((V_\mathit{1},Z)\) and \(\bar{q}\) is known.

Turning to the variance term in ?? , \(\mathop{\mathrm{\mathbb{V}}}\left[\bar{w}_{n,i}(v_\mathit{1},x,V,Z)\right]=\frac{1}{h^{2d}p_{V_\mathit{1}X}^2(v_\mathit{1},x)}\mathop{\mathrm{\mathbb{V}}}\left[\bar{g}_h(v_\mathit{1},x,V,Z)\right]\). By ?? , \[\begin{align} (\alpha^2/4)\mathop{\mathrm{\mathbb{V}}}\left[\bar{g}_h(v_\mathit{1},x,V,Z)\right]\leq \int \mathbb{E}\left[K_h^2(v_\mathit{1},x,V_\mathit{1},Z)m^2(V,Z)\right]\,\mathrm{d}P_{V_{\mathit{1}}X}(v_{\mathit{1}},x) \\ = \mathbb{E}\left[m^2(V,Z)\int {K_h^2(v_\mathit{1},x,V_\mathit{1},Z)}\,\mathrm{d}P_{V_{\mathit{1}}X}(v_{\mathit{1}},x)\right] = h^d \mathbb{E}m^2(V,Z)K^2(V_\mathit{1},Z), \end{align}\] where the last step is by a change of variables and using that \(p_{V_{\mathit{1}}X}\) integrates to one. Hence, if \(\left\lVert{\frac{1}{p_{V_\mathit{1}X}}}\right\rVert_{\infty}<\infty\) and \(P_{VZ} (mK)^2<\infty\), then \(\frac{1}{n^2}\sum_{i\in[n]}\int \bar{\sigma}_i^2(v_\mathit{1},x) \,\mathrm{d}P_{V_{\mathit{1}}X}(v_{\mathit{1}},x)\lesssim\frac{1}{\alpha^2 nh^d}\).

The bias term in ?? satisfies \[\begin{align} \frac{1}{n}\sum_{i\in[n]}P_{VX}[w_{n,i}(v_\mathit{1},x,V,X,\vartheta)]-\theta(v_\mathit{1},x)=\frac{P_{VX}K_h(v_\mathit{1},x, V_\mathit{1}, X)m(V,X)}{h^d p_{V_\mathit{1}X}(v_\mathit{1},x)}-\mu_{\mathcal{X}}(v_\mathit{1},x) \\ =\frac{P_{VX}K_h(v_\mathit{1},x, V_\mathit{1}, X)\mu_{\mathcal{X}}(V_\mathit{1},X)-h^d p_{V_\mathit{1}X}(v_\mathit{1},x)\mu_{\mathcal{X}}(v_\mathit{1},x)}{h^d p_{V_\mathit{1}X}(v_\mathit{1},x)}, \end{align}\] by the tower property of expectations and the definition of \(\mu_{\mathcal{X}}\). Let \(g(w)\mathrel{\vcenter{:}}= p_{V_\mathit{1}X}(w)\mu_{\mathcal{X}}(w)\) for \(w\mathrel{\vcenter{:}}=(v_\mathit{1},x)\). If \(g\) has \(\beta'\) continuous derivatives which are uniformly bounded, then a change-of-variables and a Taylor-expansion argument show that, if \(\left\lVert{\frac{1}{p_{V_\mathit{1}X}}}\right\rVert_{\infty}<\infty\), then the \(P_{VX}\)-integrated square of the last display is of the order \(h^{2\beta'}\). Then the bias term in ?? is \(O_{{P_{VZ}}}\left(h^{2\beta'}\right)\).

Conclude that \[\begin{align} P_{VX}({\check{\mu}_{\mathcal{X}}}-\mu_{\mathcal{X}})^2=O_{{P_{VZ}}}\left(\frac{1}{\alpha^2}\left(\frac{1}{nh^d}+h^{2\beta}\right)\right) + O_{{P_{VZ}}}\left(\frac{1}{\alpha^2 n h^d}\right)+O_{{P_{VZ}}}\left(h^{2\beta'}\right), \end{align}\] which, apart from the \(\alpha^{-2}\) factor deriving from privacy, is the usual nonprivate error rate of kernel estimators for \(\beta=\beta'\). The privacy mechanism does introduce a bias through the estimation of the secondary-nuisance parameter \(p_{V_\mathit{1}X}\) because of the \(\alpha^{-1}\) factors in ?? .

Example 12 (name=Regression \(\mu_{\mathcal{X}}\) — Orthogonal Series Estimator). Let \((\varphi_j)\), \(j=1,2,\ldots\), be an orthonormal basis in \(L_2(P_{V_\mathit{1}X})\), that is, \(P_{V_\mathit{1}X}(\varphi_j\varphi_k)=\mathbb{1}_{k=j}\) for all \(j,k\geq1\) and \(\mu = \sum_{j=1}^\infty c_j \varphi_j\) for projection coefficients \(c_j\mathrel{\vcenter{:}}= P_{V_\mathit{1}X}(\mu\varphi_j)\) for all \(\mu\in L_2(P_{V_\mathit{1}X})\). If \((\varphi_j)\) were known, one could estimate \(\mu_{\mathcal{X}}\) in the nonprivate setting by \(\sum_{j=1}^J \hat{c}_j \varphi_j\) with \(\hat{c}_j \mathrel{\vcenter{:}}=\frac{1}{n}\sum_{i\in[n]} m(V_i',X_i')\varphi_j(V_{\mathit{1}i}', X_i')\) and a positive integer \(J=J_n\) tending to infinity. However, we cannot construct a basis \((\varphi_j)\) in \(L_2(P_{V_\mathit{1}X})\), because \(P_{V_\mathit{1}X}\) itself is unknown. A remedy to this is to assume that \(\mu_{\mathcal{X}}\) is in the smaller space \(L_2(\nu_{V_{\mathit{1}}X})\subseteq L_2(P_{V_\mathit{1}X})\), where \(\nu_{V_{\mathit{1}}X}\) is the dominating measure of \(P_{V_\mathit{1}X}\). For instance, with \((V_\mathit{1},X)\) taking values in a subspace of \(\mathbb{R}^d\) and \(\nu_{V_{\mathit{1}}X}\) the Lebesgue measure, this assumption necessitates that \(|\mu_{\mathcal{X}}(v_\mathit{1},x)|\) decay to zero as \(v_\mathit{1}\) or \(x\) tends to infinity. Under this assumption, we can write \(\mu_{\mathcal{X}}=\sum_{j=1}^\infty a_j \phi_j\) for \(a_j\mathrel{\vcenter{:}}=\nu_{V_{\mathit{1}}X}(\mu\phi_j)\) and an orthonormal basis \((\phi_j)\) in \(L_2(\nu_{V_{\mathit{1}}X})\). Correspondingly, if \({\mathcal{S}}'=((V_i',X_i'))_{i\in[n]}\) were observed, we could estimate \(\mu_{\mathcal{X}}\) by \[\begin{align} {\hat{\mu}_{\mathcal{X}}}(v_\mathit{1},x)\mathrel{\vcenter{:}}=\sum_{j=1}^J \hat{a}_j \phi_j(v_\mathit{1},x), \quad \hat{a}_j\mathrel{\vcenter{:}}=\frac{1}{n}\sum_{i\in[n]} \frac{m(V_i',X_i')\phi_j(V_{\mathit{1}i}', X_i')}{\hat{p}_{V_\mathit{1}X}(V_{\mathit{1}i}',X_i')}, \end{align}\] for \(J=J_n\) tending to infinity with \(n\). We inversely weight with the estimated density \(\hat{p}_{V_\mathit{1}X}\) to correct for having the basis in \(L_2(\nu_{V_{\mathit{1}}X})\) but using the empirical version of \(P_{V_\mathit{1}X}\) to compute the projection coefficients, since \(P_{V_\mathit{1}X}\left(\frac{\mu\phi_j}{p_{V_\mathit{1}X}}\right)= \nu_{V_{\mathit{1}}X}(\mu\phi_j)=a_j\). We can rewrite \({\hat{\mu}_{\mathcal{X}}}\) in the form 32 as \[\begin{align} {\hat{\mu}_{\mathcal{X}}}(v_\mathit{1},x)&= \frac{1}{n}\sum_{i\in[n]} w_{n,i}(v_\mathit{1}, x, V_{i}',X_i',\hat{p}_{V_\mathit{1}X}), \\ w_{n,i}(v_\mathit{1}, x, v',x', p_{V_\mathit{1}X})&\mathrel{\vcenter{:}}=\sum_{j=1}^J \frac{m(v',x')\phi_j(v_\mathit{1}',x')}{p_{V_\mathit{1}X}(v_\mathit{1}',x')}\phi_j(v_\mathit{1},x), \end{align}\] and modify it according to 33 for private estimation from \(\bar{\mathcal{S}}'=((V_i',Z_i'))_{i\in[n]}\) as \[\begin{align} {\check{\mu}_{\mathcal{X}}}(v_\mathit{1},x)&\mathrel{\vcenter{:}}=\frac{1}{n}\sum_{i\in[n]} \bar w_{n,i}(v_\mathit{1}, x, V_{i}',Z_i',\check{p}_{V_\mathit{1}X}), \\ \bar w_{n,i}(v_\mathit{1}, x, v',z',\check{p}_{V_\mathit{1}X}) &= \left(Q_\mathcal{X}^{-1}(v',x')\mapsto \sum_{j=1}^J \frac{m(v',x')\phi_j(v_\mathit{1}',x')}{\check{p}_{V_\mathit{1}X}(v_\mathit{1}',x')}\phi_j(v_\mathit{1},x) \right)(v',z') \\ &=\sum_{j=1}^J \left(Q_\mathcal{X}^{-1}(v',x')\mapsto \frac{m(v',x')\phi_j(v_\mathit{1}',x')}{\check{p}_{V_\mathit{1}X}(v_\mathit{1}',x')}\right)(v',z')\phi_j(v_\mathit{1},x), \\ &\eqqcolon\sum_{j=1}^J (Q_\mathcal{X}^{-1}\check{\varrho}_j)(v',z')\phi_j(v_\mathit{1},x), \end{align}\] by the linearity of \(Q_\mathcal{X}^{-1}\), where \(\check p_{V_\mathit{1}X}\) is an arbitrary estimator of \(p_{V_\mathit{1}X}\) computed from \(\bar{\mathcal{S}}'\). Next, we study \({\check{\mu}_{\mathcal{X}}}\) in relation to 3, assuming that \(Q\in\mathcal{Q}_{\delta}\).

The term \(T_{n}\) in ?? is \[\begin{align} T_n(v_\mathit{1},x)&=\frac{1}{n}\sum_{i\in[n]} \bar w_{n,i}(v_\mathit{1}, x, V_{i}',Z_i',\check{p}_{V_\mathit{1}X})-\bar w_{n,i}(v_\mathit{1}, x, V_{i}',Z_i',p_{V_\mathit{1}X}) \nonumber \\ &=\sum_{j=1}^J \left(\frac{1}{n}\sum_{i\in[n]}(Q_\mathcal{X}^{-1}(\check{\varrho}_j-\varrho_j))(V_i',Z_i')\right) \phi_j(v_\mathit{1},x)\eqqcolon\sum_{j=1}^J \Delta_{n,j}\phi_j(v_\mathit{1},x), \end{align}\] where \(\varrho_j(v',x')\mathrel{\vcenter{:}}=\frac{m(v',x')\phi_j(v_\mathit{1}',x')}{p_{V_\mathit{1}X}(v_\mathit{1}',x')}\). The \(\Delta_{n,j}\) do not depend on \((v_\mathit{1},x)\). Because \(T_n^2\) is nonnegative, \[\begin{align} P_{VX}T_n^2\leq \left\lVert{p_{V_\mathit{1}X}}\right\rVert_{\infty}\sum_{j=1}^J\sum_{k=1}^J\Delta_{n,j}\Delta_{k,n}P_{VX}\left(\frac{\phi_j\phi_k}{p_{V_\mathit{1}X}}\right) = \left\lVert{p_{V_\mathit{1}X}}\right\rVert_{\infty}\sum_{j=1}^J \Delta_{n,j}^2 \end{align}\] as \(P_{VX}\left(\frac{\phi_j\phi_k}{p_{V_\mathit{1}X}}\right)=\nu_{V_{\mathit{1}}X}(\phi_j\phi_k)=\mathbb{1}_{j=k}\) by the orthonormality of the \((\phi_j)\). Assume that \[\begin{align} \left\lVert{\frac{m}{p_{V_\mathit{1}X}}}\right\rVert_{\infty}\left\lVert{p_{V_\mathit{1}X}}\right\rVert_{\infty}<\infty, \quad \sum_{j=1}^J\left\lVert{\phi_j}\right\rVert_{\infty}^2=O\left(J\right). \end{align}\] As \(\left\lVert{Q_\mathcal{X}^{-1}h}\right\rVert_{\infty}\leq \frac{2}{\alpha} \left\lVert{h}\right\rVert_{\infty}\), \(P_{VX}T_n^2=O\left(\frac{J}{\alpha^2}\left\lVert{\frac{\check{p}_{V_\mathit{1}X} - p_{V_\mathit{1}X}}{\check{p}_{V_\mathit{1}X}}}\right\rVert_{\infty}^2\right)\).

To bound the variance term in ?? , we use ?? under \(Q\in\mathcal{Q}_{\delta}\): \[\begin{align} \mathop{\mathrm{\mathbb{V}}}\left[\bar{w}_{n,i}(v_\mathit{1}, x,V,Z,p_{V_\mathit{1}X})\right]\leq \mathbb{E}w_{n,i}^2(v_\mathit{1}, x,V,Z,p_{V_\mathit{1}X}) \\ =\mathbb{E}\left[ \frac{m^2(V,Z)p_{V_\mathit{1}Z}(V_\mathit{1}, Z)p_{V_\mathit{1}X}(v_\mathit{1}, x)}{p_{V_\mathit{1}X}^2(V_\mathit{1}, Z)}\frac{\left(\sum_{j=1}^J \phi_j(V_\mathit{1}, Z)\phi_j(v_\mathit{1}, x)\right)^2}{p_{V_\mathit{1}Z}(V_\mathit{1}, Z)p_{V_\mathit{1}X}(v_\mathit{1}, x)}\right] \\ \leq \left\lVert{\frac{m^2p_{V_\mathit{1}Z}}{p_{V_\mathit{1}X}^2}}\right\rVert_{\infty}\left\lVert{p_{V_\mathit{1}X}}\right\rVert_{\infty}\sum_{j=1}^J\sum_{k=1}^J \mathbb{E}\left[\frac{\phi_j\phi_k}{p_{V_\mathit{1}Z}}(V,Z)\right]\frac{\phi_j\phi_k}{p_{V_\mathit{1}X}}(v_\mathit{1}, x). \end{align}\] Here the expectation equals \(\mathbb{1}_{j=k}\), and likewise \(P_{V_\mathit{1}X}\left(\frac{\phi_j\phi_k}{p_{V_\mathit{1}X}}\right)=\mathbb{1}_{j=k}\). Hence, if \(\left\lVert{\frac{m^2p_{V_\mathit{1}Z}}{p_{V_\mathit{1}X}^2}}\right\rVert_{\infty}\left\lVert{p_{V_\mathit{1}X}}\right\rVert_{\infty}<\infty\), then \(\int \mathop{\mathrm{\mathbb{V}}}\left[\bar{w}_{n,i}(v_\mathit{1}, x,V,Z,p_{V_\mathit{1}X})\right] \,\mathrm{d}P_{V_\mathit{1}X}(v_\mathit{1}, x)\lesssim J\).

To bound the bias in ?? , first write by the tower property of expectations, expanding \(\mu_{\mathcal{X}}=\sum_{j=1}^\infty a_j\phi_j\) for \(a_j=P_{VX}(\mu_{\mathcal{X}}\phi_j)\), \[\begin{align} \frac{1}{n}\sum_{i\in[n]}P_{VX}[w_{n,i}(v_\mathit{1},x,V,X,\vartheta)] &= P_{VX}\left[\sum_{j=1}^J \frac{m(V,X)\phi_j(V_\mathit{1},X)}{p_{V_\mathit{1}X}(V_\mathit{1}X)}\phi_j(v_\mathit{1},x)\right] \\ &=P_{VX}\left[\left(\sum_{j=1}^J \frac{\phi_j(V_\mathit{1},X)}{p_{V_\mathit{1}X}(V_\mathit{1}X)}\phi_j(v_\mathit{1},x)\right)\sum_{j=1}^\infty a_j \phi_j(V_\mathit{1},X)\right] \\ &=\sum_{j=1}^J a_j \phi_j(v_\mathit{1},x), \end{align}\] where we used the orthonormality of \((\phi_j)\). Hence, \[\left(\sum_{j=1}^J a_j \phi_j(v_\mathit{1},x)-\mu_{\mathcal{X}}(v_\mathit{1},x)\right)^2=\left(\sum_{j=J+1}^\infty a_j \phi_j(v_\mathit{1},x)\right)^2.\] A nonnegative function, its integral is bounded by \[\begin{align} P_{VX}\left(\sum_{j=J+1}^\infty a_j \phi_j(V_\mathit{1},X)\right)^2\leq \left\lVert{p_{V_\mathit{1}X}}\right\rVert_{\infty}\nu_{V_{\mathit{1}}X}\left(\sum_{j=J+1}^\infty a_j \phi_j(V_\mathit{1},X)\right)^2 \leq \left\lVert{p_{V_\mathit{1}X}}\right\rVert_{\infty}\sum_{j=J+1}^\infty a_j^2, \end{align}\] again by the orthonormality of \((\phi_j)\). Thus, the bias in ?? is bounded by \(O\left(\sum_{j=J+1}^\infty a_j^2\right)\).

Conclude that \[\begin{align} P_{VX}({\check{\mu}_{\mathcal{X}}}-\mu_{\mathcal{X}})^2=O\left(\frac{J}{\alpha^2}\left\lVert{\frac{\check{p}_{V_\mathit{1}X} - p_{V_\mathit{1}X}}{\check{p}_{V_\mathit{1}X}}}\right\rVert_{\infty}^2\right) + O_{{P_{VZ}}}\left(\frac{J}{\alpha^2 n}\right) + O\left(\sum_{j=J+1}^\infty a_j^2\right), \label{priv:eq:ortseries95muxpriv95bound} \end{align}\qquad{(25)}\] where the first term dominates the second one. Apart from the \(\alpha^{-2}\) privacy factor in the second term, the last two term correspond to the usual bound for orthogonal series estimates under equidistant or uniformly distributed design \((V_\mathit{1},X)\). The first term derives from of the weighting correction to accommodate the unknown density \(p_{V_\mathit{1}X}\). This first term usually results in a suboptimal rate. Indeed, suppose that \(\mu_{\mathcal{X}}\) only depends on \(X\) and that \(X_i=\frac{i}{n}\). If \(\mu_{\mathcal{X}}\) belongs to a Sobolev smoothness class \(S(\beta,L)\) for an \(L>0\) and integer \(\beta>0\), so that the \((\beta-1)\)-th derivative of \(\mu_{\mathcal{X}}\) is absolutely continuous and its integrated squared \(\beta\)-th derivative is bounded by \(L\), then we expect \(P_{VX}({\hat{\mu}_{\mathcal{X}}}-\mu_{\mathcal{X}})^2=O\left(\frac{J}{n}\right)+O\left(J^{-2\beta}\right)\) for a trigonometric basis (see [51]). If \(\beta\) is known, a suitable choice of \(J\sim n^\frac{1}{2\beta+1}\) gives a rate \(n^\frac{-2\beta}{2\beta+1}\). In contrast, if \(\left\lVert{\hat{p}_{X}- p_{X}}\right\rVert_{\infty}=\left(\frac{\log n}{n}\right)^{a}\) for some \(0<a\leq\frac{1}{2}\) — a common case for nonparametric density estimation under regularity conditions — and \(\left\lVert{\frac{1}{\hat{p}_{X}}}\right\rVert_{\infty}<\infty\), then the choice of \(J\) balancing the terms in ?? in the nonprivate setting is \(J\sim \left[\left(\frac{\log n}{n}\right)^{2a}+\frac{1}{n}\right]^\frac{-1}{2\beta+1}\), yielding a rate of \(\left[\left(\frac{\log n}{n}\right)^{2a}+\frac{1}{n}\right]^\frac{2\beta}{2\beta+1}\) for \({\hat{\mu}_{\mathcal{X}}}\). This is of the order \(\left(\frac{\log n}{n}\right)^{2a\frac{2\beta}{2\beta+1}}\), so a loss of \(n^{2a}\) is incurred.

To prevent such a loss, [52] constructs an orthonormal basis in \(L_2(\mathbb{P}_n')\) for the empirical distribution \(\mathbb{P}_n'\) of \(((V_i',X_i'))_{i\in[n]}\). Such a basis is constructed using the whole sample \(((V_{\mathit{1}i}',X_i'))_{i\in[n]}\) and for that reason it is unclear how it would lend itself for transformation to the private setting and analysis by 3.

Example 13 (name=Riesz Representer \(r\) — Orthogonal Series Estimator). The Riesz representer \(r=r_{\gamma_{\mathcal{V}}(c)}\) can be estimated by orthogonal series similarly to the regression in 12. Assume that \(r\in L_2(\nu_{V_{\mathit{1}}X})\subseteq L_2(P_{V_\mathit{1}X})\), and let \((\phi_j)\), \(j=1,2,\ldots\), be an orthonormal basis in \(L_2(\nu_{V_{\mathit{1}}X})\). Then we can write \(r=\sum_{j=1}^\infty a_j \phi_j\), where the projection coefficients \(a_j\mathrel{\vcenter{:}}=\nu_{V_{\mathit{1}}X}(r_{\gamma_{\mathcal{V}}(c)}\phi_j)\) depend on \(\gamma_{\mathcal{V}}(c)\). Assume that \(\phi_j/p_{V_\mathit{1}X}\in L_2(P_{V_\mathit{1}X})\) for all \(j\geq 1\). Then, using the Riesz property “backwards,” we have \[a_j=\nu_{V_{\mathit{1}}X}(r_{\gamma_{\mathcal{V}}(c)}\phi_j)=P_{V_\mathit{1}X}\left(r_{\gamma_{\mathcal{V}}(c)}\frac{\phi_j}{p_{V_\mathit{1}X}}\right)=P_{VX}\left(f\left(V,X,\frac{\phi_j}{p_{V_\mathit{1}X}},\gamma_{\mathcal{V}}(c)\right)\right).\] This is useful, because the form of \(r_{\gamma_{\mathcal{V}}(c)}\) may be unknown, so constructing estimates \(\hat{a}_j\) from the empirical version of \(P_{VX}\left(r_{\gamma_{\mathcal{V}}(c)}\frac{\phi_j}{p_{V_\mathit{1}X}}\right)\) may not be feasible; but \(f\) is known, so estimates can be derived from the rightmost side of the display. Indeed, we can estimate \(r\) in the nonprivate setting in the form 32 as \[\begin{align} \hat{r}(v_\mathit{1},x)&\mathrel{\vcenter{:}}=\frac{1}{n}\sum_{i\in[n]} w_{n,i}(v_\mathit{1}, x, V_{i}',X_i',\hat{p}_{V_\mathit{1}X}, \hat{\gamma}_{\mathcal{V}}(c)), \\ w_{n,i}(v_\mathit{1}, x, v',x', p_{V_\mathit{1}X}, \gamma)&\mathrel{\vcenter{:}}=\sum_{j=1}^J f\left(v',x', \frac{\phi_j}{p_{V_\mathit{1}X}},\gamma \right)\phi_j(v_\mathit{1},x), \end{align}\] for a positive integer \(J=J_n\) tending to infinity with \(n\); some estimator \(\hat{\gamma}_{\mathcal{V}}(c)\) of \(\gamma_{\mathcal{V}}(c)\), such as 66 , and \(\hat{p}_{V_\mathit{1}X}\), an estimator of \(p_{V_\mathit{1}X}\), both computed from \(\mathcal{S}'=((V_i',X_i'))_{i\in[n]}\).

Transferring \(\hat{r}\) to the private setting according to 33 , we obtain \[\begin{align} \check{r}(v_\mathit{1},x)&\mathrel{\vcenter{:}}=\frac{1}{n}\sum_{i\in[n]} \bar w_{n,i}(v_\mathit{1}, x, V_{i}',Z_i',\check{p}_{V_\mathit{1}X}, \check{\gamma}_{\mathcal{V}}(c)), \\ \bar w_{n,i}(v_\mathit{1}, x, v',z', p_{V_\mathit{1}X}, \gamma) &= \sum_{j=1}^J \bar{f}\left(v',z',\frac{\phi_j}{ p_{V_\mathit{1}X}},\gamma\right)\phi_j(v_\mathit{1},x), \end{align}\] where \(\bar{f}(\cdot,\mu,\gamma)=(Q_\mathcal{X}^{-1}(v,x)\mapsto f(v,x, \mu, \gamma))(\cdot)\), and \(\check{\gamma}_{\mathcal{V}}(c)\) is an estimator of \(\gamma_{\mathcal{V}}(c)\), such as 24 , and \(\check{p}_{V_\mathit{1}X}\) of \(p_{V_\mathit{1}X}\); all computed from \(\bar{\mathcal{S}}'=((V_i',Z_i'))_{i\in[n]}\). Let us turn to the implications of 3, which are similar to those in 12, except now we have two secondary-nuisance parameters \(\vartheta= (\gamma_{\mathcal{V}}(c),p_{V_\mathit{1}X})\). We assume that \(Q\in\mathcal{Q}_{\delta}\) in 15 .

The term \(T_n\) in ?? is \[\begin{align} T_{n}(v_\mathit{1}, x)=&\, \left\{\frac{1}{n}\sum_{i\in[n]} \bar w_{n,i}(v_\mathit{1}, x, V_{i}',Z_i',\check{p}_{V_\mathit{1}X}, \check{\gamma}_{\mathcal{V}}(c)) - \frac{1}{n}\sum_{i\in[n]} \bar w_{n,i}(v_\mathit{1}, x, V_{i}',Z_i',\check{p}_{V_\mathit{1}X}, \gamma_{\mathcal{V}}(c)) \right\} \\ &+ \left\{\frac{1}{n}\sum_{i\in[n]} \bar w_{n,i}(v_\mathit{1}, x, V_{i}',Z_i',\check{p}_{V_\mathit{1}X}, \gamma_{\mathcal{V}}(c)) - \frac{1}{n}\sum_{i\in[n]} \bar w_{n,i}(v_\mathit{1}, x, V_{i}',Z_i', p_{V_\mathit{1}X}, \gamma_{\mathcal{V}}(c))\right\} \\ \eqqcolon&\, T_{n1}(v_\mathit{1}, x) + T_{n2}(v_\mathit{1}, x). \end{align}\] By the mean-value theorem, \(T_{n1}(v_\mathit{1}, x)=(\check{\gamma}_{\mathcal{V}}(c)-\gamma_{\mathcal{V}}(c))\sum_{j=1}^J D_{n,j} \phi_j(v_\mathit{1}, x)\), where the \(D_{n,j}\mathrel{\vcenter{:}}=\frac{1}{n}\sum_{i\in[n]} \partial_\gamma \bar{f} \left(V_i',Z_i', \frac{\phi_j}{\check{p}_{V_\mathit{1}}X} ,\tilde{\gamma}_{\mathcal{V}}(c) \right)\) do not depend on \((v_\mathit{1}, x)\), and \(\tilde{\gamma}_{\mathcal{V}}(c)\) is some value between \(\check{\gamma}_{\mathcal{V}}(c)\) and \(\gamma_{\mathcal{V}}(c)\). Because \(V_\mathit{2}\) is discretely distributed, \(\check{\gamma}_{\mathcal{V}}(c)-\gamma_{\mathcal{V}}(c)=O_{{P_{VZ}}}\left(n^{-1/2}\right)\) for any reasonable estimator such as 24 . Then, along arguments in 12, \[P_{VX}T_{n1}^2 \leq O_{{P_{VZ}}}\left(1/n\right) \left\lVert{p_{V_\mathit{1}X}}\right\rVert_{\infty} \sum_{j=1}^J D_{n,j}^2\] by the orthonormality of \((\phi_j)\). Assume \(\sum_{j=1}^J D_{n,j}^2=O_{{P_{VZ}}}\left(\frac{J}{\alpha^2}\right)\), which is reasonable for \(\partial_\gamma f\) bounded in all its arguments and privacy mechanism 15 . Then, by arguments in 12, \(P_{VX}T_{n1}^2=O_{{P_{VZ}}}\left(\frac{J}{\alpha^2n}\right)\) provided \(\left\lVert{p_{V_\mathit{1}X}}\right\rVert_{\infty}<\infty\).

By the linearity of \(f\) and thus \(\bar{f}\) in the regression argument, \(T_{2n}(v_\mathit{1}, x)=\sum_{j=1}^J \Delta_{n,j} \phi_j(v_\mathit{1}, x)\) for \(\Delta_{n,j}\mathrel{\vcenter{:}}=\frac{1}{n}\sum_{i\in[n]} \bar{f}\left(V_i',Z_i',\phi_j\cdot\left(\frac{1}{ \check{p}_{V_\mathit{1}X}}-\frac{1}{p_{V_\mathit{1}X}}\right),\gamma_{\mathcal{V}}(c)\right)\). Assuming \(\left\lVert{p_{V_\mathit{1}X}}\right\rVert_{\infty}<\infty\) and \[\sum_{j=1}^J \Delta_{n,j}^2=O_{{P_{VZ}}}\left(\frac{J}{\alpha^2}\left\lVert{\frac{\check{p}_{V_\mathit{1}X} - p_{V_\mathit{1}X}}{\check{p}_{V_\mathit{1}X}}}\right\rVert_{\infty}^2\right),\] we obtain \[P_{VX}T_n^2=O_{{P_{VZ}}}\left(\frac{J}{\alpha^2n}\right)+O_{{P_{VZ}}}\left(\frac{J}{\alpha^2}\left\lVert{\check{p}_{V_\mathit{1}X}-p_{V_\mathit{1}X}}\right\rVert_{\infty}^2\right)=O_{{P_{VZ}}}\left(\frac{J}{\alpha^2}\left\lVert{\check{p}_{V_\mathit{1}X}-p_{V_\mathit{1}X}}\right\rVert_{\infty}^2\right).\]

We bound the variance term in ?? by ?? under \(Q\in\mathcal{Q}_{\delta}\). Assume that \(\left\lVert{p_{V_\mathit{1}X}}\right\rVert_{\infty}<\infty\) and \[\begin{align} \sum_{j=1}^J \mathbb{E}f^2\left(V,Z,\frac{\phi_j}{p_{V_\mathit{1}X}},\gamma_{\mathcal{V}}(c)\right) = O\left(J\right). \end{align}\] Then the arguments in 12 give \(\int \mathop{\mathrm{\mathbb{V}}}\left[\bar{w}_{n,i}(v_\mathit{1}, x,V,Z,p_{V_\mathit{1}X},\gamma_{\mathcal{V}}(c))\right] \,\mathrm{d}P_{V_\mathit{1}X}(v_\mathit{1}, x)\lesssim J\).

The bias in ?? can be bounded as in 12, yielding the bound \(O\left(\sum_{j=J+1}^\infty a_j^2\right)\) for \(a_j= \nu_{V_{\mathit{1}}X}(r_{\gamma_{\mathcal{V}}(c)}\phi_j)\).

Conclude that \[\begin{align} P_{VZ}(\check{r}-r)^2=O_{{P_{VZ}}}\left(\frac{J}{\alpha^2}\left\lVert{\frac{\check{p}_{V_\mathit{1}X} - p_{V_\mathit{1}X}}{\check{p}_{V_\mathit{1}X}}}\right\rVert_{\infty}^2\right)+ O_{{P_{VZ}}}\left(\frac{J}{\alpha^2n}\right) + O\left(\sum_{j=J+1}^\infty a_j^2\right), \end{align}\] where the first term dominates the second one. See 12 for a discussion of these terms; in particular, on how the estimated secondary-nuisance \(p_{V_\mathit{1}X}\) of the first term leads to a rate slower than what only the last two terms would imply.

12 Nonprivate Estimation↩︎

Our aim is private inference, but given that our parameter class is novel, we also present results on nonprivate inference in 12.1, proven in 12.2.

12.1 Results↩︎

In this section, we estimate \(\chi(P_{VX})\) of 3 in the nonprivate setting, from random samples from \((V,X)\sim P_{VX}\) in the nonparametric model 1 . For estimation, we further require the approximability conditions \[\begin{align} \label{priv:eq:sq95continu}\begin{aligned} \mathbb{E}\left[\{f(V,X,\mu,\gamma)-f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c))\}^2\right]&\to 0 \\ \mathbb{E}\left[\{\partial_\gamma f(V,X,\mu,\gamma)-\partial_\gamma f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c))\}^2\right]&\to 0 \end{aligned} \end{align}\tag{56}\] as \(\rho((\mu,\gamma),(\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c)))\to0\). Also note that the linearity 5 implies the \(P_{VX}\)-a.s.linearity of \(\mu\mapsto \frac{\partial^j f}{\partial \gamma^j}(V,X,\mu,\bar{\gamma})\) for any integer \(j\geq 1\) for which the derivative exists. This follows directly from the definition of the derivative.

Recall the one-step estimator \[\begin{align} \hat{\chi}_n= \chi(\hat{P}_{VX})+\mathbb{P}_n\hat{\tilde{\chi}} = \chi(\hat{P}_{VX})+\frac{1}{n}\sum_{i\in[n]}\hat{\tilde{\chi}}(V_i,X_i). \end{align}\] from 8 . Here, the estimator of the influence function ?? is \[\begin{align} \label{priv:eq:eif95chi95hat} \begin{aligned} \hat{\tilde{\chi}}(v,x)\mathrel{\vcenter{:}}=&\, \hat{r}(v_{\mathit{1}},x)(m(v,x)-{\hat{\mu}_{\mathcal{X}}}(v_{\mathit{1}},x))+\frac{\mathbb{1}_{v_{\mathit{2}}=c}}{\hat{p}_{V_{\mathit{2}}}(c)}(g(v,x)-\hat{\gamma}_{\mathcal{V}}(c))\hat{e}\\ &+ f(v,x, {\hat{\mu}_{\mathcal{X}}}, \hat{\gamma}_{\mathcal{V}}(c))-\chi(\hat{P}_{VX}), \end{aligned} \end{align}\tag{57}\] where \(\hat{r},{\hat{\mu}_{\mathcal{X}}}\), taking values in \(L_2(P_{V_{\mathit{1}}X})\), are some estimators of \(r,\mu_{\mathcal{X}}\), respectively; \(\hat{p}_{V_{\mathit{2}}}(c)\), taking values in \(\mathbb{R}\), is some estimator of \(p_{V_{\mathit{2}}}(c)\); and \(\hat{e}\), taking values in \(\mathbb{R}\), is some estimator of \[\begin{align} e\mathrel{\vcenter{:}}=\mathbb{E}\partial_\gamma f(V,X, \mu_{\mathcal{X}}, \gamma_{\mathcal{V}}(c)).\label{priv:eq:e95def} \end{align}\tag{58}\] Note that we use the same initial estimator \(\chi(\hat{P}_{VX})\) in \(\hat{\tilde{\chi}}\), which we may set to \[\begin{align} \chi(\hat{P}_{VX})=\mathbb{P}_nf(V,X, {\hat{\mu}_{\mathcal{X}}}, \hat{\gamma}_{\mathcal{V}}(c))=\frac{1}{n}\sum_{i\in[n]} f(V_i,X_i, {\hat{\mu}_{\mathcal{X}}}, \hat{\gamma}_{\mathcal{V}}(c)) \label{priv:eq:plugin95chihat}. \end{align}\tag{59}\] Therefore, \[\begin{align} \hat{\chi}_n =&\, \chi(\hat{P}_{VX})+\mathbb{P}_n\hat{\tilde{\chi}} \\ =&\, \frac{1}{n}\sum_{i\in[n]}\left\{ \hat{r}(v_{\mathit{1}i},X_i)(m(V_i,X_i)-{\hat{\mu}_{\mathcal{X}}}(v_{\mathit{1}i},X_i))+\frac{\mathbb{1}_{V_{\mathit{2}i}=c}}{\hat{p}_{V_{\mathit{2}}}(c)}(g(V_i,X_i)-\hat{\gamma}_{\mathcal{V}}(c))\hat{e}\right. \\ &+\left. \vphantom{\frac{\mathbb{1}_{V_{\mathit{2}i}=c}}{\hat{p}_{V_{\mathit{2}}}(c)}} f(V_i,X_i, {\hat{\mu}_{\mathcal{X}}}, \hat{\gamma}_{\mathcal{V}}(c)) \right\}. \end{align}\]

We assume that the estimators \[\begin{align} \label{priv:eq:nuisance} \begin{aligned} \hat{\eta}&\mathrel{\vcenter{:}}=(\hat{r},{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c),\hat{p}_{V_{\mathit{2}}}(c), \hat{e})\in L_2(P_{V_{\mathit{1}}X})\times L_2(P_{V_{\mathit{1}}X}) \times\Gamma\times\mathbb{R}\times\mathbb{R}\,\text{ of } \\ \eta&\mathrel{\vcenter{:}}=(r,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c), p_{V_{\mathit{2}}}(c), e) \end{aligned} \end{align}\tag{60}\] are computed from random samples from \(P_{VX}\) which are independentof \[\mathcal{S}\mathrel{\vcenter{:}}=\mathcal{(}(V_i,X_i))_{i\in[n]}.\] Specifically, we assume that there are two more random samples \[\begin{align} \mathcal{S}'\mathrel{\vcenter{:}}=((V_i',X_i'))_{i\in[n]}\,\text{ and }\, \mathcal{S}''\mathrel{\vcenter{:}}=((V_i'',X_i''))_{i\in[n]} \end{align}\] from \(P_{VX}\), with \(\mathcal{S},\mathcal{S}',\mathcal{S}''\) mutually independent, where \(\mathcal{S'}\) is used for the estimation of \((r,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c), p_{V_{\mathit{2}}}(c))\), and \(\mathcal{S}',\mathcal{S}''\) are used for the estimation of \(e\) in 58 as \[\begin{align} \hat{e}\mathrel{\vcenter{:}}=\mathbb{P}_n'' \partial_\gamma f(V,X,{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c))\mathrel{\vcenter{:}}=\frac{1}{n}\sum_{i\in[n]} \partial_\gamma f(V_i'',X_i'',{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c)). \label{priv:eq:expderivhat} \end{align}\tag{61}\] With this estimation strategy, we can establish the consistency of \(\hat{e}\), and hence of \(\hat{\tilde{\chi}}\) and \(\hat{\chi}_n\) in turn, without additional regularity conditions. Decompose \[\begin{align} \sqrt{n}(\hat{\chi}_n - \chi(P_{VX})) &= \sqrt{n}\mathbb{P}_n\tilde{\chi}+\sqrt{n}(\mathbb{P}_n-P_{VX})(\hat{\tilde{\chi}}-\tilde{\chi})+\sqrt{n}R_n, \tag{62} \\ R_n&\mathrel{\vcenter{:}}=\chi(\hat{P}_{VX})-\chi(P_{VX})+P_{VX} \hat{\tilde{\chi}}, \tag{63} \end{align}\] where we used that \(P_{VX}\tilde{\chi}=0\) by \(\tilde{\chi}\) being the influence function. The term \(\sqrt{n}\mathbb{P}_n\tilde{\chi}\overset{P_{VX}}{\rightsquigarrow}\mathcal{N}(0,P_{VX}\tilde{\chi}^2)\) by the standard central limit theorem. The second term in 62 is called the empirical process term and is vanishing as \(o_{P_{VX}}\left(1\right)\) under consistent estimators \(\hat{\eta}\) and stochastic boundedness conditions.

Assumption 4 (Consistent Estimators). It holds that \[\begin{align} \left\lVert{\hat{r}-r}\right\rVert_{L_2(P_{VX})}&=o_{P_{VX}}\left(1\right) \label{priv:eq:consistency95riesz95l2}, \\ \hat{\gamma}_{\mathcal{V}}(c)-\gamma_{\mathcal{V}}(c)&=o_{P_{VX}}\left(1\right), \label{priv:eq:consistency95gammaVhat} \\ \hat{p}_{V_2}(c)-p_{V_2}(c)&=o_{P_{VX}}\left(1\right). \label{priv:eq:consistency95phat} \end{align}\] {#eq: sublabel=eq:priv:eq:consistency95riesz95l2,eq:priv:eq:consistency95gammaVhat,eq:priv:eq:consistency95phat} Further, it either holds that \[\begin{align} \left\lVert{m-\mu_{\mathcal{X}}}\right\rVert_{\infty}&=O\left(1\right), \label{priv:eq:consistency95bigobound} \\ \left\lVert{\mu_{\mathcal{X}}-{\hat{\mu}_{\mathcal{X}}}}\right\rVert_{\infty}&=o_{P_{VX}}\left(1\right), \label{priv:eq:consistency95mu95supn} \end{align}\] {#eq: sublabel=eq:priv:eq:consistency95bigobound,eq:priv:eq:consistency95mu95supn} or that \[\begin{align} \left\lVert{m-{\hat{\mu}_{\mathcal{X}}}}\right\rVert_{\infty}&=O_{{P_{VX}}}\left(1\right), \label{priv:eq:consistency95bigobound95hat} \\ \left\lVert{\mu_{\mathcal{X}}-{\hat{\mu}_{\mathcal{X}}}}\right\rVert_{L_2(P_{VX})}&=o_{P_{VX}}\left(1\right), \label{priv:eq:consistency95mu95l2} \\ P_{VX}\left(\left\{(V_{\mathit{1}},X)\in \mathfrak{V}_{\mathit{1}}\times\mathfrak{X}: |r(V_{\mathit{1}},X)|>\bar{R} \right\}\right)&=0 \label{priv:eq:consistency95riesz95bound} \end{align}\] {#eq: sublabel=eq:priv:eq:consistency95bigobound95hat,eq:priv:eq:consistency95mu95l2,eq:priv:eq:consistency95riesz95bound} for some constant \(\bar{R}<\infty\), where ?? may be replaced by \(\left\lVert{r}\right\rVert_{\infty}<\infty\).

Lemma 9 (Vanishing Empirical Process Term). Assume that 4 holds for \(\hat{\eta}\) in 60 . Then \(\hat{e}-e=o_{P_{VX}}\left(1\right)\) and \((\mathbb{P}_n-P_{VX})(\hat{\tilde{\chi}}-\tilde{\chi})=o_{P_{VX}}\left(n^{-1/2}\right)\).

Hence, if the nuisance parameters in \(\eta\) are consistently estimated and bound conditions apply, the behaviour of \(\sqrt{n}(\hat{\chi}_n-\chi(P_{VX}))\) is governed by the second-order bias term \(R_n\) in 63 . We show that our class of parameters 3 enjoys a rate-double-robustness property, therefore \(R_n\) exhibits a product-structure of estimation errors.

1 implies that the bias in 63 is \[\begin{align} \label{priv:eq:bias95r} \begin{aligned} R_n =&\;\chi(\hat{P}_{VX})-\chi(P_{VX})+P_{VX}\hat{\tilde{\chi}} \\ =&\;-P_{VX}(r-\hat{r})(\mu_{\mathcal{X}}-{\hat{\mu}_{\mathcal{X}}}) +(\gamma_{\mathcal{V}}(c)-\hat{\gamma}_{\mathcal{V}}(c))\left(\frac{p_{V_{\mathit{2}}}(c)}{\hat{p}_{V_{\mathit{2}}}(c)}\hat{e}-e'\right) \\ &-(\gamma_{\mathcal{V}}(c)-\hat{\gamma}_{\mathcal{V}}(c))^2\frac{P_{VX}\partial_\gamma^2 f(V,X,{\hat{\mu}_{\mathcal{X}}},\tilde{\gamma}_{\mathcal{V}}(c))}{2}, \end{aligned} \end{align}\tag{64}\] for some \(\tilde{\gamma}_{\mathcal{V}}(c)\) between \(\gamma_{\mathcal{V}}(c)\) and \(\hat{\gamma}_{\mathcal{V}}(c)\), and \[\begin{align} e' \mathrel{\vcenter{:}}=&\, P_{VX}\partial_\gamma f(V,X,{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c)). \label{priv:eq:derivhatprime} \end{align}\tag{65}\] Because \(V_{\mathit{2}}\) is distributed on a finite set, any reasonable estimators of \(\gamma_{\mathcal{V}}(c)\), \(p_{V_{\mathit{2}}}(c)\) are \(\sqrt{n}\)-consistent; for instance \[\begin{align} \hat{\gamma}_{\mathcal{V}}(c) \mathrel{\vcenter{:}}=\frac{1}{N_c}\sum_{i\in[n]}\mathbb{1}_{V_{\mathit{2}i}'=c} g(V_i',X_i'), \quad \hat{p}_{V_{\mathit{2}}}(c)\mathrel{\vcenter{:}}= N_c/n, \quad N_c\mathrel{\vcenter{:}}=\sum_{i\in[n]}\mathbb{1}_{V_{\mathit{2}i}'=c} \label{priv:eq:gammaVhat} \end{align}\tag{66}\] satisfy \(\hat{\gamma}_{\mathcal{V}}(c)-\gamma_{\mathcal{V}}(c)=O_{{P_{VX}}}\left(n^{-1/2}\right)\), \(\hat{p}_{V_{\mathit{2}}}(c)-p_{V_{\mathit{2}}}(c)=O_{{P_{VX}}}\left(n^{-1/2}\right)\) by the standard central limit theorem. Suppose that \(\hat{e}-e'=o_{P_{VX}}\left(1\right)\) and that \(P_{VX} \partial_\gamma^2 f(V,X,{\hat{\mu}_{\mathcal{X}}},\tilde{\gamma}_{\mathcal{V}}(c))=O_{{P_{VX}}}\left(1\right)\). It follows from 64 that the bias is then \[\begin{align} R_n = -P_{VX}(r-\hat{r})(\mu_{\mathcal{X}}-{\hat{\mu}_{\mathcal{X}}}) + o_{P_{VX}}\left(n^{-1/2}\right). \label{priv:eq:chihat95dr} \end{align}\tag{67}\] Hence, \(R_n\) is ultimately determined by the product of the estimation errors of the Riesz representer \(r\) and the regression function \(\mu_{\mathcal{X}}\). For example, we find that the average treatment effect on the treated is double robust. This is aligned with [4], and is an improvement on [3], who too, establish asymptotic normality, but not efficiency, as this parameter is not natively included in their class.

Our results amounts to the asymptotic efficiency of \(\hat{\chi}_n\) under fast enough estimation rates, boundedness conditions, and the estimation strategy using independent samples.

Assumption 5 (Rates of Estimators). For \(\hat{\eta}\) in 60 , \[\begin{align} P_{VX}(r-\hat{r})(\mu_{\mathcal{X}}-{\hat{\mu}_{\mathcal{X}}})&=o_{P_{VX}}\left(n^{-1/2}\right), \\ \hat{\gamma}_{\mathcal{V}}(c)-\gamma_{\mathcal{V}}(c)&=O_{{P_{VX}}}\left(n^{-1/2}\right), \\ \hat{p}_{V_{\mathit{2}}}(c)-p_{V_{\mathit{2}}}(c)&=O_{{P_{VX}}}\left(n^{-1/2}\right), \\ P_{VX} \partial_\gamma^2 f(V,X,{\hat{\mu}_{\mathcal{X}}},\tilde{\gamma}_{\mathcal{V}}(c))&=O_{{P_{VX}}}\left(1\right). \end{align}\]

Corollary 6 (Asymptotic Efficiency of \(\hat{\chi}_n\)). If [priv:ass:consistent_nuisance,priv:ass:rates_nuisance] hold, then \(\sqrt{n}(\hat{\chi}_n-\chi(P_{VX}))\overset{P_{VX}}{\rightsquigarrow}\mathcal{N}(0, P_{VX}\tilde{\chi}^2)\) as \(n\to\infty\).

Conveniently and expectedly, when \((X,V_{\mathit{1}})\) is discretely distributed, the same limit is achievable without the use of independent samples \(\mathcal{S},\mathcal{S}',\mathcal{S}''\). Indeed, the plug-in estimator \(\chi(\hat{P}_{VX})\) is asymptotically efficient when we compute the estimators \(({\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}})\) and \(\chi(\hat{P}_{VX})\) on the same sample \(\mathcal{S}\), provided some boundedness conditions hold.

Proposition 4 (Asymptotic Efficiency of the Plug-in Estimator for Discrete \((X,V_{\mathit{1}})\)). Suppose that \(\mathfrak{X}\) and \(\mathfrak{V}_{\mathit{1}}\) are finite, and define \[\begin{align} \chi(\hat{P}_{VX})\mathrel{\vcenter{:}}=&\, \frac{1}{n}\sum_{i\in[n]} f(V_i,X_i,{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c)), \\ {\hat{\mu}_{\mathcal{X}}}(v_{\mathit{1}},x)\mathrel{\vcenter{:}}=&\, \frac{1}{N_{v_{\mathit{1}}x}}\sum_{i\in[n]}\mathbb{1}_{V_{\mathit{1}i}=v_{\mathit{1}}, X_i=x}m(V_i,X_i), \quad N_{v_{\mathit{1}}x}\mathrel{\vcenter{:}}=\sum_{i\in[n]}\mathbb{1}_{V_{\mathit{1}i}=v_{\mathit{1}}, X_i=x}, \\ \hat{\gamma}_{\mathcal{V}}(c) \mathrel{\vcenter{:}}=&\, \frac{1}{N_c}\sum_{i\in[n]}\mathbb{1}_{V_{\mathit{2}i}=c} g(V_i,X_i), \quad N_c\mathrel{\vcenter{:}}=\sum_{i\in[n]}\mathbb{1}_{V_{\mathit{2}i}=c}. \end{align}\] Assume that \(\left\lVert{\mu_{\mathcal{X}}}\right\rVert_{\infty}<\infty\) and that there exists an \(\hat{r}:\mathfrak{V}_1\times\mathfrak{X}\to\mathbb{R}\) such that \(\left\lVert{\hat{r}-r}\right\rVert_{\infty}=o_{P_{VX}}\left(1\right)\), where \(r\) is the Riesz representer of \(\mu\mapsto \mathbb{E}f(V,X,\mu,\gamma_{\mathcal{V}}(c))\) \(=\chi(P_{VX})\) as before. Further assume that for every fixed \(\mu\in L_2(P_{V_{\mathit{1}}X})\) and for every \(\tilde{\gamma}_{\mathcal{V}}(c)\) between \(\gamma_{\mathcal{V}}(c)\) and \(\hat{\gamma}_{\mathcal{V}}(c)\), the bound conditions \[\begin{align} (\mathbb{P}_n-P_{VX})\partial_\gamma f(V,X,\mu,\tilde{\gamma}_{\mathcal{V}}(c))&=O_{{P_{VX}}}\left(1\right),\label{priv:eq:emproc95discrete95bound1} \\ (\mathbb{P}_n-P_{VX})\partial_\gamma^2 f(V,X,\mu,\tilde{\gamma}_{\mathcal{V}}(c))&=O_{{P_{VX}}}\left(1\right),\label{priv:eq:emproc95discrete95bound2} \\ P_{VX}\partial_\gamma f(V,X,\mu_{\mathcal{X}},\tilde{\gamma}_{\mathcal{V}}(c))&=O_{{P_{VX}}}\left(1\right), \label{priv:eq:emproc95discrete95bound3} \\ P_{VX}\partial_\gamma^2 f(V,X,\mu_{\mathcal{X}},\tilde{\gamma}_{\mathcal{V}}(c))&=O_{{P_{VX}}}\left(1\right), \label{priv:eq:emproc95discrete95bound4} \\ P_{VX}\partial_\gamma^2 f(V,X,{\hat{\mu}_{\mathcal{X}}},\tilde{\gamma}_{\mathcal{V}}(c))&=O_{{P_{VX}}}\left(1\right) \label{priv:eq:emproc95discrete95bound5} \end{align}\] {#eq: sublabel=eq:priv:eq:emproc95discrete95bound1,eq:priv:eq:emproc95discrete95bound2,eq:priv:eq:emproc95discrete95bound3,eq:priv:eq:emproc95discrete95bound4,eq:priv:eq:emproc95discrete95bound5} hold. Then \(\sqrt{n}(\chi(\hat{P}_{VX})-\chi(P_{VX}))\overset{P_{VX}}{\rightsquigarrow}\mathcal{N}(0, P_{VX}\tilde{\chi}^2)\) as \(n\to\infty\).

Note that for particular forms of \(f\), the conditions of 4 are easier to verify. Namely, if \(f(v,x,\mu,\gamma)\) factors as \[f(v,x,\mu,\gamma)=f_1(v,x,\mu)f_2(\gamma)\] where \(\frac{\partial ^2 f_2}{\partial\gamma^2}\) exists and is continuous and \((v,x)\mapsto f_1(v,x,\mu)\) belongs to \(L_2(P_{VX})\) for any \(\mu\in L_2(P_{V_{\mathit{1}}X})\), then the boundedness conditions ?? ,?? ,?? , ?? are easily met by the standard central limit theorem and the continuous mapping theorem. For instance, this includes the average treatment effect on the treated in 3.

12.2 Proofs↩︎

In this section, we prove the results in 12.1: [priv:lem:chi_eff_empprocess,priv:cor:eff_chihat,priv:prop:asymeff_plugin_discretecovar].

Proof of 9. Throughout, we apply that if a \(P_{VX}\)-random function \(\hat{q} \in L_2(P_{VX})\) is independent of the random sample generating the process \(\mathbb{P}_n\), then \[\begin{align} \label{priv:eq:meansq95empproc} \begin{aligned} \int(\hat{q}(v,x)-q(v,x))^2 \,\mathrm{d}P_{VX}(v,x)=o_{P_{VX}}\left(1\right) \text{ implies } \\ \sqrt{n}(\mathbb{P}_n-P_{VX})(\hat{q} -q)=o_{P_{VX}}\left(1\right); \end{aligned} \end{align}\tag{68}\] which follows from Markov’s inequality and the dominated convergence theorem.

By the definitions ?? and 57 , \[\begin{align} \hat{\tilde{\chi}}(v,x)-\tilde{\chi}(v,x) =&\, T_1(v,x)+T_2(v,x)+T_3(v,x)+T_4, \label{priv:eq:emproc95decomp} \\ T_1(v,x)\mathrel{\vcenter{:}}=&\, \hat{r}(v_{\mathit{1}},x)(m(v,x)-{\hat{\mu}_{\mathcal{X}}}(v_{\mathit{1}},x)) \nonumber \\ &-r(v_{\mathit{1}},x)(m(v,x)-\mu_{\mathcal{X}}(v_{\mathit{1}},x)), \nonumber \\ T_2(v,x)\mathrel{\vcenter{:}}=&\, \frac{\mathbb{1}_{v_{\mathit{2}}=c}}{\hat{p}_{V_{\mathit{2}}}(c)}(g(v,x)-\hat{\gamma}_{\mathcal{V}}(c))\hat{e}-\frac{\mathbb{1}_{v_{\mathit{2}}=c}}{p_{V_{\mathit{2}}}(c)}(g(v,x)-\gamma_{\mathcal{V}}(c))e, \nonumber \\ T_3(v,x)\mathrel{\vcenter{:}}=&\, f(v,x, {\hat{\mu}_{\mathcal{X}}}, \hat{\gamma}_{\mathcal{V}}(c)) -f(v,x, \mu_{\mathcal{X}}, \gamma_{\mathcal{V}}(c)), \nonumber \\ T_4\mathrel{\vcenter{:}}=&\, -\chi(\hat{P}_{VX})+\chi(P_{VX}). \nonumber \end{align}\tag{69}\] As \(T_4\) is constant, not depending on \((v,x)\), \((\mathbb{P}_n-P_{VX})T_4=0\). It remains to show \((\mathbb{P}_n-P_{VX})T_j=o_{P_{VX}}\left(n^{-1/2}\right)\) for \(j=1,2,3\) by the linearity of the process \(\mathbb{P}_n-P_{VX}\).

Term \(T_1.\,\,\) Suppressing the arguments, write \[\begin{align} T_1=\hat{r}(m-{\hat{\mu}_{\mathcal{X}}})-r(m-\mu_{\mathcal{X}}) &= (\hat{r}-r+r)(m-{\hat{\mu}_{\mathcal{X}}})-r(m-\mu_{\mathcal{X}}) \\ &= (\hat{r}-r)(m-{\hat{\mu}_{\mathcal{X}}}) + r(\mu_{\mathcal{X}}-{\hat{\mu}_{\mathcal{X}}}). \end{align}\] By 4, \(\left\lVert{m-{\hat{\mu}_{\mathcal{X}}}}\right\rVert_{\infty}=O_{{P_{VX}}}\left(1\right)\); either directly by ?? , or by ?? and ?? , noting that \(\left\lVert{m-{\hat{\mu}_{\mathcal{X}}}}\right\rVert_{\infty}\leq \left\lVert{m-\mu_{\mathcal{X}}}\right\rVert_{\infty}+\left\lVert{\mu_{\mathcal{X}}-{\hat{\mu}_{\mathcal{X}}}}\right\rVert_{\infty}=O_{{P_{VX}}}\left(1\right)+o_{P_{VX}}\left(1\right)=O_{{P_{VX}}}\left(1\right)\). Then the \(L_2(P_{V_{\mathit{1}}X})\)-convergence ?? of \(r\) implies that \((\mathbb{P}_n-P_{VX})((\hat{r}-r)(m-{\hat{\mu}_{\mathcal{X}}}))=o_{P_{VX}}\left(n^{-1/2}\right)\) by 68 as \[\begin{align} P_{VX}((\hat{r}-r)^2(m-{\hat{\mu}_{\mathcal{X}}})^2) \\ \leq\left\lVert{m-{\hat{\mu}_{\mathcal{X}}}}\right\rVert_{\infty}^2 P_{VX}(\hat{r}-r)^2=O_{{P_{VX}}}\left(1\right)o_{P_{VX}}\left(1\right)=o_{P_{VX}}\left(1\right) \end{align}\] since \(\left\lVert{|q|^2}\right\rVert_{\infty}=\left\lVert{q}\right\rVert_{\infty}^2\).

By 4, either ?? , or ?? and ?? hold. In the former case, \[\begin{align} P_{VX}(r^2(\mu_{\mathcal{X}}-{\hat{\mu}_{\mathcal{X}}})^2)\leq\left\lVert{\mu_{\mathcal{X}}-{\hat{\mu}_{\mathcal{X}}}}\right\rVert_{\infty}^2 P_{VX}r^2=o_{P_{VX}}\left(1\right), \end{align}\] by 68 because \(r\in L_2(P_{V_{\mathit{1}}X})\). In the latter case, letting \[B\mathrel{\vcenter{:}}=\left\{(V_{\mathit{1}},X)\in \mathfrak{V}_{\mathit{1}}\times\mathfrak{X}: |r(V_{\mathit{1}},X)|>\bar{R} \right\}\] with complement \(B^C\), we have \[\begin{align} P_{VX}(r^2(\mu_{\mathcal{X}}-{\hat{\mu}_{\mathcal{X}}})^2) &= \int_{B^C} r^2(\mu_{\mathcal{X}}-{\hat{\mu}_{\mathcal{X}}})^2 \,\mathrm{d}P_{VX}+ \int_{B} r^2(\mu_{\mathcal{X}}-{\hat{\mu}_{\mathcal{X}}})^2 \,\mathrm{d}P_{VX} \\ &\leq \bar{R}^2 \left\lVert{\mu_{\mathcal{X}}-{\hat{\mu}_{\mathcal{X}}}}\right\rVert_{L_2(P_{V_{\mathit{1}X}})}^2+0=o_{P_{VX}}\left(1\right), \end{align}\] since by ?? , \(P_{VX}(B)=0\) and \({\hat{\mu}_{\mathcal{X}}}\) is \(L_2(P_{V_{\mathit{1}X}})\)-convergent by ?? . Thus, \((\mathbb{P}_n-P_{VX})(r(\mu_{\mathcal{X}}-{\hat{\mu}_{\mathcal{X}}}))=o_{P_{VX}}\left(n^{-1/2}\right)\) by 68 . Conclude that \((\mathbb{P}_n-P_{VX})T_1=o_{P_{VX}}\left(n^{-1/2}\right)\).

Term \(T_2.\,\,\) By the mean-value theorem there exists \((\tilde{\gamma}_{\mathcal{V}}(c), \tilde{p}_{V_{\mathit{2}}}(c), \tilde{e})\) between \[\begin{align} (\gamma_{\mathcal{V}}(c), p_{V_{\mathit{2}}}(c), e) \text{ and } (\hat{\gamma}_{\mathcal{V}}(c), \hat{p}_{V_{\mathit{2}}}(c), \hat{e}) \end{align}\] such that \[\begin{align} T_2(v,x)=&\,\frac{\mathbb{1}_{v_{\mathit{2}}=c}}{\hat{p}_{V_{\mathit{2}}}(c)}(g(v,x)-\hat{\gamma}_{\mathcal{V}}(c))\hat{e}-\frac{\mathbb{1}_{v_{\mathit{2}}=c}}{p_{V_{\mathit{2}}}(c)}(g(v,x)-\gamma_{\mathcal{V}}(c))e\\ =&\,-\frac{\mathbb{1}_{v_{\mathit{2}}=c}}{\tilde{p}_{V_{\mathit{2}}}(c)}\tilde{e}(\hat{\gamma}_{\mathcal{V}}(c)-\gamma_{\mathcal{V}}(c)) \\ &-\frac{\mathbb{1}_{v_{\mathit{2}}=c}}{\tilde{p}_{V_{\mathit{2}}}(c)^2}(g(v,x)-\tilde{\gamma}_{\mathcal{V}}(c))\tilde{e}(\hat{p}_{V_{\mathit{2}}}(c)-p_{V_{\mathit{2}}}(c)) \\ &+\frac{\mathbb{1}_{v_{\mathit{2}}=c}}{\tilde{p}_{V_{\mathit{2}}}(c)}(g(v,x)-\tilde{\gamma}_{\mathcal{V}}(c))(\hat{e}-e). \end{align}\] By the standard central limit theorem, \(\sqrt{n}(\mathbb{P}_n-P_{VX})\mathbb{1}_{V_{\mathit{2}}=c}=O_{{P_{VX}}}\left(1\right)\). By the linearity of the process \(\mathbb{P}_n-P_{VX}\), \[\begin{align} \sqrt{n}(\mathbb{P}_n-P_{VX})\big[\mathbb{1}_{V_{\mathit{2}}=c}g(V,X)-\mathbb{1}_{V_{\mathit{2}}=c}\tilde{\gamma}_{\mathcal{V}}(c)\big] \\ = \sqrt{n}(\mathbb{P}_n-P_{VX})\mathbb{1}_{V_{\mathit{2}}=c}g(V,X) - \tilde{\gamma}_{\mathcal{V}}(c)\sqrt{n}(\mathbb{P}_n-P_{VX})\mathbb{1}_{V_{\mathit{2}}=c} \\ =(1-\tilde{\gamma}_{\mathcal{V}}(c))O_{{P_{VX}}}\left(1\right)=O_{{P_{VX}}}\left(1\right) \end{align}\] again by the standard central limit theorem and ?? . Suppose that \(\hat{e}-e=o_{P_{VX}}\left(1\right)\), which we show later. Then by ?? and ?? , \((\mathbb{P}_n-P_{VX})T_2=o_{P_{VX}}\left(n^{-1/2}\right)\).

Term \(T_3.\,\,\) Recall that \(T_3(v,x)=f(v,x, {\hat{\mu}_{\mathcal{X}}}, \hat{\gamma}_{\mathcal{V}}(c))-f(v,x, \mu_{\mathcal{X}}, \gamma_{\mathcal{V}}(c))\). The continuity 56 , together with the consistency of \(\hat{\gamma}_{\mathcal{V}}\) and \({\hat{\mu}_{\mathcal{X}}}\) (?? and ?? or ?? ) imply that \(\int T_3(v,x)^2\,\mathrm{d}P_{VX}(v,x)=o_{P_{VX}}\left(1\right)\) by the continuous mapping theorem. Conclude by 68 that \((\mathbb{P}_n-P_{VX})T_3=o_{P_{VX}}\left(n^{-1/2}\right)\).

Consistency of \(\hat{e}.\,\,\) By the definition of \(e\) and \(\hat{e}\), \[\begin{align} \hat{e}-e=&\, \mathbb{P}_n''\partial_\gamma f(V,X,{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c)) - P_{VX}\partial_\gamma f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c)) \\ =&\, \mathbb{P}_n''\partial_\gamma f(V,X,{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c))-P_{VX}\partial_\gamma f(V,X,{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c)) \\ &+ P_{VX}\partial_\gamma f(V,X,{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c))- P_{VX}\partial_\gamma f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c)) \\ =&\, (\mathbb{P}_n''-P_{VX})\partial_\gamma f(V,X,{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c)) \\ &+ P_{VX}\big[\partial_\gamma f(V,X,{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c))-\partial_\gamma f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c)) \big] \\ =&\, (\mathbb{P}_n''-P_{VX})\partial_\gamma f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c)) \\ &+(\mathbb{P}_n''-P_{VX})\big[\partial_\gamma f(V,X,{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c))-\partial_\gamma f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c))\big] \\ &+ P_{VX}\big[\partial_\gamma f(V,X,{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c))-\partial_\gamma f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c)) \big]. \end{align}\] Here, the first term is \(O_{{P_{VX}}}\left(n^{-1/2}\right)=o_{P_{VX}}\left(1\right)\) by the standard central limit theorem, and the second and third term are \(o_{P_{VX}}\left(1\right)\) by the continuity 56 along the same arguments concerning \(T_3\) above. ◻

Proof of 6. Follows from [priv:lem:chi_eff_empprocess,priv:thm:dr], noting that, for \(e'\) in 65 , \[\hat{e}-e'=(\mathbb{P}_n''-P_{VX})\partial_\gamma f(V,X,{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c))\] is \(o_{P_{VX}}\left(1\right)\) by the consistency proof of \(\hat{e}\) in 9. ◻

Proof of 4. We start with an auxiliary result. Define \[\begin{align} \hat{\tilde{\chi}}(v,x)\mathrel{\vcenter{:}}=&\, \hat{r}(v_{\mathit{1}},x)(m(v,x)-{\hat{\mu}_{\mathcal{X}}}(v_{\mathit{1}},x))+\frac{\mathbb{1}_{v_{\mathit{2}}=c}}{\hat{p}_{V_{\mathit{2}}}(c)}(g(v,x)-\hat{\gamma}_{\mathcal{V}}(c))\hat{e}\nonumber \\ &+ f(v,x, {\hat{\mu}_{\mathcal{X}}}, \hat{\gamma}_{\mathcal{V}}(c))-\chi(\hat{P}_{VX}) \\ \hat{p}_{V_{\mathit{2}}}(c)\mathrel{\vcenter{:}}=&\, N_c/n. \end{align}\] First we show that \[\begin{align} \mathbb{P}_n\hat{\tilde{\chi}} =\mathbb{P}_n\left[\hat{r}(V_{\mathit{1}},X)(m(V,X)-{\hat{\mu}_{\mathcal{X}}}(V_{\mathit{1}},X)\right] \\ + \frac{\hat{e}}{\hat{p}_{V_{\mathit{2}}}(c)}\mathbb{P}_n\left[\mathbb{1}_{V_{\mathit{2}}=c}g(V,X)-\hat{\gamma}_{\mathcal{V}}(c)\mathbb{1}_{V_{\mathit{2}}=c}\right] +\mathbb{P}_n\left[f(V,X,{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c))-\chi(\hat{P}_{VX}))\right] \end{align}\] is zero. We apply an ‘empirical tower property’ to the first term to get \[\begin{align} \mathbb{P}_n\hat{r}(V_{\mathit{1}},X)(m(V,X)-{\hat{\mu}_{\mathcal{X}}}(V_{\mathit{1}},X)) = \frac{1}{n}\sum_{i\in[n]}\hat{r}(V_{\mathit{1}i},X_i)(m(V_i,X_i)-{\hat{\mu}_{\mathcal{X}}}(V_{\mathit{1}i},X_i)) \\ = \frac{1}{n}\sum_{(v_{\mathit{1}}, x)\in \mathfrak{V}_{\mathit{1}}\times\mathfrak{X}}\sum_{i: (V_{\mathit{1}i},X_i)=(v_{\mathit{1}}, x)}\hat{r}(V_{\mathit{1}i},X_i)(m(V_i,X_i)-{\hat{\mu}_{\mathcal{X}}}(V_{\mathit{1}i},X_i)) \\ = \frac{1}{n}\sum_{(v_{\mathit{1}}, x)\in \mathfrak{V}_{\mathit{1}}\times\mathfrak{X}} \left\{\hat{r}(v_{\mathit{1}}, x) \vphantom{ \left[\left(\sum_{i: (V_{\mathit{1}i},X_i)=(v_{\mathit{1}}, x)}m(V_i,X_i)\right) -\left( \sum_{i: (V_{\mathit{1}i},X_i)=(v_{\mathit{1}}, x)}{\hat{\mu}_{\mathcal{X}}}(v_{\mathit{1}}, x)\right) \right]} \right. \\ \left. \times \left[\left(\sum_{i: (V_{\mathit{1}i},X_i)=(v_{\mathit{1}}, x)}m(V_i,X_i)\right) -\left( \sum_{i: (V_{\mathit{1}i},X_i)=(v_{\mathit{1}}, x)}{\hat{\mu}_{\mathcal{X}}}(v_{\mathit{1}}, x)\right) \right]\right\} \\ = \frac{1}{n}\sum_{(v_{\mathit{1}}, x)\in \mathfrak{V}_{\mathit{1}}\times\mathfrak{X}}\hat{r}(v_{\mathit{1}}, x)\left[N_{v_{\mathit{1}}x}{\hat{\mu}_{\mathcal{X}}}(v_{\mathit{1}}, x)-N_{v_{\mathit{1}}x}{\hat{\mu}_{\mathcal{X}}}(v_{\mathit{1}}, x) \right] = 0. \end{align}\] The second term satisfies \[\begin{align} \mathbb{P}_n\left[\mathbb{1}_{V_{\mathit{2}}=c}g(V,X)-\hat{\gamma}_{\mathcal{V}}(c)\mathbb{1}_{V_{\mathit{2}}=c}\right]&=\mathbb{P}_n\mathbb{1}_{V_{\mathit{2}}=c}g(V,X) - \hat{\gamma}_{\mathcal{V}}(c)\mathbb{P}_n\mathbb{1}_{V_{\mathit{2}}=c} \\ &= N_c \hat{\gamma}_{\mathcal{V}}(c)-\hat{\gamma}_{\mathcal{V}}(c)N_c=0 \end{align}\] by the definition of \(\hat{\gamma}_{\mathcal{V}}(c), N_c\). The last term satisfies \[\mathbb{P}_n\left[f(V,X,{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c))-\chi(\hat{P}_{VX}))\right]=\mathbb{P}_nf(V,X,{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c))-\chi(\hat{P}_{VX})=0\] by the definition of \(\chi(\hat{P}_{VX})\) as \(\chi(\hat{P}_{VX})\) is constant with respect to the empirical measure \(\mathbb{P}_n\). Conclude that \(\mathbb{P}_n\hat{\tilde{\chi}}=0\).

Now we show asymptotic efficiency. Using that \(\mathbb{P}_n\hat{\tilde{\chi}}=0\), write \[\begin{align} \sqrt{n}(\chi(\hat{P}_{VX})-\chi(P_{VX})) &= \sqrt{n}\mathbb{P}_n\tilde{\chi }+ \sqrt{n}(\mathbb{P}_n-P_{VX})(\hat{\tilde{\chi}}-\tilde{\chi})+\sqrt{n} R_n, \\ R_n &= \chi(\hat{P}_{VX})-\chi(P_{VX}) + P_{VX}\hat{\tilde{\chi}}, \end{align}\] for \(\tilde{\chi}\) in ?? . The first term satisfies \(\sqrt{n}\mathbb{P}_n\tilde{\chi }\overset{P_{VX}}{\rightsquigarrow}\mathcal{N}(0, P_{VX}\tilde{\chi}^2)\). Because the arguments in 1 also apply when \(({\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}})\) and \(\chi(\hat{P}_{VX})\) are computed from the same sample \(\mathcal{S}\), we have \(\sqrt{n}R_n=o_{P_{VX}}\left(1\right)\). This follows from ?? of 1, because \({\hat{\mu}_{\mathcal{X}}}\), \(\hat{\gamma}_{\mathcal{V}}(c)\) and \(\hat{p}_{V_{\mathit{2}}}(c)\) are asymptotically normal, \(\hat{r}\) is consistent , \(P_{VX}\partial_\gamma^2 f(V,X,{\hat{\mu}_{\mathcal{X}}},\tilde{\gamma}_{\mathcal{V}}(c)=O_{{P_{VX}}}\left(1\right)\) by ?? , and, as we show at the end of this proof, \(\hat{e}-e'=o_{P_{VX}}\left(1\right)\) for \(e'=P_{VX}\partial_\gamma f(V,X,{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c))\).

It remains to show \(\sqrt{n}(\mathbb{P}_n-P_{VX})(\hat{\tilde{\chi}}-\tilde{\chi})=o_{P_{VX}}\left(1\right)\). Because 9 assumes that \(({\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}})\) are computed from a sample independent from what generates \(\mathbb{P}_n\), we need to adapt the arguments therein. As in 69 , write \[\begin{align} \hat{\tilde{\chi}}(v,x)-\tilde{\chi}(v,x) =&\, T_1(v,x)+T_2(v,x)+T_3(v,x)+T_4, \label{priv:eq:emproc95decomp95discrete} \\ T_1(v,x)\mathrel{\vcenter{:}}=&\, \hat{r}(v_{\mathit{1}},x)(m(v,x)-{\hat{\mu}_{\mathcal{X}}}(v_{\mathit{1}},x)) \nonumber \\ &-r(v_{\mathit{1}},x)(m(v,x)-\mu_{\mathcal{X}}(v_{\mathit{1}},x)), \nonumber \\ T_2(v,x)\mathrel{\vcenter{:}}=&\,\frac{\mathbb{1}_{v_{\mathit{2}}=c}}{\hat{p}_{V_{\mathit{2}}}(c)}(g(v,x)-\hat{\gamma}_{\mathcal{V}}(c))\hat{e}-\frac{\mathbb{1}_{v_{\mathit{2}}=c}}{p_{V_{\mathit{2}}}(c)}(g(v,x)-\gamma_{\mathcal{V}}(c))e, \nonumber \\ T_3(v,x)\mathrel{\vcenter{:}}=&\, f(v,x, {\hat{\mu}_{\mathcal{X}}}, \hat{\gamma}_{\mathcal{V}}(c)) -f(v,x, \mu_{\mathcal{X}}, \gamma_{\mathcal{V}}(c)), \nonumber \\ T_4\mathrel{\vcenter{:}}=&\, -\chi(\hat{P}_{VX})+\chi(P_{VX}). \nonumber \end{align}\tag{70}\] Again, \(T_4\) being a constant under \(\mathbb{P}_n-P_{VX}\), \(\sqrt{n}(\mathbb{P}_n-P_{VX})T_4=0\), so we need show \((\mathbb{P}_n-P_{VX})T_j=o_{P_{VX}}\left(n^{-1/2}\right)\) for \(j=1,2,3\).

Term \(T_1.\,\,\) Write \[\begin{align} \sqrt{n}(\mathbb{P}_n-P_{VX})T_1(V,X) = \sqrt{n}(\mathbb{P}_n-P_{VX})\left[m(V,X)(\hat{r}(V_{\mathit{1}},X)-r(V_{\mathit{1}},X))\right] \nonumber \\ +\sqrt{n}(\mathbb{P}_n-P_{VX})\left[{\hat{\mu}_{\mathcal{X}}}(V_{\mathit{1}},X)(\hat{r}(V_{\mathit{1}},X)-r(V_{\mathit{1}},X))\right] \tag{71} \\ +\sqrt{n}(\mathbb{P}_n-P_{VX})\left[r(V_{\mathit{1}},X)(\mu_{\mathcal{X}}(V_{\mathit{1}},X)-{\hat{\mu}_{\mathcal{X}}}(V_{\mathit{1}},X))\right]. \tag{72} \end{align}\] Because \(\mathfrak{V}_\mathit{1}\times\mathfrak{X}\) is finite, we can write \[\begin{align} r(v_{\mathit{1}},x)&=\sum_{(\bar{v}_{\mathit{1}},\bar{x})\in \mathfrak{V}_\mathit{1}\times\mathfrak{X}}\varrho_{\bar{v}_{\mathit{1}},\bar{x}}\mathbb{1}_{(\bar{v}_{\mathit{1}},\bar{x})}(v_{\mathit{1}},x), \quad \varrho_{\bar{v}_{\mathit{1}},\bar{x}}\mathrel{\vcenter{:}}= r(\bar{v}_{\mathit{1}},\bar{x}), \\ \hat{r}(v_{\mathit{1}},x)&=\sum_{(\bar{v}_{\mathit{1}},\bar{x})\in \mathfrak{V}_\mathit{1}\times\mathfrak{X}}\hat{\varrho}_{\bar{v}_{\mathit{1}},\bar{x}}\mathbb{1}_{(\bar{v}_{\mathit{1}},\bar{x})}(v_{\mathit{1}},x), \quad \hat{\varrho}_{\bar{v}_{\mathit{1}},\bar{x}}\mathrel{\vcenter{:}}=\hat{r}(\bar{v}_{\mathit{1}},\bar{x}). \end{align}\] But then \[\begin{align} \sqrt{n}(\mathbb{P}_n-P_{VX})\left[m(V,X)(\hat{r}(V_{\mathit{1}},X)-r(V_{\mathit{1}},X))\right] \\ =\sqrt{n}(\mathbb{P}_n-P_{VX})\left[m(V,X)\left(\sum_{(\bar{v}_{\mathit{1}},\bar{x})\in \mathfrak{V}_\mathit{1}\times\mathfrak{X}}(\hat{\varrho}_{\bar{v}_{\mathit{1}},\bar{x}}-\varrho_{\bar{v}_{\mathit{1}},\bar{x}})\mathbb{1}_{(\bar{v}_{\mathit{1}},\bar{x})}(V_{\mathit{1}},X)\right)\right] \\ = \sum_{(\bar{v}_{\mathit{1}},\bar{x})\in \mathfrak{V}_\mathit{1}\times\mathfrak{X}}(\hat{\varrho}_{\bar{v}_{\mathit{1}},\bar{x}}-\varrho_{\bar{v}_{\mathit{1}},\bar{x}})\sqrt{n}(\mathbb{P}_n-P_{VX})\left[ m(V,X)\mathbb{1}_{(\bar{v}_{\mathit{1}},\bar{x})}(V_{\mathit{1}},X)\right], \end{align}\] which is \(o_{P_{VX}}\left(1\right)\), because \(\hat{r}\) is consistent, \[\sqrt{n}(\mathbb{P}_n-P_{VX})\left[ m(V,X)\mathbb{1}_{(\bar{v}_{\mathit{1}},\bar{x})}(V_{\mathit{1}},X)\right]=O_{{P_{VX}}}\left(1\right)\] by the standard central limit theorem, and \(|\mathfrak{V}_\mathit{1}\times\mathfrak{X}|\) is finite. The terms 71 , 72 can be handled similarly because \(\left\lVert{{\hat{\mu}_{\mathcal{X}}}}\right\rVert_{\infty}\leq \left\lVert{{\hat{\mu}_{\mathcal{X}}}-\mu_{\mathcal{X}}}\right\rVert_{\infty}+\left\lVert{\mu_{\mathcal{X}}}\right\rVert_{\infty}=o_{P_{VX}}\left(1\right)+O\left(1\right)=O_{{P_{VX}}}\left(1\right)\) by assumption. Thus \(\sqrt{n}(\mathbb{P}_n-P_{VX})T_1=o_{P_{VX}}\left(1\right)\).

Term \(T_2.\,\,\) Same arguments as in 9 apply, yielding \(\sqrt{n}(\mathbb{P}_n-P_{VX})T_2\) \(=o_{P_{VX}}\left(1\right)\), because \(\hat{e}-e=o_{P_{VX}}\left(1\right)\) — which we show at the end of this proof —, and \(\hat{\gamma}_{\mathcal{V}}(c),\hat{p}_{V_{\mathit{2}}}(c)\) are consistent.

Term \(T_3.\,\,\) Write \[\begin{align} T_3(v,x)=&\,f(v,x,{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c))-f(v,x,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c)) \nonumber \\ =&\,f(v,x,{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c))-f(v,x,\mu_{\mathcal{X}},\hat{\gamma}_{\mathcal{V}}(c)) \nonumber \\ &+f(v,x,\mu_{\mathcal{X}},\hat{\gamma}_{\mathcal{V}}(c))-f(v,x,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c)). \label{priv:eq:emproc95decomp95discrete95t3952} \end{align}\tag{73}\] As with \(r\) before, represent \(\mu_{\mathcal{X}}(v_{\mathit{1}},x)=\sum_{(\bar{v}_{\mathit{1}},\bar{x})\in \mathfrak{V}_\mathit{1}\times\mathfrak{X}}\lambda_{\bar{v}_{\mathit{1}},\bar{x}}\mathbb{1}_{(\bar{v}_{\mathit{1}},\bar{x})}(v_{\mathit{1}},x)\) for some parameter \(\lambda\in\mathbb{R}^{|\mathfrak{V}_\mathit{1}\times\mathfrak{X}|}\) with \(\lambda_{\bar{v}_{\mathit{1}},\bar{x}}\mathrel{\vcenter{:}}=\mu_{\mathcal{X}}(\bar{v}_{\mathit{1}},\bar{x})\); similarly, write \[\begin{align} {\hat{\mu}_{\mathcal{X}}}(v_{\mathit{1}},x)&=\sum_{(\bar{v}_{\mathit{1}},\bar{x})\in \mathfrak{V}_\mathit{1}\times\mathfrak{X}}\hat{\lambda}_{\bar{v}_{\mathit{1}},\bar{x}}\mathbb{1}_{(\bar{v}_{\mathit{1}},\bar{x})}(v_{\mathit{1}},x), \\ \hat{\lambda}_{\bar{v}_{\mathit{1}},\bar{x}}&\mathrel{\vcenter{:}}={\hat{\mu}_{\mathcal{X}}}(\bar{v}_{\mathit{1}},\bar{x}). \end{align}\] By the linearity 5 of \(f\) and the mean value theorem, \[\begin{align} \sqrt{n}(\mathbb{P}_n-P_{VX})\left[f(V,X,{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c))-f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c))\right] \\ =\sqrt{n}(\mathbb{P}_n-P_{VX})\left[f(V,X,{\hat{\mu}_{\mathcal{X}}}-\mu_{\mathcal{X}},\hat{\gamma}_{\mathcal{V}}(c))\right] \\ =\sqrt{n}(\mathbb{P}_n-P_{VX})\left[\sum_{(\bar{v}_{\mathit{1}},\bar{x})\in \mathfrak{V}_\mathit{1}\times\mathfrak{X}}(\hat{\lambda}_{\bar{v}_{\mathit{1}},\bar{x}}-\lambda_{\bar{v}_{\mathit{1}},\bar{x}}) f(V,X,\mathbb{1}_{\bar{v}_{\mathit{1}},\bar{x}},\hat{\gamma}_{\mathcal{V}}(c))\right] \\ = \sum_{(\bar{v}_{\mathit{1}},\bar{x})\in \mathfrak{V}_\mathit{1}\times\mathfrak{X}}(\hat{\lambda}_{\bar{v}_{\mathit{1}},\bar{x}}-\lambda_{\bar{v}_{\mathit{1}},\bar{x}})\sqrt{n}(\mathbb{P}_n-P_{VX})\left[f(V,X,\mathbb{1}_{\bar{v}_{\mathit{1}},\bar{x}},\hat{\gamma}_{\mathcal{V}}(c))\right] \\ = \sum_{(\bar{v}_{\mathit{1}},\bar{x})\in \mathfrak{V}_\mathit{1}\times\mathfrak{X}}(\hat{\lambda}_{\bar{v}_{\mathit{1}},\bar{x}}-\lambda_{\bar{v}_{\mathit{1}},\bar{x}})\sqrt{n}(\mathbb{P}_n-P_{VX})\left[f(V,X,\mathbb{1}_{\bar{v}_{\mathit{1}},\bar{x}},\gamma_{\mathcal{V}}(c))\right] \\ +\sum_{(\bar{v}_{\mathit{1}},\bar{x})\in \mathfrak{V}_\mathit{1}\times\mathfrak{X}}(\hat{\lambda}_{\bar{v}_{\mathit{1}},\bar{x}}-\lambda_{\bar{v}_{\mathit{1}},\bar{x}})(\mathbb{P}_n-P_{VX})\left[\partial_\gamma f(V,X,\mathbb{1}_{\bar{v}_{\mathit{1}},\bar{x}},\tilde{\gamma}_{\mathcal{V}}(c))\right] \\ \times \sqrt{n}(\hat{\gamma}_{\mathcal{V}}(c)-\gamma_{\mathcal{V}}(c)) \end{align}\] for some \(\tilde{\gamma}_{\mathcal{V}}(c)\) between \(\gamma_{\mathcal{V}}(c)\) and \(\hat{\gamma}_{\mathcal{V}}(c)\). But this is \(o_{P_{VX}}\left(1\right)\) by the standard central limit theorem, and because \(\sqrt{n}(\hat{\gamma}_{\mathcal{V}}(c)-\gamma_{\mathcal{V}}(c))=O_{{P_{VX}}}\left(1\right)\), \({\hat{\mu}_{\mathcal{X}}}\) being consistent and \((\mathbb{P}_n-P_{VX})\left[\partial_\gamma f(V,X,\mathbb{1}_{\bar{v}_{\mathit{1}},\bar{x}},\tilde{\gamma}_{\mathcal{V}}(c))\right]=O_{{P_{VX}}}\left(1\right)\) by ?? . For the term 73 , apply the mean-value theorem twice to get \[\begin{align} T_{3,1}(v,x)\mathrel{\vcenter{:}}= f(v,x,\mu_{\mathcal{X}},\hat{\gamma}_{\mathcal{V}}(c))-f(v,x,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c)) \\ =(\hat{\gamma}_{\mathcal{V}}(c)-\gamma_{\mathcal{V}}(c))\left\{\partial_\gamma f(v,x,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c))+(\tilde{\gamma}_{\mathcal{V}}(c)-\gamma_{\mathcal{V}}(c))\partial_\gamma^2 f(v,x,\mu_{\mathcal{X}},\tilde{\gamma}_{\mathcal{V}}'(c))\right\} \end{align}\] for some \(\tilde{\gamma}_{\mathcal{V}}'(c)\) between \(\gamma_{\mathcal{V}}(c)\) and \(\hat{\gamma}_{\mathcal{V}}(c)\). Thus \(\sqrt{n}(\mathbb{P}_n-P_{VX})T_{3,1}=o_{P_{VX}}\left(1\right)\) by the standard central limit theorem, \(\sqrt{n}\)-consistency of \(\hat{\gamma}_{\mathcal{V}}(c)\), and stochastic boundedness ?? of the second derivative. Hence, \(\sqrt{n}(\mathbb{P}_n-P_{VX})T_3=o_{P_{VX}}\left(1\right)\).

Consistency of \(\hat{e}.\,\,\) Write \[\begin{align} \hat{e}- e=&\, \mathbb{P}_n\partial_\gamma f(V,X,{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c))-P_{VX}\partial_\gamma f(V,X,{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c)) \nonumber \\ =&\, (\mathbb{P}_n- P_{VX}) \partial_\gamma f(V,X,{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c)) \tag{74} \\ &+ P_{VX}\left[ \partial_\gamma f(V,X,{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c))- \partial_\gamma f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c))\right]. \tag{75} \end{align}\] As \(\mu\mapsto \partial_\gamma f(v,x,\mu,\bar{\gamma})\) is linear, 74 is \(o_{P_{VX}}\left(1\right)\) along similar arguments as \(T_3\) above; because we only need consistency, we only need the existence and stochastic boundedness ?? of the second derivative; no need for higher order derivatives. For the term 75 , write it as \[\begin{align} P_{VX}\left[ \partial_\gamma f(V,X,{\hat{\mu}_{\mathcal{X}}},\hat{\gamma}_{\mathcal{V}}(c))- \partial_\gamma f(V,X,\mu_{\mathcal{X}},\hat{\gamma}_{\mathcal{V}}(c))\right] \\ +P_{VX}\left[ \partial_\gamma f(V,X,\mu_{\mathcal{X}},\hat{\gamma}_{\mathcal{V}}(c))- \partial_\gamma f(V,X,\mu_{\mathcal{X}},\gamma_{\mathcal{V}}(c))\right]. \end{align}\] By the linearity of \(\mu\mapsto \partial_\gamma f(v,x,\mu,\hat{\gamma}_{\mathcal{V}}(c))\), the first term here is \(o_{P_{VX}}\left(1\right)\), following a mean-value expansion, by the consistency of \({\hat{\mu}_{\mathcal{X}}}\) and the assumed boundedness ?? of \(P_{VX}\partial_\gamma f(V,X,\mu_{\mathcal{X}},\hat{\gamma}_{\mathcal{V}}(c))\). The second term is \(o_{P_{VX}}\left(1\right)\) by the same arguments under assumption ?? for \(P_{VX}\partial_\gamma^2 f(V,X,\mu_{\mathcal{X}},\hat{\gamma}_{\mathcal{V}}(c))\). Hence, \(\hat{e}-e=o_{P_{VX}}\left(1\right)\). Finally, note that \(\hat{e}-e'=o_{P_{VX}}\left(1\right)\) too as we claimed, because it is equal to 74 . ◻

References↩︎

[1]
R. F. Barber and J. C. Duchi, “Privacy and Statistical Risk: Formalisms and Minimax Bounds.” arXiv, Dec. 2014, Accessed: Aug. 23, 2024. [Online]. Available: https://arxiv.org/abs/1412.4451.
[2]
C. Dwork, F. McSherry, K. Nissim, and A. Smith, Calibrating Noise to Sensitivity in Private Data Analysis,” in Theory of Cryptography, vol. 3876, D. Hutchison, T. Kanade, J. Kittler, J. M. Kleinberg, F. Mattern, J. C. Mitchell, M. Naor, O. Nierstrasz, C. Pandu Rangan, B. Steffen, M. Sudan, D. Terzopoulos, D. Tygar, M. Y. Vardi, G. Weikum, S. Halevi, and T. Rabin, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2006, pp. 265–284.
[3]
A. Rotnitzky, E. Smucler, and J. M. Robins, “Characterization of Parameters With a Mixed Bias Property,” Biometrika, vol. 108, no. 1, pp. 231–238, Mar. 2021, doi: 10.1093/biomet/asaa054.
[4]
V. Chernozhukov, W. K. Newey, and R. Singh, “Automatic Debiased Machine Learning of Causal and Structural Effects,” Econometrica, vol. 90, no. 3, pp. 967–1027, 2022, doi: 10.3982/ECTA18515.
[5]
L. P. Hansen, “Large Sample Properties of Generalized Method of Moments Estimators,” Econometrica, vol. 50, no. 4, p. 1029, Jul. 1982, doi: 10.2307/1912775.
[6]
W. K. Newey and D. McFadden, Chapter 36: Large Sample Estimation and Hypothesis Testing,” in Handbook of Econometrics, vol. 4, Elsevier, 1994, pp. 2111–2245.
[7]
S. L. Warner, “Randomized Response: A Survey Technique for Eliminating Evasive Answer Bias,” Journal of the American Statistical Association, vol. 60, no. 309, pp. 63–69, Mar. 1965, doi: 10.1080/01621459.1965.10480775.
[8]
A. Evfimievski, J. Gehrke, and R. Srikant, “Limiting Privacy Breaches in Privacy Preserving Data Mining,” in Proceedings of the twenty-second ACM SIGMOD-SIGACT-SIGART symposium on Principles of database systems, Jun. 2003, pp. 211–222, doi: 10.1145/773153.773174.
[9]
D. Desfontaines and B. Pejó, SoK: Differential Privacies.” arXiv, Nov. 2022, Accessed: May 23, 2024. [Online]. Available: https://arxiv.org/abs/1906.01337.
[10]
H. Asi, J. C. Duchi, and O. Javidbakht, “Element Level Differential Privacy: The Right Granularity of Privacy.” arXiv, Dec. 2019, Accessed: Sep. 13, 2024. [Online]. Available: https://arxiv.org/abs/1912.04042.
[11]
C. Gentry, “A Fully Homomorphic Encryption Scheme,” PhD thesis, Stanford University, 2009.
[12]
Y. Yang et al., “A Comprehensive Survey on Secure Outsourced Computation and Its Applications,” IEEE Access, vol. 7, pp. 159426–159465, Oct. 2019, doi: 10.1109/ACCESS.2019.2949782.
[13]
A. Smith, “Efficient, Differentially Private Point Estimators.” arXiv, Sep. 2008, Accessed: May 17, 2023. [Online]. Available: https://arxiv.org/abs/0809.4794.
[14]
O. Sheffet, “Differentially Private Ordinary Least Squares,” in Proceedings of the 34th International Conference on Machine Learning, Aug. 2017, vol. 70, pp. 3105–3114.
[15]
D. Alabi, A. McMillan, J. Sarathy, A. Smith, and S. Vadhan, “Differentially Private Simple Linear Regression.” arXiv, Jul. 2020, Accessed: May 23, 2023. [Online]. Available: https://arxiv.org/abs/2007.05157.
[16]
Y. Jiang, Y. Liu, X. Yan, A.-S. Charest, L. Kong, and B. Jiang, “Analysis of Differentially Private Synthetic Data: A Measurement Error Approach,” in AAAI Conference on Artificial Intelligence, 2024.
[17]
J. Drechsler, I. Globus-Harris, A. McMillan, J. Sarathy, and A. Smith, “Non-Parametric Differentially Private Confidence Intervals for the Median.” arXiv, Jul. 2021, Accessed: May 17, 2023. [Online]. Available: https://arxiv.org/abs/2106.10333.
[18]
P.-L. Loh and M. J. Wainwright, “High-Dimensional Regression With Noisy and Missing Data: Provable Guarantees With Nonconvexity,” The Annals of Statistics, vol. 40, no. 3, Jun. 2012, doi: 10.1214/12-AOS1018.
[19]
J. Acharya, Z. Sun, and H. Zhang, Hadamard Response: Estimating Distributions Privately, Efficiently, and With Little Communication,” in Proceedings of the twenty-second international conference on artificial intelligence and statistics, 2019, vol. 89, pp. 1120–1129.
[20]
T. B. Berrett, L. Györfi, and H. Walk, “Strongly Universally Consistent Nonparametric Regression and Classification With Privatised Data,” Electronic Journal of Statistics, vol. 15, no. 1, Jan. 2021, doi: 10.1214/21-EJS1845.
[21]
L. P. Barnes, W.-N. Chen, and A. Özgür, “Fisher Information Under Local Differential Privacy,” IEEE Journal on Selected Areas in Information Theory, vol. 1, no. 3, pp. 645–659, 2020, doi: 10.1109/JSAIT.2020.3039461.
[22]
M. J. Kusner, Y. Sun, K. Sridharan, and K. Q. Weinberger, “Private causal inference,” in Proceedings of the 19th international conference on artificial intelligence and statistics, 2016, vol. 51, pp. 1308–1317.
[23]
Y. Zhu, L. Gultchin, A. Gretton, M. Kusner, and R. Silva, “Causal Inference With Treatment Measurement Error: A Nonparametric Instrumental Variable Approach.” arXiv, Jun. 2022, Accessed: Jun. 23, 2023. [Online]. Available: https://arxiv.org/abs/2206.09186.
[24]
Y. Ohnishi and J. Awan, “Locally Private Causal Inference for Randomized Experiments.” arXiv, 2023, doi: 10.48550/ARXIV.2301.01616.
[25]
A. Agarwal and R. Singh, “Causal Inference With Corrupted Data: Measurement Error, Missing Values, Discretization, and Differential Privacy.” arXiv, Feb. 2024, Accessed: Aug. 26, 2024. [Online]. Available: https://arxiv.org/abs/2107.02780.
[26]
L. Steinberger, “Efficiency in Local Differential Privacy.” arXiv, Jan. 2023, Accessed: Mar. 27, 2023. [Online]. Available: https://arxiv.org/abs/2301.10600.
[27]
J. C. Duchi, M. I. Jordan, and M. J. Wainwright, “Minimax Optimal Procedures for Locally Private Estimation,” Journal of the American Statistical Association, vol. 113, no. 521, pp. 182–201, Jan. 2018, doi: 10.1080/01621459.2017.1389735.
[28]
J. C. Duchi and F. Ruan, “The Right Complexity Measure in Locally Private Estimation: It Is Not the Fisher Information,” The Annals of Statistics, vol. 52, no. 1, Feb. 2024, doi: 10.1214/22-AOS2227.
[29]
C. Butucea, A. Rohde, and L. Steinberger, “Interactive versus noninteractive locally differentially private estimation: Two elbows for the quadratic functional,” The Annals of Statistics, vol. 51, no. 2, Apr. 2023, doi: 10.1214/22-AOS2254.
[30]
K. Chaudhuri, C. Monteleoni, and A. D. Sarwate, “Differentially Private Empirical Risk Minimization.” arXiv, Feb. 2011, Accessed: May 24, 2023. [Online]. Available: https://arxiv.org/abs/0912.0071.
[31]
D. Kifer, A. Smith, and A. Thakurta, “Private Convex Empirical Risk Minimization and High-Dimensional Regression,” in Proceedings of the 25th Annual Conference on Learning Theory, Jun. 2012, vol. 23, pp. 25.1–25.40.
[32]
R. Bassily, A. Smith, and A. Thakurta, “Differentially Private Empirical Risk Minimization: Efficient Algorithms and Tight Error Bounds.” arXiv, Oct. 2014, Accessed: May 24, 2023. [Online]. Available: https://arxiv.org/abs/1405.7085.
[33]
K. Fukuchi, Q. K. Tran, and J. Sakuma, Differentially Private Empirical Risk Minimization With Input Perturbation,” in Discovery Science, vol. 10558, A. Yamamoto, T. Kida, T. Uno, and T. Kuboyama, Eds. Cham: Springer International Publishing, 2017, pp. 82–90.
[34]
J. Lei, “Differentially Private M-Estimators,” in NIPS, 2011.
[35]
A. B. Slavkovic and R. Molinari, “Perturbed M-Estimation: A Further Investigation of Robust Statistics for Differential Privacy.” 2021.
[36]
P. Mangold, A. Bellet, J. Salmon, and M. Tommasi, High-Dimensional Private Empirical Risk Minimization by Greedy Coordinate Descent,” in Proceedings of the 26th international conference on artificial intelligence and statistics, 2023, vol. 206, pp. 4894–4916.
[37]
H. Asi and J. C. Duchi, “Near Instance-Optimality in Differential Privacy.” arXiv, May 2020, Accessed: Sep. 13, 2024. [Online]. Available: https://arxiv.org/abs/2005.10630.
[38]
E. Bolthausen, A. W. van der Vaart, and E. Perkins, Semiparametric Statistics,” in Lectures on Probability Theory and Statistics, vol. 1781, P. Bernard, J.-M. Morel, F. Takens, and B. Teissier, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2002.
[39]
A. Van Der Vaart, “Higher Order Tangent Spaces and Influence Functions,” Statist. Sci., vol. 29, no. 4, Nov. 2014, doi: 10.1214/14-STS478.
[40]
E. H. Kennedy, Semiparametric Doubly Robust Targeted Double Machine Learning: A Review,” in Handbook of Statistical Methods for Precision Medicine, 1st ed., Boca Raton: Chapman and Hall/CRC, 2024, pp. 207–236.
[41]
R. Piziak and P. L. Odell, Matrix Theory, 1st ed. Chapman and Hall/CRC, 2007.
[42]
J. Neyman, “On the Applications of the Theory of Probability to Agricultural Experiments,” PhD thesis, University of Warsaw, 1924.
[43]
D. B. Rubin, “Estimating Causal Effects of Treatments in Randomized and Nonrandomized Studies,” Journal of Educational Psychology, vol. 66, no. 5, pp. 688–701, 1974, doi: 10.1037/h0037350.
[44]
J. Hahn, “On the Role of the Propensity Score in Efficient Semiparametric Estimation of Average Treatment Effects,” Econometrica, vol. 66, no. 2, p. 315, Mar. 1998, doi: 10.2307/2998560.
[45]
K. Ray and A. Van Der Vaart, “Semiparametric Bayesian causal inference,” The Annals of Statistics, vol. 48, no. 5, Oct. 2020, doi: 10.1214/19-AOS1919.
[46]
A. W. van der Vaart, Asymptotic Statistics. Cambridge University Press, 1998.
[47]
D. Pollard, “Asymptopia: An Exposition of Statistical Asymptotic Theory,” 2005.
[48]
A. D. Polânin and A. V. Manzhirov, Handbook of Integral Equations, 1st ed. Boca Raton, London, New York, Washington: CRC press, 1998.
[49]
R. Kress, Linear Integral Equations, 3rd ed. New York: Springer, 2014.
[50]
M. Reed and B. Simon, Methods of Modern Mathematical Physics. New York: Academic Press, 1972.
[51]
A. B. Tsybakov, Introduction to Nonparametric Estimation. New York, NY: Springer New York, 2009.
[52]
M. Kohler, “Multivariate Orthogonal Series Estimates for Random Design Regression,” Journal of Statistical Planning and Inference, vol. 138, no. 10, pp. 3217–3237, Oct. 2008, doi: 10.1016/j.jspi.2008.01.011.

  1. Their sufficient conditions for a product-form bias [3] stipulate a parameter structure where both factors in the product are ratios of two regressions, with the same denominator in both. In general, the variationally dependent denominators do not naturally translate into double-robustness. An exception is when the denominator and the nominator are chosen so that the resulting ratio in each factor is itself a regression function. But then these parameters are strictly included in our class.↩︎

  2. Some other privacy notions are element-level privacy [10] or homomorphic encryption [11], [12].↩︎

  3. The existence of an \(L_Q\) satisfying 13 is not necessary (but clearly sufficient) for the identification of some parameters — e.g.when \(\chi\) is the functional of only the marginal \(P_V\).↩︎

  4. See 1 in 10.1 for further semiparametric properties.↩︎

  5. If \(\mathcal{S}=\mathcal{(}(V_i,X_i))_{i\in[n]}\), \(\mathcal{S}'= ((V_i',X_i'))_{i\in[n]}\), \(\mathcal{S}''=((V_i'',X_i''))_{i\in[n]}\) are three, mutually independent, random samples from \(P_{VX}\), then, given a \(Q\in\mathcal{Q}_{\psi}\), the samples \(\bar{\mathcal{S}}',\bar{\mathcal{S}}''\) are obtained by drawing \(Z_i'{\vert\,}(\mathcal{S},\bar{\mathcal{S}},\mathcal{S}',\mathcal{S}'')\sim Q(\cdot{\vert\,}X_i')\) and \(Z_i''{\vert\,}(\mathcal{S},\bar{\mathcal{S}},\mathcal{S}',\bar{\mathcal{S}}',\mathcal{S}'')\sim Q(\cdot{\vert\,}X_i'')\) for all \(i\in[n]\).↩︎