Local increment inference for time-inhomogeneous drift in Gaussian processes


Abstract

We study statistical inference for deterministic drift structures in Gaussian process models under high-frequency observations. The observed process consists of a centered stationary Gaussian component combined with a broad class of time-inhomogeneous deterministic drifts. To estimate the drift parameter, we introduce a least squares-type contrast based on first-order increments. We establish consistency and asymptotic normality under weak dependence conditions on the Gaussian component. A central feature of the framework is that the rate of convergence of the estimator depends jointly on the local roughness of the Gaussian noise and the long-time information accumulation structure generated by the drift. The theory accommodates a wide range of drift families, including integrable, polynomial-type, and periodic structures. In particular, different drift densities produce qualitatively different statistical regimes, including non-standard rates of convergence and accelerated rates for persistent or growing deterministic structures.

Keywords: Gaussian processes, high-frequency data, time-inhomogeneous drift, contrast-based estimation, method of moments
MSC2010: 62M10; 62F12, 60G15.

1 Introduction↩︎

Gaussian processes (GPs) provide a flexible probabilistic framework for modeling time-dependent phenomena with uncertainty quantification. Because of their analytical tractability and nonparametric structure, they are widely used in statistics, machine learning, spatial statistics, and time series analysis. See, for example, Ibragimov and Rozanov [1] and Rasmussen and Williams [2].

In many practical situations, observed data exhibit both deterministic long-term trends and stochastic short-term fluctuations. This naturally leads to models consisting of a time-inhomogeneous drift combined with a stationary Gaussian component.

In this paper, we consider a stochastic process of the form \[\begin{align} X_t = Z_t + \int_0^t \mu(s)\,\mathrm{d}s, \quad t\ge0, \label{true} \end{align}\tag{1}\] where \(\mu\) is a deterministic mean density \(\mu\), say drift, and \(Z=(Z_t)_{t\ge0}\) is a centered stationary Gaussian process with covariance kernel \(K\) satisfying \[\mathbb{E}[Z_tZ_s]=K(|t-s|), \quad t,s\ge0.\] The process is observed at time points \(t_i=ih_n\;(i=1,2,\dots,n)\) under high-frequency sampling: \[h_n\to0, \quad T_n:=nh_n\to\infty\quad (n\to \infty).\] The primary objective of the present paper is statistical inference for the drift \(\mu\) from high-frequency observations of the process \(X\).

In many empirical analyses of Gaussian process models, the observed data are first centered by subtracting the sample mean and then treated as approximately mean-zero observations. Such preprocessing is convenient for covariance analysis and kernel estimation, particularly in stationary settings.

However, this empirical centering procedure may fundamentally alter the underlying deterministic structure when the drift component is nonstationary. Indeed, subtracting the sample mean removes not only global location information but also part of the low-frequency deterministic signal contained in the drift itself. Consequently, the resulting centered process may no longer preserve the original statistical structure of the drift component.

From a mathematical viewpoint, empirical centering implicitly assumes that the deterministic component behaves asymptotically as a constant average level over the observation horizon. This assumption becomes problematic for time-inhomogeneous drifts, especially under high-frequency asymptotics with diverging observation horizons. Therefore, statistical inference for nonstationary drift structures should be formulated directly without relying on empirical mean removal.

Our approach is based on the local increments \[\Delta_i^n X := X_{t_i}-X_{t_{i-1}},\quad i=1,2,\dots,n,\] and a a least squares-type estimation, which is a standard methodology to estimate the trend of processes.

So far, parameter estimation for Gaussian processes has been extensively studied. Classical approaches include maximum likelihood estimation, composite likelihood methods, and frequency-domain procedures; see Rasmussen and Williams [2], Varin et al. [3], Whittle [4], Fukasawa and Takabatake [5], and Takabatake [6], among others. Most existing studies primarily focus on covariance structures, spectral characteristics, or likelihood approximation techniques for approximately mean-zero Gaussian processes.

Statistical inference for deterministic mean functions has received relatively less attention, particularly under high-frequency asymptotics with diverging observation horizons. Recent work by Kobayashi et al. [7] studied maximum likelihood estimation for mean functions of Gaussian processes under small noise asymptotics through continuous-time likelihood analysis.

In contrast, the present paper develops a local increment-based approach under high-frequency discrete observations. In the author’s experience, specifying the mean function is often essential for predicting future phenomena from real data via Gaussian process models. Rather than treating the deterministic trend as a nuisance component to be removed by empirical centering, we regard it as the primary statistical signal of interest.

Our methodology is closely related to local inference methods for diffusion-type processes. In diffusion statistics, local increment structures play a fundamental role in constructing tractable estimators under high-frequency observations; see, for example, Kessler [8] and references therein. The present paper extends a similar local inference philosophy to Gaussian process models with deterministic nonstationary drifts. In particular, the framework accommodates a broad class of drift families, including integrable, polynomial-type, and periodic structures, among others.

A key feature of the proposed approach is that the asymptotic analysis is formulated directly through local covariance structures of increments, without relying on global ergodic representations or mixing-based asymptotic arguments. This allows us to treat a broad class of Gaussian processes in a unified framework.

The asymptotic behavior of the estimator is governed jointly by the information accumulation structure of the drift family and the dependence accumulation structure of the Gaussian increments. The resulting rate of convergence depends explicitly on how rapidly the deterministic signal accumulates over the observation horizon and on the local dependence structure of the Gaussian noise. Moreover, even non-smooth covariance kernels can be treated naturally within the present local framework.

The remainder of this paper is organized as follows. Section 2 introduces a parametric model and its assumptions. Section 3.1 presents the consistency and asymptotic normality results together with several representative examples illustrating different information accumulation regimes. Section 4 concludes the paper with several remarks and possible future extensions. Proofs of the main results are collected in Section 5.

Notation↩︎

  • The random vector \(Z\) follows the normal distribution with mean vector \(m\) and covariance matrix \(\Sigma\), we write \(Z\sim {\cal N}(m,\Sigma)\).

  • For a centered stationary Gaussian process \(Z=(Z_t)_{t\ge 0}\) with the kernel function \(K\), we write \(Z\sim GP(0,K)\).

  • For a subet \(S\subset \mathbb{R}^p\), \(\overline{S}\) is the closure of \(S\) w.r.t. the Euclidian norm.

  • For two nonnegative sequences \(a_n\) and \(b_n\), we write \(a_n\asymp b_n\) if there exist constants \(0<c<C<\infty\) such that \(cb_n\le a_n\le Cb_n\) for all sufficiently large \(n\). Moreover, we write \(a_n\sim b_n\) if \(a_n/b_n\to1.\)

  • For a function \(f(x,y):\mathbb{R}^d\times\mathbb{R}^{d'}\to\mathbb{R}\) and \(x=(x_1,\dots,x_d)\), \[\partial_x f = \left(\frac{\partial f}{\partial x_1},\dots,\frac{\partial f}{\partial x_d}\right)\in\mathbb{R}^d, \qquad \partial_x^2 f = \left(\frac{\partial^2 f}{\partial x_i\partial x_j}\right)_{1\le i,j\le d}\in\mathbb{R}^{d\otimes d},\] provided the derivatives exist.

2 Models and Assumptions↩︎

For the true process 1 , we consider a parametric model: \[\begin{align} X_t = Z_t + \int_0^t \mu_\xi(s)\,\mathrm{d}s, \quad X_0 = Z_0, \label{model} \end{align}\tag{2}\] where \(\mu_\xi:[0,\infty)\to\mathbb{R}\) is a deterministic drift depending on an unknown parameter \(\xi\in \mathbb{R}^p\), and \(Z=(Z_t)_{t\ge0}\sim GP(0,K)\) is a centered stationary Gaussian process with covariance kernel \(K\).

The primary objective of this paper is statistical inference for the drift parameter \(\xi\) from discrete observations of \(X\). The covariance structure of the Gaussian component is treated mainly as a nuisance component, representing dependent noise rather than the primary inferential target.

We observe the process at \(t_i = ih_n\;(i=1,2,\dots,n)\), with high-frequency sampling scheme: \[\begin{align} h_n \to 0, \quad T_n:=nh_n \to \infty, \quad n\to\infty. \label{sampling} \end{align}\tag{3}\]

We assume that the drift belongs to a parametric family \[\left\{\mu_\xi:[0,\infty)\to\mathbb{R}\,\Big|\, \xi\in\overline{\Xi}, \int_0^t |\mu_\xi(s)|\,\mathrm{d}s < \infty\;for any t>0\right\},\] where \(\Xi\subset\mathbb{R}^p\) is an open bounded convex set. The true parameter is denoted by \[\xi_0\in\Xi,\quad \mu_{\xi_0}\equiv \mu.\]

We shall make the following assumptions for the Gaussian noise \(Z\) with the notation such as \[\Delta_i^n Z := Z_{t_i}- Z_{t_{i-1}}.\]

A 1. There exists constants \(C>0\) and \(\alpha \in(0,2]\) such that \[\sup_{1\le i\le n}\sum_{j=1}^n \left|\mathrm{Cov}(\Delta_i^nZ,\Delta_j^nZ)\right|\le C h_n^\alpha.\]

A 2. There exist constants \(\beta\in(0,2]\) and \(c_K>0\) such that \[K(0)-K(t) = c_K t^\beta + o(t^\beta), \quad t\downarrow0.\]

Assumption A1 is a short-range dependence condition for local increments. It controls the cumulative dependence of \(\Delta_i^nZ\) on the other increments at the same sampling scale. For instance, suppose that \(K\in C^2([0,\infty))\) and \(\int_0^\infty |\partial_t^2K(t)|\,\mathrm{d}t<\infty\). Then Taylor expansion yields \[\begin{align} |\mathrm{Cov}(\Delta_i^nZ,\Delta_j^nZ)|\le C h_n^2 \sup_{u\in I_{ij}^n} |\partial_t^2K(u)|, \label{short-range-c2} \end{align}\tag{4}\] where \(I_{ij}^n\) is an interval between \((|i-j|-1)h_n\) and \((|i-j|+1)h_n\). Consequently, A1 holds with \(\alpha=2\).

On the other hand, the Ornstein-Uhlenbeck kernel \(K(t)=\sigma^2\exp(-\lambda |t|)\), we have \(\beta=1\) and \(K(0)-K(t)\sim \sigma^2\lambda t\) as \(t\downarrow0\). Moreover, for \(i\ne j\), \[\begin{align} \mathrm{Cov}(\Delta_i^nZ,\Delta_j^nZ)=\sigma^2\left(2e^{-\lambda |i-j|h_n}-e^{-\lambda |i-j-1|h_n}-e^{-\lambda |i-j+1|h_n}\right). \end{align}\] For \(|i-j|\ge1\), the right-hand side is bounded by \(C h_n^2 e^{-c|i-j|h_n}\). Therefore, \[\begin{align} \sum_{j\ne i}\left|\mathrm{Cov}(\Delta_i^nZ,\Delta_j^nZ)\right|\le C h_n^2\sum_{k=1}^\infty e^{-ckh_n}\le C h_n.\label{short-range-ou} \end{align}\tag{5}\] Thus A1 holds with \(\alpha=1\). Therefore, the present framework naturally accommodates a broad class of Gaussian processes without requiring global smoothness assumptions on the covariance kernel.

Assumption A2 characterizes the local covariance behavior of the Gaussian process. Under this condition, the increment variance satisfies \[\mathrm{Var}(Z_{t+h}-Z_t) = 2c_K h^\beta + o(h^\beta), \quad h\downarrow0.\] The parameter \(\beta\) controls the small-scale behavior of the Gaussian increments. Larger values of \(\beta\) correspond to smoother sample paths and smaller local increment variances. It also determines the asymptotic scaling of the local Gaussian contrast introduced later. This formulation is actually flexible and includes both smooth and non-smooth Gaussian kernels. For instance, kernels satisfying \(K\in C^2([0,\infty))\) with \(\partial_t^2K(0)\neq0\) correspond to the case \(\beta=2\). On the other hand, non-smooth models such as the Ornstein-Uhlenbeck process satisfy \(\beta=1\).

Remark 1. The exponents \(\alpha\) in A1 and \(\beta\) in A2 describe different aspects of the Gaussian noise structure. More precisely, the exponent \(\alpha\) controls the accumulation of dependence among Gaussian increments through the covariance summability condition in A1. On the other hand, \(\beta\) describes the local roughness of the Gaussian process through the scaling of the increment variance near the origin. From the viewpoint of statistical inference, the consistency condition is essentially determined by \(\alpha\), whereas the asymptotic normalization in the central limit theorem is governed by \(\beta\). However, in many standard stationary Gaussian models, including Ornstein-Uhlenbeck and short-memory fractional Brownian motions, one naturally has \(\alpha=\beta\), as in the examples described above.

Next, we shall make assumptions on the parametric drift family.

B 1. For every \(t\ge0\), the map \(\xi\mapsto\mu_\xi(t)\) is continuous on \(\overline{\Xi}\).

B 2. For every \(t\ge0\), the map \(\xi\mapsto\mu_\xi(t)\) is twice continuously differentiable on \(\overline{\Xi}\).

B 3. There exists a constant \(C>0\) such that \[\sup_{\xi\in\overline{\Xi}}\sup_{t\ge0}|\partial_t\mu_\xi(t)|\le C.\]

B 4. There exists a deterministic sequence \(a_n\to\infty\) and a nonnegative continuous function \(Q:\Xi\times\Xi\to[0,\infty)\) such that \[\sup_{\xi\in\overline{\Xi}}\left|\frac{1}{a_n}\sum_{i=1}^n\bigl(\mu_\xi(t_{i-1})-\mu_{\xi_0}(t_{i-1})\bigr)^2-Q(\xi,\xi_0)\right|\to0,\quad n\to \infty,\] and that the limit funciton \(Q\) meets the following condition: \[Q(\xi,\xi_0)=0 \iff \xi=\xi_0\quad on \overline{\Xi}.\]

B 5. There exist a deterministic sequence \(b_n\to\infty\) and a measurable function \(M:[0,\infty)\to[0,\infty)\) such that \[\sup_{\xi\in\overline{\Xi}} |\partial_\xi\mu_\xi(t)| \le M(t), \quad t\ge0,\] and \[\frac{1}{b_n}\sum_{i=1}^n M(t_{i-1})^2 = O(1)\quad n\to \infty.\]

Assumption B4 describes the identifiability of drift parameters through discrete observations. The normalization sequence \(a_n\) represents the effective accumulation rate of statistical information generated by the drift structure. Depending on the long-time behavior of the drift, different accumulation rates may arise. Several representative examples are discussed later.

Assumption B5 controls the parameter sensitivity of the drift family. The normalization sequence \(b_n\) characterizes the quadratic accumulation rate of local parameter sensitivity over the observation horizon.

2.0.0.1 LSE-type contrast functions:

Since the local increments is given as follows: for \(i=1,\dots,n\), \[\Delta_i^n X := X_{t_i}-X_{t_{i-1}} = \int_{t_{i-1}}^{t_i}\mu_\xi(s)\,\mathrm{d}s + \Delta_i^n Z,\] and it follows by the local integrabiolity of \(\mu_\xi\) that \[\int_{t_{i-1}}^{t_i}\mu_\xi(s)\,\mathrm{d}s = \mu_\xi(t_{i-1})h_n + O(h_n),\] we have \[\Delta_i^n Z =\Delta_i^n X - \mu_\xi(t_{i-1})h_n + o(h_n),\quad n\to \infty.\] This local decomposition forms the basis of our estimation procedure. Moreover, by Assumption A2, \(\mathrm{Var}(\Delta_i^n Z) = 2(K(0)-K(h_n))\), which implies that \(\Delta_i^n X - \mu_\xi(t_{i-1})h_n\sim {\cal N}(0, 2[K(0)-K(h_n)])\). Since this variance does not depend on \(\xi\), the corresponding local quasi-log-likelihood reduces to a least squares-type contrast. This motivates the following contrast function: \[\begin{align} H_n(\xi) := \sum_{i=1}^n\bigl(\Delta_i^n X-\mu_\xi(t_{i-1})h_n\bigr)^2. \label{contrast} \end{align}\tag{6}\] The estimator is then defined by \[\begin{align} \widehat{\xi}_n \in \arg\min_{\xi\in\overline{\Xi}} H_n(\xi). \label{estimator} \end{align}\tag{7}\]

3 Main Results and Examples↩︎

3.1 Main Theorems↩︎

We first establish the consistency of the estimator \(\widehat{\xi}_n\) defined in 7 . For this purpose, we analyze the asymptotic behavior of the least squares-type contrast function 6 .

The following theorem gives consistency of the estimator.

Theorem 1. Suppose Assumptions A1, B1, and B3–B5 hold. Moreover, assume that \[\begin{align} a_n h_n^{2-\alpha}\to\infty;\quad \frac{nh_n^2}{a_n}\to 0;\quad b_n=O(a_n). \label{balance-beta} \end{align}\qquad{(1)}\] as \(n\to\infty\). Then, \(\widehat{\xi}_n\) has the weak consistency: \[\widehat{\xi}_n\to^p\xi_0.\]

Remark 2. The sequences \(a_n\) and \(b_n\) describe different aspects of the drift structure. More precisely, \(a_n\) in B4 represents the quadratic separation scale of the drift family and determines the effective information accumulation for parameter identification. In contrast, \(b_n\) in B5 measures the accumulated local sensitivity of the drift family with respect to the parameter. In general, one need not have \(a_n\asymp b_n\); see Example 7.

The condition \(b_n=O(a_n)\) is a technical condition ensuring the uniform negligibility of the stochastic term \[\sup_{\xi\in\overline{\Xi}}\left|\frac{1}{a_nh_n}\sum_{i=1}^n d_i(\xi)\Delta_i^n Z\right|\stackrel{p}{\longrightarrow}0;\] see the proof of Theorem 1. In many standard examples, including directly Riemann integrable, polynomial-type, and periodic drift families described below, one typically has \(a_n\asymp b_n\).

Remark 3. Unlike classical drift estimation problems for ergodic diffusion processes, the present framework does not require ergodicity of the underlying Gaussian process itself. Indeed, the deterministic drift structure enters the local increment contrast through the non-random quantities \[\mu_\xi(t_{i-1})-\mu_{\xi_0}(t_{i-1}),\] so that the information accumulation mechanism is deterministic rather than path-dependent. Consequently, the consistency proof is based on the deterministic separation structure in B4 together with covariance control for weakly dependent Gaussian increments through A1.

We next establish asymptotic normality. To formulate the limiting covariance structure, define \[\begin{align} I_n := \frac{1}{a_n}\sum_{i=1}^n\partial_\xi\mu_{\xi_0}(t_{i-1})\partial_\xi\mu_{\xi_0}(t_{i-1})^\top. \label{information} \end{align}\tag{8}\] where \(a_n\) is the sequence given in B4, and we shall make further assumptions as follows.

B 6. There exists a positive definite matrix \(I(\xi_0)\) such that \[I_n\to I(\xi_0),\quad n\to \infty.\]

B 7. There exists a positive definite matrix \(\Gamma(\xi_0)\) such that \[\frac{1}{a_nh_n^\beta}\sum_{i,j=1}^n\partial_\xi\mu_{\xi_0}(t_{i-1})\partial_\xi\mu_{\xi_0}(t_{j-1})^\top\mathrm{Cov}(\Delta_i^nZ,\Delta_j^nZ)\to\Gamma(\xi_0),\quad n\to \infty.\]

B 8. There exists a neighborhood \(U\) of \(\xi_0\) and a measurable function \(N:[0,\infty)\to[0,\infty)\) such that \[\sup_{\xi\in U}|\partial_\xi^2\mu_\xi(t)|\le N(t), \quad t\ge0,\] and \[\frac{1}{a_n}\sum_{i=1}^n N(t_{i-1})^2=O(1),\quad n\to \infty.\]

Assumption B6 requires that the normalized quadratic sensitivity matrix 8 converges to the nondegenerate matrix \(I(\xi_0)\), which plays the role of a Fisher-type information matrix associated with the drift family. The quantities appearing in B7 and B8 are normalized by \(a_n\), reflecting the fact that both the asymptotic score variance and the Hessian are governed by the same global information accumulation scale.

The next theorem gives asymptotic normality of the estimator.

Theorem 2. Suppose Assumptions A1, A2, and B1–B8 hold. Let \(a_n\) and \(b_n\) be the sequences appearing in Assumptions B4 and B5, respectively, and assume that \(b_n=O(a_n)\). Furthermore, assume that \[\begin{align} a_n h_n^{2-(\alpha\wedge \beta)}\to\infty;\quad \frac{nh_n^2}{a_n}\to0;\quad nh_n^{4-\beta}\to0. \label{asymp46normal-rate} \end{align}\qquad{(2)}\] as \(n\to \infty\). Then, \[\sqrt{a_nh_n^{2-\beta}}(\widehat{\xi}_n-\xi_0)\stackrel{d}{\longrightarrow}{\cal N}\left(0,I(\xi_0)^{-1}\Gamma(\xi_0)I(\xi_0)^{-1}\right).\]

Remark 4. Rate conditions ?? are simplified when \(\alpha=\beta\) as follows: \[a_nh_n^{2-\alpha}\to\infty;\quad nh_n^2\to 0.\]

Remark 5. The normalization rate \(\sqrt{a_n h_n^{2-\beta}}\) shows explicitly how the rate of convergence depends both on the information accumulation structure of the drift family and on the dependence accumulation exponent \(\beta\) of the Gaussian noise.

Remark 6. An interesting feature of the present framework is that the weak consistency and asymptotic normality are governed by different properties of the Gaussian noise. More precisely, consistency depends on the increment dependence accumulation rate \(\alpha\) in  A1, whereas asymptotic normality is determined by the local roughness exponent \(\beta\) in A2 In many standard Gaussian process models, one typically has \(\alpha=\beta\). However, separating these quantities allows the framework to accommodate more general dependence structures.

3.2 Examples↩︎

We next illustrate how the asymptotic behavior of the estimator depends on both the local roughness of the Gaussian noise and the information accumulation structure of the drift family.

We first present several examples of covariance kernels satisfying Assumption A2, where the local behavior of the kernel determines the scaling exponent \(\beta\) and the local variance constant \(c_K\) appearing in the asymptotic theory. We then introduce examples of drift families satisfying the identifiability and information accumulation conditions. Finally, we combine these examples to verify the information-type assumptions appearing in the asymptotic normality theorem and to derive explicit forms of the asymptotic variance matrix \(\Gamma(\xi_0)\).

3.2.1 Examples of covariance kernels↩︎

We first present several examples of covariance kernels satisfying Assumption A2.

Example 1 (Gaussian kernel). Consider the Gaussian kernel \[K(t)=a\exp\left(-\frac{b}{2}t^2\right), \quad a,b>0,\] which satisfies A1 with \(\alpha=2\) due to 4 . Moreover, since \[K(0)-K(t)=a\left(1-\exp\left(-\frac{b}{2}t^2\right)\right)=\frac{ab}{2}t^2+o(t^2), \quad t\downarrow0,\] A2 holds with \[\alpha = \beta=2, \quad c_K=\frac{ab}{2}.\]

Example 2 (Matérn kernel). Consider the Matérn kernel \[K(t)=a\frac{2^{1-\nu}}{\Gamma(\nu)}\left(\sqrt{2\nu}\,b|t|\right)^\nu B_\nu\left(\sqrt{2\nu}\,b|t|\right),\] where \(a,b,\nu>0\) and \(B_\nu\) denotes the modified Bessel function of the second kind.

If \(\nu>1\), then \[K(0)-K(t)=ab^2\frac{\nu}{2\nu-1}t^2+o(t^2), \quad t\downarrow0,\] and hence, with 4 , \[\alpha = \beta=2,\quad c_K=ab^2\frac{\nu}{2\nu-1}.\]

If \(\nu\in(0,1)\), \[K(0)-K(t)\asymp |t|^{2\nu}, \quad t\downarrow0,\] so that \[\alpha = \beta=2\nu.\]

Example 3 (Rational quadratic kernel). Consider the rational quadratic kernel \[K(t)=a\left(1+\frac{b^2t^2}{2\gamma}\right)^{-\gamma}, \quad a,b,\gamma>0.\] Using Taylor expansion, \[K(0)-K(t)=\frac{ab^2}{2}t^2+o(t^2), \quad t\downarrow0.\] Hence, with 4 , \[\alpha = \beta=2, \quad c_K=\frac{ab^2}{2}.\]

Example 4 (Ornstein-Uhlenbeck kernel). Consider the exponential kernel \[K(t)=ae^{-b|t|}, \quad a,b>0.\] Note that \(\alpha=1\) from 5 , and that \[K(0)-K(t)=a(1-e^{-b|t|})=ab|t|+o(|t|), \quad t\downarrow0.\] Therefore, \[\alpha = \beta=1, \quad c_K=ab.\] This example illustrates that the present framework naturally includes non-smooth Gaussian kernels.

3.2.2 Examples of drift accumulation rates↩︎

We next consider several representative drift families satisfying the series of Assumptions B4 and B5.

Example 5 (Directly Riemann integrable drifts). We first consider the directly Riemann integrable (DRI) case, which corresponds to localized drift signals.

Definition 1. A measurable function \(f:[0,\infty)\to\mathbb{R}\) is said to be directly Riemann integrable (DRI) if \[\lim_{\delta\downarrow0}\delta\sum_{k=1}^\infty\sup_{t\in[(k-1)\delta,k\delta]}|f(t)|=\lim_{\delta\downarrow0}\delta\sum_{k=1}^\infty\inf_{t\in[(k-1)\delta,k\delta]}|f(t)|<\infty.\] See Feller [9] for general properties of directly Riemann integrable functions.

Suppose that \(t\mapsto\bigl(\mu_\xi(t)-\mu_{\xi_0}(t)\bigr)^2\) is directly Riemann integrable. Then, \[\sum_{i=1}^n\bigl(\mu_\xi(t_{i-1})-\mu_{\xi_0}(t_{i-1})\bigr)^2\sim h_n^{-1}\int_0^\infty\bigl(\mu_\xi(t)-\mu_{\xi_0}(t)\bigr)^2\,\mathrm{d}t,\] and B4 holds with \(a_n=h_n^{-1}\). Suppose moreover that \(t\mapsto\sup_{\xi\in\overline{\Xi}}|\partial_\xi\mu_\xi(t)|^2\) is directly Riemann integrable. Then, \[\sum_{i=1}^n\sup_{\xi\in\overline{\Xi}}|\partial_\xi\mu_\xi(t_{i-1})|^2=O(h_n^{-1}),\] and B5 holds with \(b_n=h_n^{-1}\). Therefore, \[b_n=O(a_n).\]

The rate of convergence becomes \[\sqrt{a_nh_n^{2-\beta}}=h_n^{(1-\beta)/2}.\] Typical examples include localized Gaussian drifts: \(\mu_\xi(t)=e^{-\xi t^2}\), exponentially decaying drifts: \(\mu_\xi(t)=e^{-\xi t}\), and compactly supported drifts.

In particular, for smooth Gaussian kernels satisfying \(\beta=2\), \[\widehat{\xi}_n-\xi_0=O_p(h_n^{1/2}),\] while for the Ornstein-Uhlenbeck kernel satisfying \(\beta=1\), and \(a_nh_n^{2-\beta}=1\), so that the consistency condition fails. Thus the present theorem does not guarantee consistency in the O-U case.

The information matrix and asymptotic covariance matrix depend on the covariance kernel and will be discussed in the combined examples below.

Remark 7. Directly Riemann integrable drifts correspond to localized deterministic signals whose magnitude gradually vanishes over time. Consequently, additional observations collected over a long horizon contribute diminishing amounts of statistical information for drift estimation. This explains why the resulting rate of convergence is substantially slower than in persistent signal models such as periodic drifts, described below. From an applied viewpoint, such drifts may be suitable for modeling transient trend phenomena exhibiting diminishing marginal effects over time.

Example 6 (Polynomial-type drifts). Consider a drift family satisfying \[|\mu_\xi(t)-\mu_{\xi_0}(t)|\asymp(1+t)^\alpha,\quad t\to\infty,\] for some \(\alpha\in\mathbb{R}\). Then, \[\sum_{i=1}^n\bigl(\mu_\xi(t_{i-1})-\mu_{\xi_0}(t_{i-1})\bigr)^2\asymp h_n^{-1}\int_0^{T_n}(1+t)^{2\alpha}\,\mathrm{d}t.\] If \(\alpha\neq-1/2\), it follows that \[a_n\asymp h_n^{-1}T_n^{2\alpha+1}.\] Suppose moreover that \[\sup_{\xi\in\overline{\Xi}}|\partial_\xi\mu_\xi(t)|^2\asymp(1+t)^{2\gamma},\quad t\to\infty,\] for some \(\gamma\in\mathbb{R}\). Then, \[\sum_{i=1}^n\sup_{\xi\in\overline{\Xi}}|\partial_\xi\mu_\xi(t_{i-1})|^2\asymp h_n^{-1}\int_0^{T_n}(1+t)^{2\gamma}\,\mathrm{d}t.\] If \(\gamma\neq-1/2\), then \(b_n\asymp h_n^{-1}T_n^{2\gamma+1}\). In particular, if \(\gamma\le\alpha\), then \[b_n=O(a_n).\]

The rate of convergence becomes \[\sqrt{a_nh_n^{2-\beta}}\asymp T_n^{\alpha+1/2}h_n^{(1-\beta)/2}.\]

In the critical case \(\alpha=-1/2\), \[a_n\asymp h_n^{-1}\log T_n,\] and therefore \[\sqrt{a_nh_n^{2-\beta}}\asymp(\log T_n)^{1/2}h_n^{(1-\beta)/2}.\] Typical examples include polynomial decay: \(\mu_\xi(t)=(1+t)^{-\xi}\), \(\xi>0\), slow logarithmic growth: \(\mu_\xi(t)=\log(1+\xi t)\), and polynomial growth: \(\mu_\xi(t)=(1+t)^\xi\), \(\xi>0\).

For example, if \[\mu_\xi(t)=\xi(1+t)^\alpha,\] then \[a_n\asymp b_n\asymp h_n^{-1}T_n^{2\alpha+1}.\]

If \(\alpha<-1/2\), the signal becomes asymptotically localized and the information accumulation is relatively slow. This regime includes directly Riemann integrable drifts.

If \(\alpha>-1/2\), the deterministic signal persists over long horizons, producing substantially faster rate of convergences.

In particular, polynomially growing drifts corresponding to \(\alpha>0\) generate accelerated information accumulation. For smooth Gaussian kernels satisfying \(\beta=2\), \[\widehat{\xi}_n-\xi_0=O_p(T_n^{-(\alpha+1/2)}h_n^{-1/2}),\] while for the Ornstein-Uhlenbeck kernel satisfying \(\beta=1\), \[\widehat{\xi}_n-\xi_0=O_p(T_n^{-(\alpha+1/2)}).\]

Example 7 (Periodic-type drifts). Consider a drift family satisfying \[|\mu_\xi(t)-\mu_{\xi_0}(t)|\asymp f_\xi(t),\] where \(f_\xi\) is a nontrivial bounded periodic function. Suppose moreover that \[\frac{1}{T}\int_0^Tf_\xi(t)^2\,\mathrm{d}t\to c(\xi,\xi_0)\in(0,\infty),\quad T\to\infty.\] Then, \[\sum_{i=1}^n\bigl(\mu_\xi(t_{i-1})-\mu_{\xi_0}(t_{i-1})\bigr)^2\asymp h_n^{-1}T_n=n.\] Hence, \(a_n\asymp n\). Suppose moreover that \[\sup_{\xi\in\overline{\Xi}}|\partial_\xi\mu_\xi(t)|\le M(t),\] where \(M\) is bounded and periodic. Then, \[\sum_{i=1}^n\sup_{\xi\in\overline{\Xi}}|\partial_\xi\mu_\xi(t_{i-1})|^2\asymp n,\] and \(b_n\asymp n\). Therefore, \[b_n=O(a_n).\] The rate of convergence becomes \[\sqrt{a_nh_n^{2-\beta}}=\sqrt{nh_n^{2-\beta}}.\] In particular, for smooth Gaussian kernels satisfying \(\beta=2\), \[\widehat{\xi}_n-\xi_0=O_p(n^{-1/2}),\] while for the Ornstein-Uhlenbeck kernel satisfying \(\beta=1\), \[\widehat{\xi}_n-\xi_0=O_p((nh_n)^{-1/2}).\]

For example, consider the periodic drift family \[\mu_\xi(t)=\xi_1\sin(\omega t)+\xi_2\cos(\omega t),\] where \(\xi=(\xi_1,\xi_2)\in\Xi\subset\mathbb{R}^2\) and \(\omega>0\) is known. Then, \[\mu_\xi(t)-\mu_{\xi_0}(t)=(\xi_1-\xi_{0,1})\sin(\omega t)+(\xi_2-\xi_{0,2})\cos(\omega t),\] and hence \[\sum_{i=1}^n\bigl(\mu_\xi(t_{i-1})-\mu_{\xi_0}(t_{i-1})\bigr)^2\asymp n.\] Moreover, \[\partial_\xi\mu_\xi(t)=\bigl(\sin(\omega t),\cos(\omega t)\bigr),\] so that \[\sum_{i=1}^n\|\partial_\xi\mu_\xi(t_{i-1})\|^2\asymp n.\] Therefore, \[a_n\asymp b_n\asymp n.\]

However, estimation of frequency can be excluded from our assumptions. Consider \[\mu_\xi(t)=\sin(\xi t),\quad \xi\in\Xi\subset(0,\infty).\] Then, for \(\xi\neq\xi_0\), \[\frac{1}{T}\int_0^T\bigl(\sin(\xi t)-\sin(\xi_0 t)\bigr)^2\,\mathrm{d}t\to1,\quad T\to\infty.\] Thus the above condition holds with \(c(\xi,\xi_0)=1\). Moreover, \[\sum_{i=1}^n\sup_{\xi\in\overline{\Xi}}|\partial_\xi\mu_\xi(t_{i-1})|^2\asymp h_n^{-1}T_n^3.\] Thus B5 holds with \(b_n\asymp h_n^{-1}T_n^3\). In particular, \[b_n\neq O(a_n),\] since \(a_n\asymp n=h_n^{-1}T_n.\) Therefore, although the separation condition holds, the consistency theorem is not directly applicable in this case. The same phenomenon appears for other trigonometric drift families such as \(\mu_\xi(t)=\cos(\xi t)\), and \(\mu_\xi(t)=A\sin(\xi t+\phi)\), when the frequency parameter is unknown.

Unlike directly Riemann integrable drifts, periodic-type drifts continue generating information over arbitrarily long observation horizons. As a result, the information accumulation is substantially faster than in localized signal models. However, it does not necessarily mean that the consistency itself fails, and we may require a different local analysis or a refined parametrization.

3.2.3 Combined examples for asymptotic variance↩︎

We finally combine the covariance kernel examples and the drift accumulation structures introduced above, and verify the assumptions appearing in the asymptotic normality theorem.

The key point is that \(a_n\) is determined by the separation order of the drift family, while \(b_n\) is determined by the growth order of the derivative field \(t\mapsto\partial_\xi\mu_{\xi_0}(t)\).

Under the local kernel condition, \[\mathrm{Cov}(\Delta_i^nZ,\Delta_i^nZ)=2\{K(0)-K(h_n)\}\sim2c_Kh_n^\beta.\]

If the off-diagonal covariance terms are of smaller order or asymptotically absorbed into the same limit, then the covariance sum appearing in Assumption B7 is governed by the order \[h_n^\beta\sum_{i=1}^n\partial_\xi\mu_{\xi_0}(t_{i-1})\partial_\xi\mu_{\xi_0}(t_{i-1})^\top.\]

Thus, once Assumption B5 gives \[\sum_{i=1}^n\partial_\xi\mu_{\xi_0}(t_{i-1})\partial_\xi\mu_{\xi_0}(t_{i-1})^\top=O(b_n),\] the normalization in Assumption B7 is naturally determined by the product of the derivative-growth order and the local increment variance order.

Example 8 (Polynomial drift with Gaussian and Ornstein-Uhlenbeck kernels). Consider the polynomial drift \[\mu_\xi(t)=\xi(1+t)^m,\quad \xi\in\Xi\subset\mathbb{R},\] where \(m\in\mathbb{R}\). By Example 6,

  • If \(m>-1/2\): \(a_n=b_n=h_n^{-1}T_n^{2m+1}\), and \[I(\xi_0)=\lim_{n\to\infty}\frac{1}{a_n}\sum_{i=1}^n(1+t_{i-1})^{2m}=\frac{1}{2m+1}.\]

  • If \(m=-1/2\): \(a_n=b_n=h_n^{-1}\log T_n\), and \[I(\xi_0)=1.\]

  • If \(m<-1/2\) (DRI case): \(a_n=b_n=h_n^{-1}\), and \[I(\xi_0)=\int_0^\infty(1+t)^{2m}\,\mathrm{d}t=\frac{1}{-2m-1}.\]

For the Gaussian kernel in Example 1, we have \(\beta=2\) and \(c_K=ab/2\). Hence, \[\Gamma(\xi_0)=abI(\xi_0),\] and

  • If \(m\ge-1/2\): \[T_n^{m+1/2}h_n^{1/2}(\widehat{\xi}_n-\xi_0)\stackrel{d}{\longrightarrow}\mathcal{N}\left(0,(abI(\xi_0))^{-1}\right),\]

  • If \(m<-1/2\) (DRI case): \[h_n^{-1/2}(\widehat{\xi}_n-\xi_0)\stackrel{d}{\longrightarrow}\mathcal{N}\left(0,(abI(\xi_0))^{-1}\right).\]

For the Ornstein-Uhlenbeck kernel in Example 4, we have \(\beta=1\) and \(c_K=ab\). Hence, \[\Gamma(\xi_0)=2abI(\xi_0).\]

  • If \(m>-1/2\): \[T_n^{m+1/2}(\widehat{\xi}_n-\xi_0)\stackrel{d}{\longrightarrow}\mathcal{N}\left(0,(2abI(\xi_0))^{-1}\right).\]

  • If \(m=-1/2\): \[(\log T_n)^{1/2}(\widehat{\xi}_n-\xi_0)\stackrel{d}{\longrightarrow}\mathcal{N}\left(0,(2abI(\xi_0))^{-1}\right).\]

  • If \(m<-1/2\) (DRI case): \(a_nh_n^{2-\beta}=1\), so that the consistency condition fails.

Example 9 (Periodic drift with Gaussian and Ornstein-Uhlenbeck kernels). Consider the periodic drift \[\mu_\xi(t)=\xi_1\sin(\omega t)+\xi_2\cos(\omega t),\] where \(\xi=(\xi_1,\xi_2)\in\Xi\subset\mathbb{R}^2\) and \(\omega>0\) is known. By Example 7, \(a_n=b_n=n.\) Moreover, \[I(\xi_0):=\lim_{n\to\infty}\frac{1}{n}\sum_{i=1}^n\partial_\xi\mu_\xi(t_{i-1})\partial_\xi\mu_\xi(t_{i-1})^\top = \frac{1}{2} \begin{pmatrix} 1 & 0 \\ 0 & 1 \end{pmatrix}.\]

For the Gaussian kernel in Example 1, we have \(\beta=2\) and \(c_K=ab/2\). Hence, \[\Gamma(\xi_0)=ab\,I(\xi_0),\] and \[\sqrt{n}(\widehat{\xi}_n-\xi_0)\stackrel{d}{\longrightarrow}\mathcal{N}\left(0,(ab\,I(\xi_0))^{-1}\right).\]

For the Ornstein-Uhlenbeck kernel in Example 4, we have \(\beta=1\) and \(c_K=ab\). Hence, \[\Gamma(\xi_0)=2ab\,I(\xi_0),\] and \[\sqrt{nh_n}(\widehat{\xi}_n-\xi_0)\stackrel{d}{\longrightarrow}\mathcal{N}\left(0,(2ab\,I(\xi_0))^{-1}\right).\]

In particular, both kernels yield consistency and asymptotic normality since \[a_nh_n^{2-\beta}\to\infty.\]

These examples show that \(b_n\) is not imposed independently of the drift family. Rather, \(b_n\) is determined by the growth order of the derivative field in Assumption B5. Once \(b_n\) is identified, the score covariance order is given by \(b_nh_n^\beta\), where \(\beta\) is determined by the local behavior of the covariance kernel. For example, \(b_nh_n^2\) arises for the Gaussian kernel, whereas \(b_nh_n\) arises for the Ornstein–Uhlenbeck kernel.

4 Cncluding remarks↩︎

In this paper, we studied statistical inference for deterministic drift structures in Gaussian process models under high-frequency observations. The proposed methodology is based on local increment information and a least squares-type contrast constructed from adjacent increments. Unlike many existing approaches in Gaussian process statistics, the present framework treats the deterministic drift itself as the primary inferential target rather than a nuisance component to be removed by empirical centering.

A central feature of the theory is that the asymptotic behavior of the estimator is governed jointly by the information accumulation structure of the drift and the small-time covariance structure of the Gaussian component. This viewpoint naturally yields different statistical regimes according to the long-time behavior of the deterministic drift. In particular, directly Riemann integrable drifts correspond to relatively weak information accumulation, whereas periodic or growing drifts may produce substantially faster rate of convergences.

The framework developed here is intentionally formulated under relatively general assumptions. In particular, directly Riemann integrable drifts appear only as one representative example rather than a fundamental assumption of the theory. This allows the present approach to accommodate a broad class of nonstationary deterministic structures beyond localized transient trends.

4.0.0.1 Future works:

Several possible extensions of the present framework remain for future investigation.

First, it would be interesting to incorporate trajectory fitting-type estimators based on deterministic approximations of stochastic dynamics; see Kasonga [10]. Instead of approximating the local drift contribution by the first-order Euler expansion, \[\int_{t_{i-1}}^{t_i} \mu_\xi(s)\,\mathrm{d}s \approx \mu_\xi(t_{i-1})h_n,\] one may fit the observed trajectory to the solution of a deterministic differential equation generated by the drift itself. Such an approach may reduce discretization bias and improve finite-sample performance, especially when the drift function is sufficiently smooth. However, since the asymptotic fluctuation in the present setting is mainly governed by the stochastic variability of the Gaussian increments, trajectory fitting methods are not expected to alter the asymptotic rate of convergence itself. A rigorous comparison between local contrast methods and trajectory fitting procedures will be studied elsewhere.

Second, although the present paper focuses primarily on drift estimation, extension to simultaneous inference for both drift and covariance parameters remains an important open problem. In particular, after obtaining a consistent estimator of the drift parameter, one may construct residual-based estimators for covariance or kernel parameters through residual processes of the form \[\widehat{Z}_t = X_t - \int_0^t \mu_{\widehat{\xi}_n}(s)\,\mathrm{d}s.\] Such procedures would naturally involve ergodic properties and long-time averaging behavior of the residual process. The asymptotic interaction between the first-stage drift estimation error and the second-stage kernel estimation remains to be clarified. A systematic two-step asymptotic theory in this direction would provide a more unified framework for simultaneous inference on deterministic and stochastic structures.

Third, although the present paper considers stationary Gaussian processes, the local contrast construction itself does not fundamentally rely on Gaussianity. It is therefore natural to investigate extensions to more general ergodic stochastic processes, including ergodic diffusion processes and related semimartingale models. In such settings, local increment structures may still provide tractable estimating equations without requiring global likelihood evaluation. The corresponding asymptotic theory would likely require mixing properties, ergodic theorems, and martingale approximation techniques adapted to non-Gaussian dynamics. We hope that the present work provides a useful starting point for further investigation of drift inference under dependent stochastic environments.

5 Proofs↩︎

5.1 Proof of Theorem 1↩︎

Define the normalized contrast difference \[\begin{align} \overline{H}_n(\xi):=\frac{1}{a_nh_n^2}\bigl(H_n(\xi)-H_n(\xi_0)\bigr). \label{normalized-contrast} \end{align}\tag{9}\] Since the normalization factor does not depend on \(\xi\), the estimator \(\widehat{\xi}_n\) is equivalently characterized as a minimizer of \(\overline{H}_n(\xi)\). Set \[\begin{align} d_i(\xi)&:=\mu_\xi(t_{i-1})-\mu_{\xi_0}(t_{i-1}); \quad r_i^n:=\int_{t_{i-1}}^{t_i}\{\mu_{\xi_0}(s)-\mu_{\xi_0}(t_{i-1})\}\,\mathrm{d}s, \end{align}\] so that \[\Delta_i^nX-h_n\mu_\xi(t_{i-1})=\Delta_i^nZ+r_i^n-h_nd_i(\xi).\] Hence, \[\begin{align} H_n(\xi)-H_n(\xi_0) &=\sum_{i=1}^n\Bigl\{(\Delta_i^nZ+r_i^n-h_nd_i(\xi))^2-(\Delta_i^nZ+r_i^n)^2\Bigr\} \\ &=h_n^2\sum_{i=1}^nd_i(\xi)^2-2h_n\sum_{i=1}^nd_i(\xi)\Delta_i^nZ-2h_n\sum_{i=1}^nd_i(\xi)r_i^n. \end{align}\] Therefore, \[\begin{align} \overline{H}_n(\xi) &=\frac{1}{a_n}\sum_{i=1}^nd_i(\xi)^2-\frac{2}{a_nh_n}\sum_{i=1}^nd_i(\xi)\Delta_i^nZ-\frac{2}{a_nh_n}\sum_{i=1}^nd_i(\xi)r_i^n \\ &=: J_{1,n}(\xi) + J_{2,n}(\xi) + J_{3,n}(\xi). \end{align}\]

As for \(J_{1,n}\): Assymption B4 yields \[\sup_{\xi\in \Xi}\left|J_{1,n}(\xi) -Q(\xi,\xi_0)\right|\to0.\]

As for \(J_{2,n}\): Assumptions A1 and B4 imply that, for each fixed \(\xi\in\Xi\), \[\begin{align} \mathrm{Var}\left(\sum_{i=1}^nd_i(\xi)\Delta_i^nZ\right) &\le\sum_{i,j=1}^n |d_i(\xi)d_j(\xi)|\,|\mathrm{Cov}(\Delta_i^nZ,\Delta_j^nZ)| \\ &\le Ch_n^\alpha\sum_{i=1}^nd_i(\xi)^2=O(a_nh_n^\alpha). \end{align}\] Hence, \[J_{2,n}(\xi)=O_P\left((a_nh_n^{2-\alpha})^{-1/2}\right)\stackrel{p}{\longrightarrow}0.\] We next prove stochastic equicontinuity. By the mean value theorem and B5, \[\begin{align} \mathrm{Var}\left(\sum_{i=1}^n\{d_i(\xi)-d_i(\eta)\}\Delta_i^nZ\right) &\le Ch_n^\alpha |\xi-\eta|^2\sum_{i=1}^nM(t_{i-1})^2 =O(h_n^\alpha b_n|\xi-\eta|^2). \end{align}\] Since \(b_n=O(a_n)\), \[J_{2,n}(\xi)-J_{2,n}(\eta)=O_P\left(|\xi-\eta|(a_nh_n^{2-\alpha})^{-1/2}\right).\] Let \(\varepsilon>0\) be arbitrary. Since \(\overline{\Xi}\) is compact, there exist finitely many points \(\xi^{(1)},\dots,\xi^{(N)}\in\Xi\) such that \[\Xi\subset\bigcup_{k=1}^N B(\xi^{(k)},\varepsilon).\] Hence, \[\sup_{\xi\in\overline{\Xi}}|J_{2,n}(\xi)| \le\max_{1\le k\le N}|S_n(\xi^{(k)})|+\max_{1\le k\le N}\sup_{\xi\in B(\xi^{(k)},\varepsilon)}|J_{2,n}(\xi)-J_{2,n}(\xi^{(k)})|.\] Since \(J_{2,n}(\xi^{(k)})\stackrel{p}{\longrightarrow}0\) for each \(k=1,\dots,N\), we obtain \(\max_{1\le k\le N}|J_{2,n}(\xi^{(k)})|\stackrel{p}{\longrightarrow}0\). Moreover, \[\sup_{\xi\in B(\xi^{(k)},\varepsilon)}|J_{2,n}(\xi)-J_{2,n}(\xi^{(k)})|=O_P\left(\varepsilon(a_nh_n^{2-\alpha})^{-1/2}\right)=o_P(1).\] Therefore, \[\sup_{\xi\in\overline{\Xi}}|J_{2,n}(\xi)|\stackrel{p}{\longrightarrow}0.\] As for \(J_{3,n}\): Noticing by B3 that \[\begin{align} |r_i^n|\le Ch_n^2. \label{r95i94n}, \end{align}\tag{10}\] we have \[\begin{align} \sup_{\xi\in\overline{\Xi}}\left|\frac{1}{a_nh_n}\sum_{i=1}^nd_i(\xi)r_i^n\right| &\le\frac{Ch_n}{a_n}\sup_{\xi\in\overline{\Xi}}\sum_{i=1}^n|d_i(\xi)| \le\frac{Ch_n n^{1/2}}{a_n}\sup_{\xi\in\overline{\Xi}}\left(\sum_{i=1}^nd_i(\xi)^2\right)^{1/2} \\ &\le C\left(\frac{nh_n^2}{a_n}\right)^{1/2}\to0. \end{align}\] Consequently, \[\sup_{\xi\in\overline{\Xi}}|\overline{H}_n(\xi)-Q(\xi,\xi_0)|\stackrel{p}{\longrightarrow}0.\] Since \(Q\) is continuous, \(Q(\xi,\xi_0)\ge0\), and \[Q(\xi,\xi_0)=0 \iff \xi=\xi_0,\] the standard M-estimation theory, e.g., van der Vaart [11], Theorem 5.7, yields \[\widehat{\xi}_n\stackrel{p}{\longrightarrow}\xi_0.\] This completes the proof.

5.2 Proof of Theorem 2↩︎

Define \[\Psi_n(\xi):=\partial_\xi H_n(\xi);\qquad \mathcal{J}_n:=\frac{1}{2a_nh_n^2}\int_0^1\partial_\xi\Psi_n\{\xi_0+u(\widehat{\xi}_n-\xi_0)\}\,\mathrm{d}u,\] and set \[A_n':=\{\widehat{\xi}_n\in\Xi\},\quad A_n'':=\{|\det\mathcal{J}_n|>\tfrac12\det I(\xi_0)\},\quad A_n:=A_n'\cap A_n''.\] Since \(\xi_0\in\Xi\) and \(\widehat{\xi}_n\stackrel{p}{\longrightarrow}\xi_0\), we have \(P(A_n')\to1\). By Lemma 1, we have \(\mathcal{J}_n\stackrel{p}{\longrightarrow}I(\xi_0)\), so it follows that \(P(A_n'')\to1\) since \(I(\xi_0)\) is positive definite. Consequently, \[P(A_n)\to1.\] On \(A_n\), the first-order condition gives \[\Psi_n(\widehat{\xi}_n)=0.\] By the integral form of Taylor’s formula, \[0=\Psi_n(\xi_0)+\left\{\int_0^1\partial_\xi\Psi_n\{\xi_0+u(\widehat{\xi}_n-\xi_0)\}\,\mathrm{d}u\right\}(\widehat{\xi}_n-\xi_0)\quad \text{on }A_n.\] Therefore, \[\widehat{\xi}_n-\xi_0=-\mathcal{J}_n^{-1}\frac{1}{2a_nh_n^2}\Psi_n(\xi_0)\quad \text{on }A_n.\] Let \[r_n:=\sqrt{a_nh_n^{2-\beta}}.\] Then \[r_n(\widehat{\xi}_n-\xi_0)=-\mathcal{J}_n^{-1}r_n\frac{1}{2a_nh_n^2}\Psi_n(\xi_0)1_{A_n}+r_n(\widehat{\xi}_n-\xi_0)1_{A_n^c}.\] For every \(\varepsilon>0\), \[P\left(\left|r_n(\widehat{\xi}_n-\xi_0)1_{A_n^c}\right|>\varepsilon\right)\le P(A_n^c)\to0.\] Hence, \[r_n(\widehat{\xi}_n-\xi_0)=-\mathcal{J}_n^{-1}r_n\frac{1}{2a_nh_n^2}\Psi_n(\xi_0)1_{A_n}+o_P(1).\] We next identify the limit of the normalized score. Since \[\Psi_n(\xi_0)=-2h_n\sum_{i=1}^n\{\Delta_i^nX-h_n\mu_{\xi_0}(t_{i-1})\}\partial_\xi\mu_{\xi_0}(t_{i-1}),\] and \[\Delta_i^nX-h_n\mu_{\xi_0}(t_{i-1})=\Delta_i^nZ+r_i^n,\qquad r_i^n:=\int_{t_{i-1}}^{t_i}\{\mu_{\xi_0}(s)-\mu_{\xi_0}(t_{i-1})\}\,\mathrm{d}s,\] we have \[r_n\frac{1}{2a_nh_n^2}\Psi_n(\xi_0)=-\frac{1}{\sqrt{a_nh_n^\beta}}\sum_{i=1}^n\partial_\xi\mu_{\xi_0}(t_{i-1})\Delta_i^nZ-\frac{1}{\sqrt{a_nh_n^\beta}}\sum_{i=1}^n\partial_\xi\mu_{\xi_0}(t_{i-1})r_i^n.\] Since \[\sum_{i=1}^n|\partial_\xi\mu_{\xi_0}(t_{i-1})|\le n^{1/2}\left(\sum_{i=1}^nM(t_{i-1})^2\right)^{1/2}=O((nb_n)^{1/2}),\] by B5, we see with 10 that \[\left|\frac{1}{\sqrt{a_nh_n^\beta}}\sum_{i=1}^n\partial_\xi\mu_{\xi_0}(t_{i-1})r_i^n\right|\le Ch_n^{2-\beta/2}\left(\frac{nb_n}{a_n}\right)^{1/2}.\] Since \(b_n=O(a_n)\) and \(nh_n^{4-\beta}\to0\), we get \[\frac{1}{\sqrt{a_nh_n^\beta}}\sum_{i=1}^n\partial_\xi\mu_{\xi_0}(t_{i-1})r_i^n\to0.\] Therefore, \[r_n\frac{1}{2a_nh_n^2}\Psi_n(\xi_0)=-\frac{1}{\sqrt{a_nh_n^\beta}}\sum_{i=1}^n\partial_\xi\mu_{\xi_0}(t_{i-1})\Delta_i^nZ+o_P(1).\] The leading term is a centered Gaussian random vector. Moreover, by B7, \[\mathrm{Var}\left(\frac{1}{\sqrt{a_nh_n^\beta}}\sum_{i=1}^n\partial_\xi\mu_{\xi_0}(t_{i-1})\Delta_i^nZ\right)\to\Gamma(\xi_0).\] Hence, \[-\frac{1}{\sqrt{a_nh_n^\beta}}\sum_{i=1}^n\partial_\xi\mu_{\xi_0}(t_{i-1})\Delta_i^nZ\stackrel{d}{\longrightarrow}{\cal N}(0,\Gamma(\xi_0)).\] Consequently, \[r_n\frac{1}{2a_nh_n^2}\Psi_n(\xi_0)\stackrel{d}{\longrightarrow}{\cal N}(0,\Gamma(\xi_0)).\] Since \(\mathcal{J}_n\stackrel{p}{\longrightarrow}I(\xi_0)\) and \(P(A_n)\to1\), we have \[\mathcal{J}_n^{-1}1_{A_n}\stackrel{p}{\longrightarrow}I(\xi_0)^{-1}.\] By Slutsky’s theorem, \[r_n(\widehat{\xi}_n-\xi_0)\stackrel{d}{\longrightarrow}{\cal N}\left(0,I(\xi_0)^{-1}\Gamma(\xi_0)I(\xi_0)^{-1}\right).\] This proves the desired asymptotic normality.

5.3 Auxiliary lemmas↩︎

Lemma 1. Suppose the assumptions of Theorem 2. Then \[\frac{1}{2a_nh_n^2}\int_0^1 \partial_\xi\Psi_n\{\xi_0+u(\widehat{\xi}_n-\xi_0)\}\,\mathrm{d}u\stackrel{p}{\longrightarrow}I(\xi_0).\]

Proof. Put \(\xi_n(u):=\xi_0+u(\widehat{\xi}_n-\xi_0)\) for \(0\le u\le1\). Let \(U\) be the neighborhood of \(\xi_0\) in Assumption B8. Choose \(\eta>0\) such that \(\{\xi\in\mathbb{R}^p:|\xi-\xi_0|<\eta\}\subset U\). Define \[C_n:=\{\xi_n(u)\in U \text{ for all } u\in[0,1]\}.\] Since \(\sup_{0\le u\le1}|\xi_n(u)-\xi_0|\le|\widehat{\xi}_n-\xi_0|\) and \(\widehat{\xi}_n\stackrel{p}{\longrightarrow}\xi_0\), we have \(P(C_n)\to1\). Recall that \[H_n(\xi)=\sum_{i=1}^n\{\Delta_i^nX-h_n\mu_\xi(t_{i-1})\}^2.\] Then \[\Psi_n(\xi)=-2h_n\sum_{i=1}^n\{\Delta_i^nX-h_n\mu_\xi(t_{i-1})\}\partial_\xi\mu_\xi(t_{i-1}).\] Hence \[\partial_\xi\Psi_n(\xi)=2h_n^2\sum_{i=1}^n\partial_\xi\mu_\xi(t_{i-1})\partial_\xi\mu_\xi(t_{i-1})^\top-2h_n\sum_{i=1}^n\{\Delta_i^nX-h_n\mu_\xi(t_{i-1})\}\partial_\xi^2\mu_\xi(t_{i-1}).\] Therefore, \[\frac{1}{2a_nh_n^2}\int_0^1\partial_\xi\Psi_n(\xi_n(u))\,\mathrm{d}u=B_n+R_n,\] where \[B_n:=\frac{1}{a_n}\sum_{i=1}^n\int_0^1\partial_\xi\mu_{\xi_n(u)}(t_{i-1})\partial_\xi\mu_{\xi_n(u)}(t_{i-1})^\top\,\mathrm{d}u\] and \[R_n:=-\frac{1}{a_nh_n}\sum_{i=1}^n\int_0^1\{\Delta_i^nX-h_n\mu_{\xi_n(u)}(t_{i-1})\}\partial_\xi^2\mu_{\xi_n(u)}(t_{i-1})\,\mathrm{d}u.\] First, we show that \(B_n\stackrel{p}{\longrightarrow}I(\xi_0)\). On \(C_n\), by the mean value theorem, Assumption B5, and Assumption B8, we have \[\left|B_n-\frac{1}{a_n}\sum_{i=1}^n\partial_\xi\mu_{\xi_0}(t_{i-1})\partial_\xi\mu_{\xi_0}(t_{i-1})^\top\right|\le C|\widehat{\xi}_n-\xi_0|\frac{1}{a_n}\sum_{i=1}^nM(t_{i-1})N(t_{i-1}).\] By Cauchy-Schwarz inequality, \[\frac{1}{a_n}\sum_{i=1}^nM(t_{i-1})N(t_{i-1})\le\left\{\frac{1}{a_n}\sum_{i=1}^nM(t_{i-1})^2\right\}^{1/2}\left\{\frac{1}{a_n}\sum_{i=1}^nN(t_{i-1})^2\right\}^{1/2}=O(1).\] Since \(\widehat{\xi}_n\stackrel{p}{\longrightarrow}\xi_0\), it follows that \[\left|B_n-\frac{1}{a_n}\sum_{i=1}^n\partial_\xi\mu_{\xi_0}(t_{i-1})\partial_\xi\mu_{\xi_0}(t_{i-1})^\top\right|1_{C_n}\stackrel{p}{\longrightarrow}0.\] By the definition of \(I_n\) and Assumption B6, \[\frac{1}{a_n}\sum_{i=1}^n\partial_\xi\mu_{\xi_0}(t_{i-1})\partial_\xi\mu_{\xi_0}(t_{i-1})^\top=I_n\to I(\xi_0).\] Thus \(B_n1_{C_n}\stackrel{p}{\longrightarrow}I(\xi_0)\). Next, we show that \(R_n=o_P(1)\) on \(C_n\). Write \[\Delta_i^nX-h_n\mu_{\xi_n(u)}(t_{i-1})=\Delta_i^nZ+\left\{\int_{t_{i-1}}^{t_i}\mu_{\xi_0}(s)\,\mathrm{d}s-h_n\mu_{\xi_0}(t_{i-1})\right\}+h_n\{\mu_{\xi_0}(t_{i-1})-\mu_{\xi_n(u)}(t_{i-1})\}.\] Correspondingly, write \(R_n=R_{n,1}+R_{n,2}+R_{n,3}\). For the Gaussian part, on \(C_n\) we have \[\mathbb{E}[|R_{n,1}1_{C_n}|^2]\le\frac{C}{a_n^2h_n^2}\sum_{i,j=1}^nN(t_{i-1})N(t_{j-1})|\mathrm{Cov}(\Delta_i^nZ,\Delta_j^nZ)|.\] By Cauchy-Schwarz inequality and Assumption A1, \[\mathbb{E}[|R_{n,1}1_{C_n}|^2]\le C\frac{h_n^{\beta-2}}{a_n}\left\{\frac{1}{a_n}\sum_{i=1}^nN(t_{i-1})^2\right\}.\] Since \(a_nh_n^{2-\beta}\to\infty\), we obtain \(R_{n,1}1_{C_n}\stackrel{p}{\longrightarrow}0\). For the time-discretization part, Assumption B3 gives \[\left|\int_{t_{i-1}}^{t_i}\mu_{\xi_0}(s)\,\mathrm{d}s-h_n\mu_{\xi_0}(t_{i-1})\right|\le Ch_n^2.\] Hence, on \(C_n\), \[|R_{n,2}|\le Ch_n\frac{1}{a_n}\sum_{i=1}^nN(t_{i-1})\le Ch_n\left\{\frac{1}{a_n}\sum_{i=1}^nN(t_{i-1})^2\right\}^{1/2}.\] Thus \(R_{n,2}1_{C_n}\stackrel{p}{\longrightarrow}0\). For the parameter-shift part, the mean value theorem gives \[\sup_{0\le u\le1}|\mu_{\xi_n(u)}(t_{i-1})-\mu_{\xi_0}(t_{i-1})|\le M(t_{i-1})|\widehat{\xi}_n-\xi_0|.\] Therefore, on \(C_n\), \[|R_{n,3}|\le C|\widehat{\xi}_n-\xi_0|\frac{1}{a_n}\sum_{i=1}^nM(t_{i-1})N(t_{i-1}).\] By Cauchy-Schwarz inequality and the growth bounds for \(M\) and \(N\), the last factor is \(O(1)\). Since \(\widehat{\xi}_n\stackrel{p}{\longrightarrow}\xi_0\), we obtain \(R_{n,3}1_{C_n}\stackrel{p}{\longrightarrow}0\). Consequently, \(R_n1_{C_n}=o_P(1)\). Combining the convergence of \(B_n\) and \(R_n\), we obtain \[\left\{\frac{1}{2a_nh_n^2}\int_0^1\partial_\xi\Psi_n\{\xi_0+u(\widehat{\xi}_n-\xi_0)\}\,\mathrm{d}u-I(\xi_0)\right\}1_{C_n}\stackrel{p}{\longrightarrow}0.\] Since \(P(C_n)\to1\), we conclude that \[\frac{1}{2a_nh_n^2}\int_0^1\partial_\xi\Psi_n\{\xi_0+u(\widehat{\xi}_n-\xi_0)\}\,\mathrm{d}u\stackrel{p}{\longrightarrow}I(\xi_0).\] ◻


This work is partially supported by JSPS KAKENHI Grant-in-Aid for Scientific Research (B) #24K02907; (C) #24K06875; Japan Science and Technology Agency CREST #JPMJCR2115.

6 Limit theorems for stationary Gaussian processes↩︎

Lemma 2. Let \(Z=(Z_t)_{t\ge0}\sim GP(0,K)\) be centered and stationary. Assume A2 and A1. Let \(m:[0,\infty)\to\mathbb{R}\) be measurable and suppose that \[\frac{1}{a_n}\sum_{i=1}^n m(t_{i-1})^2=O(1).\] Then \[\sum_{i=1}^n m(t_{i-1})\Delta_i^nZ=O_P\left((a_nh_n^\beta)^{1/2}\right).\]

Proof. Set \(S_n(m):=\sum_{i=1}^n m(t_{i-1})\Delta_i^nZ\). Since \(Z\) is centered, we have \(\mathbb{E}[S_n(m)]=0\). Moreover, \[\mathrm{Var}(S_n(m))=\sum_{i,j=1}^n m(t_{i-1})m(t_{j-1})\mathrm{Cov}(\Delta_i^nZ,\Delta_j^nZ).\] Using the inequaity \(2|xy|\le x^2+y^2\), we obtain \[\begin{align} |\mathrm{Var}(S_n(m))|&\le \frac{1}{2}\sum_{i,j=1}^n \{m(t_{i-1})^2+m(t_{j-1})^2\}\left|\mathrm{Cov}(\Delta_i^nZ,\Delta_j^nZ)\right| \\ &\le \sum_{i=1}^n m(t_{i-1})^2\sum_{j=1}^n \left|\mathrm{Cov}(\Delta_i^nZ,\Delta_j^nZ)\right|. \end{align}\] By A1, \[|\mathrm{Var}(S_n(m))|\le C h_n^\beta\sum_{i=1}^n m(t_{i-1})^2.\] Since \(a_n^{-1}\sum_{i=1}^n m(t_{i-1})^2=O(1)\), we get \(\mathrm{Var}(S_n(m))=O(a_nh_n^\beta)\). Therefore, \(S_n(m)=O_P((a_nh_n^\beta)^{1/2})\). ◻

Lemma 3. Let \(Z=(Z_t)_{t\ge0}\sim GP(0,K)\) be centered and stationary. Assume A[as:mixing] and A2. Let \(f:\mathbb{R}\times\overline{\Theta}\to\mathbb{R}\) be continuous and suppose that \(f\) is continuously differentiable in \((x,\theta)\). Assume that there exist constants \(C>0\) and \(p\ge1\) such that \[\sup_{\theta\in\overline{\Theta}}\{|f(x,\theta)|+|\partial_xf(x,\theta)|+|\partial_\theta f(x,\theta)|\}\le C(1+|x|^p), \quad x\in\mathbb{R}.\] Then, under \(h_n\to0\) and \(nh_n\to\infty\), \[\sup_{\theta\in\overline{\Theta}}\left|\frac{1}{n}\sum_{i=1}^n f(Z_{t_{i-1}},\theta)-\int_{\mathbb{R}} f(z,\theta)\phi_{K(0)}(z)\,\mathrm{d}z\right|\stackrel{p}{\longrightarrow}0.\]

Proof. Set \(T_n:=nh_n\). First, we compare the discrete average with the continuous-time average. By the mean value theorem and the growth condition on \(\partial_x f\), \[\mathbb{E}\left|\frac{1}{n}\sum_{i=1}^n f(Z_{t_{i-1}},\theta)-\frac{1}{T_n}\int_0^{T_n}f(Z_t,\theta)\,\mathrm{d}t\right|\le \frac{1}{T_n}\sum_{i=1}^n\int_{t_{i-1}}^{t_i}\mathbb{E}\left[|Z_t-Z_{t_{i-1}}|\int_0^1 |\partial_x f(Z_{t_{i-1}}+u(Z_t-Z_{t_{i-1}}),\theta)|\,\mathrm{d}u\right]\mathrm{d}t.\] By Cauchy-Schwarz inequality and the polynomial growth condition, the right-hand side is bounded by \[C\frac{1}{T_n}\sum_{i=1}^n\int_{t_{i-1}}^{t_i}\{\mathbb{E}|Z_t-Z_{t_{i-1}}|^2\}^{1/2}\,\mathrm{d}t.\] Since \[\mathbb{E}|Z_t-Z_{t_{i-1}}|^2=2\{K(0)-K(t-t_{i-1})\},\] A2 gives \[\sup_{0\le r\le h_n}\mathbb{E}|Z_{s+r}-Z_s|^2=O(h_n^\beta).\] Therefore, \[\mathbb{E}\left|\frac{1}{n}\sum_{i=1}^n f(Z_{t_{i-1}},\theta)-\frac{1}{T_n}\int_0^{T_n}f(Z_t,\theta)\,\mathrm{d}t\right|=O(h_n^{\beta/2})\to0.\] By A[as:mixing], \(Z\) is ergodic, and hence Birkhoff’s ergodic theorem yields \[\frac{1}{T_n}\int_0^{T_n}f(Z_t,\theta)\,\mathrm{d}t\stackrel{p}{\longrightarrow}\mathbb{E}[f(Z_0,\theta)]=\int_{\mathbb{R}}f(z,\theta)\phi_{K(0)}(z)\,\mathrm{d}z\] for each fixed \(\theta\in\overline{\Theta}\). Thus the desired convergence holds pointwise in \(\theta\).

It remains to prove uniformity in \(\theta\). For \(\theta_1,\theta_2\in\overline{\Theta}\), the mean value theorem and the growth condition on \(\partial_\theta f\) imply \[\left|\frac{1}{n}\sum_{i=1}^n f(Z_{t_{i-1}},\theta_1)-\frac{1}{n}\sum_{i=1}^n f(Z_{t_{i-1}},\theta_2)\right|\le |\theta_1-\theta_2|\frac{1}{n}\sum_{i=1}^n C(1+|Z_{t_{i-1}}|^p).\] The same bound holds for the deterministic limit because \(\mathbb{E}[1+|Z_0|^p]<\infty\). Applying the pointwise result to \(1+|x|^p\) gives \[\frac{1}{n}\sum_{i=1}^n (1+|Z_{t_{i-1}}|^p)=O_P(1).\] Since \(\overline{\Theta}\) is compact, a finite covering argument gives \[\sup_{\theta\in\overline{\Theta}}\left|\frac{1}{n}\sum_{i=1}^n f(Z_{t_{i-1}},\theta)-\int_{\mathbb{R}} f(z,\theta)\phi_{K(0)}(z)\,\mathrm{d}z\right|\stackrel{p}{\longrightarrow}0.\] This proves the assertion. ◻

Lemma 4. Let \(Z=(Z_t)_{t\ge0}\sim GP(0,K)\) be centered and stationary. Assume A1 and A2. Set \[Y_i^n:=\frac{\Delta_i^nZ}{\sqrt{v_n}},\quad v_n:=2\{K(0)-K(h_n)\}.\] Let \(f:\mathbb{R}\times\overline{\Theta}\to\mathbb{R}\) be continuous and suppose that \(f(\cdot,\theta)\) is continuously differentiable for every \(\theta\in\overline{\Theta}\). Assume that there exist constants \(C>0\) and \(p\ge1\) such that \[\sup_{\theta\in\overline{\Theta}}\{|f(x,\theta)|+|\partial_xf(x,\theta)|\}\le C(1+|x|^p), \quad x\in\mathbb{R},\] and that \(n^{-1}h_n^{\alpha-\beta}\to0\). Then, \[\sup_{\theta\in\overline{\Theta}}\left|\frac{1}{n}\sum_{i=1}^n f(Y_i^n,\theta)-\int_{\mathbb{R}}f(z,\theta)\phi(z)\,\mathrm{d}z\right|\stackrel{p}{\longrightarrow}0, \quad n\to \infty.\] where \(\phi\) is the probability density function of \({\cal N}(0,1)\).

Proof. Define \[G_n(\theta):=\frac{1}{n}\sum_{i=1}^n\left\{f(Y_i^n,\theta)-\mathbb{E}[f(Y_i^n,\theta)]\right\},\quad \theta\in\overline{\Theta},\] where \(\mathbb{E}[f(Y_i^n,\theta)]=\int_{\mathbb{R}}f(z,\theta)\phi(z)\,\mathrm{d}z\) by stationarity: \(Y_i^n\sim N(0,1)\).

For centered jointly Gaussian random variables \(X,Y\) with unit variance, the Gaussian covariance inequality yields \[|\mathrm{Cov}(f(X,\theta),f(Y,\theta))|\le C|\mathrm{Cov}(X,Y)|,\] where the constant \(C\) is independent of \(\theta\in\overline{\Theta}\) by the polynomial growth condition. Hence, by A1, \[\begin{align} \mathrm{Var}(G_n(\theta)) &\le \frac{C}{n^2}\sum_{i,j=1}^n |\mathrm{Cov}(Y_i^n,Y_j^n)| = \frac{C}{n^2}\sum_{i,j=1}^n \frac{|\mathrm{Cov}(\Delta_i^nZ,\Delta_j^nZ)|}{v_n}\\ &\sim \frac{C}{n^2h_n^\beta}\sum_{i,j=1}^n |\mathrm{Cov}(\Delta_i^nZ,\Delta_j^nZ)|\le Cn^{-1}h_n^{\alpha-\beta} \to 0. \end{align}\] Thus \[G_n(\theta)\stackrel{p}{\longrightarrow}0\] for each fixed \(\theta\in\overline{\Theta}\).

To prove uniform convergence, assume additionally that \[\sup_{\theta\in\overline{\Theta}}|\partial_\theta f(x,\theta)|\le C(1+|x|^p), \quad x\in\mathbb{R}.\] Then the same argument gives \[\sup_n\mathbb{E}\left[\sup_{\theta\in\overline{\Theta}}|\partial_\theta G_n(\theta)|\right]<\infty.\] Hence the sequence \(\{G_n(\theta)\}_n\) is stochastically equicontinuous on the compact set \(\overline{\Theta}\). Combining this with the pointwise convergence yields \[\sup_{\theta\in\overline{\Theta}}|G_n(\theta)|\stackrel{p}{\longrightarrow}0.\] This proves the assertion. ◻

Corollary 1. Under the same assumptions as in Lemma 4, for any integer \(k\ge1\), \[\frac{1}{nv_n^k}\sum_{i=1}^n (\Delta_i^nZ)^{2k}\stackrel{p}{\longrightarrow}\frac{(2k)!}{2^k k!}.\] where \(v_n:=2\{K(0)-K(h_n)\}\).

Proof. Apply Lemma 4 with \(f(x)=x^{2k}\). Since \(G\sim N(0,1)\) satisfies \(\mathbb{E}[G^{2k}]=(2k)!/(2^k k!)\), the assertion follows. ◻

Lemma 5. Let \(Z=(Z_t)_{t\ge0}\sim GP(0,K)\) be centered and stationary. Assume A1 and A2. Set \(v_n:=2\{K(0)-K(h_n)\}\). Let \(f:\mathbb{R}\times\overline{\Theta}\to\mathbb{R}\) be continuous and suppose that there exist constants \(C>0\) and \(p\ge1\) such that \[\sup_{\theta\in\overline{\Theta}}\{|f(x,\theta)|+|\partial_xf(x,\theta)|+|\partial_\theta f(x,\theta)|\}\le C(1+|x|^p), \quad x\in\mathbb{R}.\] Suppose also A3. Then \[\sup_{\theta\in\overline{\Theta}}\left|\frac{1}{n}\sum_{i=1}^n f\left(\frac{\Delta_i^nX-\mu_{\xi_0}(t_{i-1})h_n}{\sqrt{v_n}},\theta\right)-\int_{\mathbb{R}}f(z,\theta)\phi(z)\,\mathrm{d}z\right|\stackrel{p}{\longrightarrow}0.\] where \(\phi\) is the probability density function of \({\cal N}(0,1)\).

Proof. Set \(Y_i^n:=\Delta_i^nZ/\sqrt{v_n}\) and \(\widetilde{Y}_i^n:=(\Delta_i^nX-\mu_{\xi_0}(t_{i-1})h_n)/\sqrt{v_n}\). By the model, \[\widetilde{Y}_i^n-Y_i^n=\frac{1}{\sqrt{v_n}}\int_{t_{i-1}}^{t_i}\{\mu_{\xi_0}(s)-\mu_{\xi_0}(t_{i-1})\}\,\mathrm{d}s.\] By A3, \[\sup_{1\le i\le n}|\widetilde{Y}_i^n-Y_i^n|\le C\frac{h_n^2}{\sqrt{v_n}}=O(h_n^{2-\beta/2})\to0.\] Using the mean value theorem and the polynomial growth condition on \(\partial_xf\), we obtain \[\sup_{\theta\in\overline{\Theta}}\frac{1}{n}\sum_{i=1}^n |f(\widetilde{Y}_i^n,\theta)-f(Y_i^n,\theta)|\stackrel{p}{\longrightarrow}0.\] Lemma 4 gives \[\sup_{\theta\in\overline{\Theta}}\left|\frac{1}{n}\sum_{i=1}^n f(Y_i^n,\theta)-\int_{\mathbb{R}}f(z,\theta)\phi(z)\,\mathrm{d}z\right|\stackrel{p}{\longrightarrow}0.\] The assertion follows. ◻

References↩︎

[1]
Ibragimov, I. A. and Rozanov, Yu. A. (1978). Gaussian Random Processes, Springer-Verlag, New York.
[2]
Rasmussen, C. E. and Williams, C. K. I. (2006). Gaussian Processes for Machine Learning, the MIT Press, Massachusetts Institute of Technology. www.GaussianProcess.org/gpml.
[3]
Varin, C., Reid, N. and Firth, D. (2011). An overview of composite likelihood methods, Statistica Sinica, 21(1), 5–42.
[4]
Whittle, P. (1953). Estimation and information in stationary time series. Arkiv för Matematik, 2(5), 423–434.
[5]
Fukasawa, M. and Takabatake, T. (2019). Asymptotically efficient estimators for self-similar stationary Gaussian noises under high-frequency observations. Bernoulli, 25(3), 1870–1900.
[6]
Takabatake, T. (2024). Quasi-likelihood analysis of fractional Brownian motion with constant drift under high-frequency observations. Statistics & Probability Letters, 207, 110006.
[7]
Kobayashi, K., Nishiwaki, Y., Shimizu, Y. and Takaoka, N. (2025). Maximum likelihood estimation of mean functions for Gaussian processes under small noise asymptotics, arXiv preprint arXiv:2507.05628v1 [math.ST].
[8]
Kessler, M. (1997). Estimation of an ergodic diffusion from discrete observations, Scandinavian Journal of Statistics, 24(2), 211–229.
[9]
Feller, W. (1971). An Introduction to Probability Theory and Its Applications, vol. 2, 2nd ed., John Wiley and Sons, New York.
[10]
Kasonga, R. A. (1990). Parameter estimation by deterministic approximation of a solution of a stochastic differential equation. Communications in Statistics, 6, (1), 59–67.
[11]
van der Vaart, A.W. (1998). Asymptotic statistics. Cambridge University Press, Cambridge.

  1. Department of Applied Mathematics, Waseda University. E-mail:shimizu@waseda.jp↩︎