January 01, 1970
We study the cubic weakly nonlinear Schrödinger equation with randomized spatially quasi-periodic initial data in higher dimensions. Under a polynomial decay assumption in Fourier space, we establish a Large Deviations Principle for rogue waves in the so-called subcritical time regime.
The proof proceeds in two main steps. We first characterize the distribution of the linear solution and establish the corresponding linear large deviations principle. The lower bound is obtained via pointwise estimates, while the upper bound follows from a combination of truncation and probabilistic arguments. The method used in this step appears to be new; compare with [1]. We then perform a detailed combinatorial analysis of the Picard iteration, deriving an effective bound for the Duhamel term and thereby establishing the nonlinear large deviations principle.
The probabilistic study of dispersive partial differential equations originates from the seminal works of Lebowitz–Rose–Speer and Bourgain on invariant measures for the nonlinear Schrödinger equation; see [2]–[4]. These developments initiated a systematic investigation of low-regularity dynamics through randomization of either forcing terms or initial data.
From a structural point of view, such problems naturally split into two classes. The first concerns stochastic partial differential equations (SPDEs), where randomness enters through singular forcing and leads to dynamics such as the KPZ equation [5] and stochastic quantization models in Euclidean quantum field theory [6]. The second class consists of deterministic dispersive equations with random initial data, which provide a probabilistic framework for studying low-regularity behavior and are closely connected to wave turbulence theory [7], [8]. In many singular settings, a common feature is that the underlying analysis requires renormalization to handle divergent nonlinear interactions.
In the past decade, there has been remarkable progress in the analysis of singular stochastic parabolic equations. The development of the theory of regularity structures by Hairer [9], [10] and paracontrolled calculus by Gubinelli–Imkeller–Perkowski [6] has led to the local well-posedness theory, essentially completing the picture in the so-called subcritical regime; see [11] for further introduction.
More recently, the emerging theories of random tensors [11] and random averaging operators [12], developed by Deng–Nahmod–Yue, can be interpreted as dispersive counterparts to the aforementioned parabolic frameworks. These theories provide a systematic frameworks to the renormalization of oscillatory multilinear interactions; see also the corresponding hyperbolic counterpart [13].
Random data theory has played a fundamental role in the study of the long-time behaviour of solutions to evolution partial differential equations. One of the most active directions in this field is the derivation of the Wave Kinetic Equation (WKE), which connects the statistical behaviour of Fourier coefficients of solutions to an effective kinetic description. Significant recent advances in this direction can be found in [14]–[18], together with the references therein.
Another interesting direction related to random data theory concerns the so-called rogue waves, a term used by oceanographers to describe isolated and large-amplitude waves, typically defined as those whose height exceeds twice the significant wave height of the ambient sea state. Empirically, such events occur more frequently than predicted by Gaussian statistics; see [1], [19], [20].
To explain this phenomenon, several mechanisms for rogue wave formation have been proposed, although no unified theory is currently available. Linear superposition and nonlinear focusing are two fundamental mechanisms widely studied in the literature. The former is a constructive interference phenomenon: the phases of many weakly interacting waves align at a certain point in space, producing a large-amplitude peak. The latter suggests that rogue waves may emerge from a significant exchange of energy between waves of different wavenumbers, leading to a substantial amplification of one component; see [21]. In this setting, rogue waves are often viewed as limiting cases of breather solutions and may also arise from soliton collisions; see [19]. These represent non-generic behaviours and may be regarded as rare phenomena, or more generally as extreme events from a probabilistic perspective; see [1].
The probabilistic perspective suggests that rogue waves provide a canonical, yet still poorly understood, example of extreme events. A central problem is therefore to quantify deviations of the wave-height distribution from Gaussianity, and thereby to estimate the likelihood of rogue wave occurrence. This remains a long-standing problem with significant implications for ships and naval structures; see [20].
In [20], Dematteis, Grafke and Vanden-Eijnden observed that rogue waves obey a large deviations principle. This phenomenon was later rigorously established by Garrido et al. for the cubic weakly nonlinear Schrödinger (NLS) equation \[\begin{align} \label{eq:snls} \mathrm{i}\partial_t u + \partial_{xx} u + {\varepsilon^{\alpha}} |u|^{2} u = 0,\quad \alpha>1, \end{align}\tag{1}\] with randomized periodic initial data of the form \[u(0,x) = \sum_{n \in \mathbb{Z}} c(n)\, g_{n}\, e^{\mathrm{i} n x}.\] Their result holds over time scales \(t \sim \varepsilon^{-\beta}\) for both subcritical and critical cases, that is, \(\alpha - 1 > \beta\) and \(\alpha - 1 = \beta > 0\) respectively. Their analysis relies on an exponential decay assumption on the deterministic Fourier components of the random initial data, namely \(c(n)=a e^{-b|n|}\) or \(c(n)=a e^{-b|n|^2}\) for all \(n \in \mathbb{Z}\), where \(a,b>0\) are fixed. This assumption plays a crucial role in their argument; see [1] for further details. Recently, this condition has been relaxed. In [22], Liang and Wang extended the result to \(\ell^1(\mathbb{Z})\) decay, while Fan and Ye [23] further improved it to \(c_n = \langle n \rangle^{-\frac{1}{2}+}\).
In contrast to standard NLS 1 , Grande studied the following beating NLS \[\begin{align} {\mathrm{i}}\partial_t u + \partial_{xx} u = 2\cos(2x)\,|u|^2 u, \end{align}\] with randomized periodic initial data supported on two Fourier modes of the form \[u(0,x) = \varepsilon \bigl(\alpha e^{\mathrm{i}x} + \beta e^{-\mathrm{i}x}\bigr),\] where \(\alpha\) and \(\beta\) are complex-valued independent Gaussian random variables with zero mean and variances \(\sigma_\alpha^2 = \mathbb{E}\bigl[|\alpha|^2\bigr]\) and \(\sigma_\beta^2 = \mathbb{E}\bigl[|\beta|^2\bigr]\). When \(\sigma^2_\alpha \neq \sigma^2_\beta\), he showed that resonant energy exchange between Fourier modes leads to a fattening of the tails of the probability distribution of the sup-norm of the solution, thereby increasing the likelihood of rogue wave formation. Consequently, a large deviations principle holds; see [21].
The above discussion has been confined to the spatially periodic setting in the study of rogue waves. As noted by Wilkening and Zhao [24], rogue waves are also intimately connected with spatially quasi-periodic dynamics arising from modulational instability (also known as Benjamin–Feir instability). Thus, extending the study of rogue waves to the spatially quasi-periodic setting under randomization remains a natural and largely unexplored direction. In addition, the deterministic Cauchy problem with spatially quasi-periodic initial data forms an important and active area in nonlinear partial differential equations; see, for example, the works of Deift [25], [26] and subsequent works citing them.
Before formulating the problem precisely, we first introduce in this subsection the randomization of spatially quasi-periodic Fourier series on \(\mathbb{R}^d\).
A function \(f:\mathbb{R}^d \to \mathbb{C}\) is said to be quasi-periodic if it is quasi-periodic in each spatial direction \(x_j\), with associated rationally independent frequency vector \(\omega_j \in \mathbb{R}^{\nu_j}\) for \(j=1,\dots,d\), that is, \[\begin{align} \label{eq:rationallyindependent} \langle n_j,\omega_j \rangle= 0 \quad \Longrightarrow \quad n_j = 0 \in \mathbb{Z}^{\nu_j}. \end{align}\tag{2}\]
Let \(\nu := \sum_{j=1}^d \nu_j\) denote the total number of frequencies. We define the frequency matrix \(\boldsymbol{\Omega} \in \mathbb{R}^{\nu \times d}\) (with \(\nu > d\)) by \[\begin{align} \label{eq:frequencymatrix} \boldsymbol{\Omega} = \mathrm{diag}\bigl(\omega_1^{\top}, \dots, \omega_d^{\top}\bigr) = \begin{pmatrix} \omega_1^{\top} & & \\ & \ddots & \\ & & \omega_d^{\top} \end{pmatrix}. \end{align}\tag{3}\] We say that \(\boldsymbol{\Omega}\) is non-resonant or rationally independent if 2 holds for each direction. Equivalently, in a more compact form, \[\boldsymbol{n}\boldsymbol{\Omega} = \boldsymbol{0} \quad \Longrightarrow \quad \boldsymbol{n} = \boldsymbol{0} \in \mathbb{Z}^{\nu},\] where \(\boldsymbol{n} = (n_{1},\dots,n_d) \in \mathbb{Z}^{\nu_1} \times \cdots \times \mathbb{Z}^{\nu_d} \simeq \mathbb{Z}^{\nu}\).
The block-diagonal structure of \(\boldsymbol{\Omega}\) reflects the independence of the frequency components across different spatial directions. In particular, initial data [eq:initialdata] is quasi-periodic in each spatial direction \(x_j\) with frequency vector \(\omega_j\). Therefore, \(f\) admits the quasi-periodic Fourier expansion \[\label{eq:fourierseries} f(\boldsymbol{x}) = \sum_{\boldsymbol{n} \in \mathbb{Z}^\nu} \hat{f}(\boldsymbol{n}) \, e^{\mathrm{i}\langle \boldsymbol{n}\boldsymbol{\Omega}, \boldsymbol{x} \rangle}, \quad \boldsymbol{x} \in \mathbb{R}^d,\tag{4}\] where the Fourier coefficients \(\hat{f}(\boldsymbol{n})\) are determined by the spatial average \[\label{eq:fouriercoefficient} \hat{f}(\boldsymbol{n}) = \lim_{L \to \infty} \frac{1}{(2L)^d} \int_{[-L,L]^d} f(\boldsymbol{x}) \, e^{-\mathrm{i}\langle \boldsymbol{n}\boldsymbol{\Omega}, \boldsymbol{x} \rangle} \, \mathrm{d}\boldsymbol{x}.\tag{5}\]
By introducing the angular variables \(\boldsymbol{y} = \boldsymbol{x}\boldsymbol{\Omega}^\top \pmod{2\pi}\) on the torus \(\mathbb{T}^\nu\), we define the generating function \(F\) as the following Fourier expansion on \[F(\boldsymbol{y}) = \sum_{\boldsymbol{n} \in \mathbb{Z}^\nu} \hat{f}(\boldsymbol{n}) e^{\mathrm{i} \langle \boldsymbol{n}, \boldsymbol{y}\rangle}, \quad \boldsymbol{y} \in \mathbb{T}^\nu,\] such that \(f(\boldsymbol{x}) = F(\boldsymbol{x}\boldsymbol{\Omega}^\top)\).
By the Birkhoff Ergodic Theorem, the rational independence of frequency matrix \(\boldsymbol{\Omega}\) ensures that spatial average in 5 coincides with the Haar integral over the torus, that is, \[\hat{f}(\boldsymbol{n}) = \frac{1}{(2\pi)^\nu} \int_{\mathbb{T}^\nu} F(\boldsymbol{y}) e^{-\mathrm{i} \langle \boldsymbol{n}, \boldsymbol{y} \rangle} \,\mathrm{d}\boldsymbol{y}.\]
In this paper, we consider the randomization of Fourier series 4 by prescribing the coefficients as \[\hat{f}(\boldsymbol{n}) = c(\boldsymbol{n}) g_{\boldsymbol{n}}, \quad \boldsymbol{n} \in \mathbb{Z}^\nu.\]
Here \(\{c(\boldsymbol{n})\}_{\boldsymbol{n}\in\mathbb{Z}^\nu}\) is a deterministic sequence satisfying the following polynomial decay condition: \[\begin{align} \label{eq:decay} |c(\boldsymbol{n})| \le \tnorm{\boldsymbol{n}}_{-\rho} = \prod_{j=1}^{d}\prod_{j'=1}^{\nu_j} (1+|n_{j,j'}|)^{-\rho_{j,j'}}, \end{align}\tag{6}\] where \[\boldsymbol{n} = (n_{1,1},\cdots,n_{1,\nu_1};\cdots;n_{d,1},\cdots,n_{d,\nu_d}) \in \mathbb{Z}^{\nu_1} \times \cdots \times \mathbb{Z}^{\nu_d},\] and \[\rho = (\rho_{1,1},\cdots,\rho_{1,\nu_1};\cdots;\rho_{d,1},\cdots,\rho_{d,\nu_d}) \quad \text{with } \rho_{j,j'} > 0 ~~\text{large enough for all indices}.\]
The random component \(\{g_{\boldsymbol{n}}\}_{\boldsymbol{n} \in \mathbb{Z}^\nu}\) constitutes a family of independent and identically distributed (i.i.d. for short) standard complex Gaussian random variables satisfying the following conditions: \[\mathbb{E}[g_{\boldsymbol{n}}] = 0, \qquad \mathbb{E}[g_{\boldsymbol{n}}g_{\boldsymbol{m}}] = 0, \qquad\mathbb{E}[g_{\boldsymbol{n}} \overline{g_{\boldsymbol{m}}}] = \delta_{\boldsymbol{n},\boldsymbol{m}},\] where \(\delta_{\boldsymbol{n},\boldsymbol{m}}\) denotes the Kronecker delta. Equivalently, \(\operatorname{Re}\, g_{\boldsymbol{n}}\) and \(\operatorname{Im}\, g_{\boldsymbol{n}}\) are independent \(\mathcal{N}_{\mathbb{R}}(0,\tfrac{1}{2})\) random variables; see 5.5.1 for further details.
This construction defines a Gaussian random field \(F\) on the torus \(\mathbb{T}^\nu\), which induces a randomized quasi-periodic field \(f\) on \(\mathbb{R}^d\), that is, \[\label{eq:randomqp95d} f(\boldsymbol{x}) = \sum_{\boldsymbol{n} \in \mathbb{Z}^\nu} c(\boldsymbol{n}) g_{\boldsymbol{n}} e^{\mathrm{i} \langle \boldsymbol{n}\boldsymbol{\Omega}, \boldsymbol{x} \rangle}.\tag{7}\]
Proposition 1 (Expectation and Variance). Let randomized field \(f(\boldsymbol{x})\) be defined as in 7 . Then the following properties hold:
The field is centered: \(\mathbb{E}[f(\boldsymbol{x})] = 0\) for all \(\boldsymbol{x} \in \mathbb{R}^d\).
If the decay rates satisfy \({\rho_{j,j'} > 1/2}\) for each \(j=1,\dots,d\) and \(j'=1,\dots,\nu_j\), then the second moment is uniformly bounded: \(\mathbb{V}[f(\boldsymbol{x})] < \infty\).
Proof. (i) By linearity of the expectation operator, \[\begin{align} \mathbb{E}[f(\boldsymbol{x})] = \sum_{\boldsymbol{n} \in \mathbb{Z}^\nu} c(\boldsymbol{n}) e^{\mathrm{i} \langle \boldsymbol{n}\boldsymbol{\Omega}, \boldsymbol{x} \rangle} \mathbb{E}[g_{\boldsymbol{n}}]. \end{align}\] Since \(\{g_{\boldsymbol{n}}\}\) are centered, that is, \(\mathbb{E}[g_{\boldsymbol{n}}] = 0\) for all \(\boldsymbol{n} \in \mathbb{Z}^\nu\), it follows that \(\mathbb{E}[f(\boldsymbol{x})] = 0\).
(ii) For the second moment, we use \(|z|^2 = z \overline{z}\). By independence and orthogonality, we obtain \[\begin{align} \mathbb{E}\left[|f(\boldsymbol{x})|^2\right] &= \mathbb{E} \left[ \sum_{\boldsymbol{n},\boldsymbol{m} \in \mathbb{Z}^\nu} c(\boldsymbol{n}) \overline{c(\boldsymbol{m})} g_{\boldsymbol{n}} \overline{g_{\boldsymbol{m}}} e^{\mathrm{i} \langle (\boldsymbol{n}-\boldsymbol{m})\boldsymbol{\Omega}, \boldsymbol{x} \rangle} \right] \nonumber \\ &= \sum_{\boldsymbol{n},\boldsymbol{m} \in \mathbb{Z}^\nu} c(\boldsymbol{n}) \overline{c(\boldsymbol{m})} e^{\mathrm{i} \langle (\boldsymbol{n}-\boldsymbol{m})\boldsymbol{\Omega}, \boldsymbol{x} \rangle} \delta_{\boldsymbol{n},\boldsymbol{m}} \nonumber \\ &= \sum_{\boldsymbol{n} \in \mathbb{Z}^\nu} |c(\boldsymbol{n})|^2. \label{eq:parseval95iso} \end{align}\tag{8}\] Substituting polynomial decay assumption 6 yields \[\begin{align} \mathbb{E}\left[|f(\boldsymbol{x})|^2\right] &= \sum_{\boldsymbol{n} \in \mathbb{Z}^\nu} |c(\boldsymbol{n})|^2 \\ &\le \sum_{\substack{n_{j,j'} \in \mathbb{Z} \\ j=1,\dots,d \\ j' = 1,\dots,\nu_j}} \prod_{j=1}^d \prod_{j'=1}^{\nu_j} \left(1+|n_{j,j'}|\right)^{-2\rho_{j,j'}} \\ &= \prod_{j=1}^d \prod_{j'=1}^{\nu_j} \sum_{n_{j,j'} \in \mathbb{Z}} \left(1+|n_{j,j'}|\right)^{-2\rho_{j,j'}}. \end{align}\] Each one-dimensional summation converges provided \(\rho_{j,j'} > 1/2\). Consequently, the product is finite, ensuring that the second moment (and thus the variance) is finite and uniform with respect to \(\boldsymbol{x}\). ◻
Proposition 2 (Tail bound for \(g_{\boldsymbol{n}}\)). Let \(0 < \delta \ll 1\), and let \[\kappa=(\kappa_{1,1},\dots,\kappa_{1,\nu_1};\dots;\kappa_{d,1},\dots,\kappa_{d,\nu_d}), \qquad \kappa_{j,j'}>0~~\text{for all}~~j=1,\cdots,d~~\text{and}~~ 1\le j^\prime\leq\nu_j.\] Define the local sets \[\Omega_{\delta,\boldsymbol{n}} = \Bigl\{ \omega \in \Omega : |g_{\boldsymbol{n}}(\omega)| > {\delta^{-\frac{1}{2}}}\tnorm{\boldsymbol{n}}_\kappa \Bigr\},\] where \[\tnorm{\boldsymbol{n}}_\kappa = \prod_{j=1}^d \prod_{j'=1}^{\nu_j} (1 + |n_{j,j'}|)^{\kappa_{j,j'}},\] and the global set \[\Omega_\delta = \bigcup_{\boldsymbol{n} \in \mathbb{Z}^\nu} \Omega_{\delta,\boldsymbol{n}}.\] Then we have \[\mathbb{P}(\Omega_\delta) =\mathcal{O}( e^{-\delta^{-1}}).\]
Proof. Step 1. Tail probability of \(g_{\boldsymbol{n}}\).
Since \(g_{\boldsymbol{n}} \sim \mathscr{N}_{\mathbb{C}}(0,1)\), we have \(|g_{\boldsymbol{n}}|^2 \sim \mathscr E(1)\). Therefore, \[\mathbb{P}(\Omega_{\delta,\boldsymbol{n}}) = \exp\bigl(-\delta^{-1}\tnorm{\boldsymbol{n}}_{\kappa}^2\bigr).\] Since \(\tnorm{\boldsymbol{n}}_{\kappa} \ge 1\) and \(\tnorm{\boldsymbol{n}}_{\kappa} = 1\) if and only if \(\boldsymbol{n}=0\), we have \[\mathbb{P}(\Omega_{\delta,\boldsymbol{0}})=e^{-\delta^{-1}} \quad\text{and}\quad \tnorm{\boldsymbol{n}}_{\kappa}^2 - 1 > 0 \;\text{for } \boldsymbol{n}\neq 0.\] Thus, \[\mathbb{P}(\Omega_\delta) \le e^{-\delta^{-1}} + \sum_{\boldsymbol{n}\neq 0} \exp\bigl(-\delta^{-1}\tnorm{\boldsymbol{n}}_{\kappa}^2\bigr),\] and equivalently, \[\mathbb{P}(\Omega_\delta) = e^{-\delta^{-1}} \Biggl( 1 + \sum_{\boldsymbol{n}\neq 0} \exp\bigl(-\delta^{-1}(\tnorm{\boldsymbol{n}}_{\kappa}^2-1)\bigr) \Biggr).\]
Step 2. Positive gap.
Since \(\tnorm{\boldsymbol{n}}_\kappa \to \infty\) as \(|\boldsymbol{n}|\to\infty\) and \(\tnorm{\boldsymbol{n}}_\kappa > 1\) for all \(\boldsymbol{n}\neq 0\), the quantity \(\tnorm{\boldsymbol{n}}_\kappa^2 - 1\) attains a positive minimum on \(\mathbb{Z}^\nu \setminus \{0\}\). Hence there exists \(c_0>0\) such that \[\tnorm{\boldsymbol{n}}_\kappa^2 - 1 \ge c_0.\] Define the level sets \[L_t = \Bigl\{ \boldsymbol{n}\in\mathbb{Z}^\nu: c_0+t \le \tnorm{\boldsymbol{n}}_{\kappa}^2 - 1 < c_0+t+1 \Bigr\}, \quad t\in\mathbb{N}.\] Then \[\sum_{\boldsymbol{n}\neq 0} e^{-\delta^{-1}(\tnorm{\boldsymbol{n}}_{\kappa}^2-1)} = \sum_{t\ge 0}\sum_{\boldsymbol{n}\in L_t} e^{-\delta^{-1}(\tnorm{\boldsymbol{n}}_{\kappa}^2-1)}.\]
Step 3. Counting argument.
For \(\boldsymbol{n}\in L_t\), we have \[\prod_{j,j'} (1+|n_{j,j'}|)^{2\kappa_{j,j'}} < c_0+t+2.\] Taking logarithms and using positivity of all \(\kappa_{j,j'}\), it follows that each coordinate satisfies \[1+|n_{j,j'}| < (c_0+t+2)^{\frac{1}{2\kappa_{j,j'}}}.\] Hence, for any fixed \(1 \leq j \leq d\) and \(1 \leq j' \leq \nu_j\), the corresponding \(n_{j,j'}\) satisfies \[\begin{align} n_{j,j'} \leq 2\left( (c_0 + t + 2)^{\frac{1}{2\kappa_{j,j'}}} - 1 \right) + 1 = 2(c_0 + t + 2)^{\frac{1}{2\kappa_{j,j'}}} - 1 < 2(c_0 + t + 2)^{\frac{1}{2\kappa_{j,j'}}}. \end{align}\] Thus \[|L_t| <2^\nu (c_0+t+2)^{\sum_{j=1}^d \sum_{j'=1}^{\nu_j} \frac{1}{2\kappa_{j,j'}}}.\]
Step 4. Summation and conclusion.
We estimate \[\sum_{t\ge 0} |L_t| e^{-\delta^{-1}(c_0+t)} \lesssim e^{-c_0\delta^{-1}} \sum_{t\ge 0} (c_0+t+2)^{\sum_{j=1}^d \sum_{j'=1}^{\nu_j} \frac{1}{2\kappa_{j,j'}}} e^{-t\delta^{-1}}.\] Since \(e^{-t\delta^{-1}}\) decays exponentially fast, \[\sum_{t\ge 0} (c_0+t+2)^{\sum_{j=1}^d \sum_{j'=1}^{\nu_j} \frac{1}{2\kappa_{j,j'}}} e^{-t\delta^{-1}} \lesssim \delta^{1+\sum_{j=1}^d \sum_{j'=1}^{\nu_j} \frac{1}{2\kappa_{j,j'}}}.\] Therefore, \[\mathbb{P}(\Omega_\delta) \lesssim e^{-\delta^{-1}} \left(1 + \delta^{1+\sum_{j=1}^d \sum_{j'=1}^{\nu_j} \frac{1}{2\kappa_{j,j'}}}e^{-c_0\delta^{-1}}\right) \lesssim e^{-\delta^{-1}}.\] ◻
Corollary 1. Assume that the following conditions hold:
(Related to initial data in Fourier space) Let \(\rho\) and \(\kappa\) be as in 6 and 2, respectively. Assume that \[\rho_{j,j'} - \kappa_{j,j'} > 2, \quad j=1,\dots,d,\; j'=1,\dots,\nu_j.\]
(Size of initial data with respect to \(\varepsilon\)) Let \(\omega \notin \Omega_{\delta}\) with \(\delta\sim\varepsilon^{1+\eta}\), where \[\label{eq:delta} 0 \le \eta < \min\left\{ 2\Bigl(\frac{1}{\nu}-\mu_1\Bigr)\eta_2-\mu_2,\; \eta_3\right\}.\tag{9}\] Here \(0<\mu_1 \ll 1/\nu, 0<\mu_2 \ll 1\), \(\eta_2\) and \(\eta_3\) are given by \[\begin{align} \eta_2 &= \sum_{j=1}^{d} \sum_{j'=1}^{\nu_j} \bigl(\rho_{j,j'} - \kappa_{j,j'} - 1\bigr), \label{eta2}\\ \eta_3 &= \frac{2}{3}(\alpha-1)-\frac{\mu_3}{3}, \quad 0<\mu_3 \ll 1. \end{align}\tag{10}\]
Then the randomized Fourier coefficients satisfy the following almost sure polynomial decay estimates: \[\begin{align} \label{eq:decayall} \bigl|c(\boldsymbol{n}) g_{\boldsymbol{n}}(\omega)\bigr| \le \varepsilon^{-\frac{1}{2}-\frac{\eta}{2}} \,\tnorm{\boldsymbol{n}}_{-(\rho-\kappa)},\quad\omega\notin\Omega_\delta, \end{align}\tag{11}\] where \[\begin{align} \label{eq:rhokappa} \rho - \kappa = \bigl( \rho_{1,1}-\kappa_{1,1}, \dots, \rho_{1,\nu_1}-\kappa_{1,\nu_1}; \dots; \rho_{d,1}-\kappa_{d,1}, \dots, \rho_{d,\nu_d}-\kappa_{d,\nu_d} \bigr). \end{align}\tag{12}\]
Remark 1 (\(\rho-\kappa\)). Roughly speaking, the condition on \(\rho\) and \(\kappa\) requires \(\rho_{j,j'} - \kappa_{j,j'}\) to be sufficiently large to ensure absolute convergence of series appearing in estimates such as 3.
Remark 2 (\(\eta\)). Regarding the assumption on \(\eta\), there are three remarks below.
The condition \(\eta\ge0\) ensures that the second term in 28 is asymptotically negligible compared to the first term. If this condition fails, for example, \(\delta = \varepsilon^{1/2}\), then \[\varepsilon \log\left(1 + \frac{\mathbb{P}(\Omega_\delta)}{\mathbb{P}(\mathscr A_\varepsilon^{(3)})}\right) \sim \varepsilon \log\left(1 + e^{-\frac{1}{\sqrt{\varepsilon}} + \frac{1}{\varepsilon}}\right) \sim1-\sqrt{\varepsilon}\to 1,\quad\varepsilon\to0^+.\] Hence, the contribution of the bad set is no longer negligible at the scale of large deviations. Similar analysis to 42 and 44 .
The condition \(\eta < 2\left(\frac{1}{\nu}-\mu_1\right)\eta_2-\mu_2\) is required in order to get the effective scale \(\mathcal{O}(\varepsilon^{-\frac{1}{2}+\frac{\mu_2}{2}})\) of \(\mathscr R_N\); see 1.
The condition \(\eta < \eta_3\) is imposed to obtain the effective scale \(\mathcal{O}(\varepsilon^{-\frac{1}{2}+\frac{\mu_3}{2}})\) for the Duhamel estimates by working on a shorter time scale, namely after reducing the original time scale by a factor of \(\varepsilon^{\frac{\eta}{2}+\frac{\mu_3}{2}}\); see 4.
Motivated by the preceding discussion, we consider the Cauchy problem on \(\mathbb{R}^d\) for a weakly nonlinear Schrödinger equation (NLS) with randomized quasi-periodic initial data:
align &_t u + u + ^ |u|^2 u = 0,>1;
&u(0, ) = _ ^ c() g_ e^ , .
Here \(u = u(t,\boldsymbol{x})\) is a complex-valued function with \((t,\boldsymbol{x}) \in \mathbb{R} \times \mathbb{R}^d\). The operator \(\partial_t\) denotes the time derivative, and \(\Delta = \sum_{j=1}^d \partial_j^2\) is the Laplacian, where \(\partial_j = \frac{\partial}{\partial x_j}\). The parameter \(0 < \varepsilon \ll 1\) measures the strength of the nonlinearity, and \(\mathrm{i}\) denotes the imaginary unit satisfying \(\mathrm{i}^2 = -1\).
For model [eq:nls]-[eq:initialdata], we study the corresponding rogue waves introduced above. Although evolution equation [eq:nls] is deterministic, randomness in initial data [eq:initialdata] induces a stochastic process \(t \mapsto u(t,\boldsymbol{x})\) for each fixed \(\boldsymbol{x} \in \mathbb{R}^d\), and a random field \(\boldsymbol{x} \mapsto u(t,\boldsymbol{x})\) for each fixed \(t \in \mathbb{R}\).
We quantify such extreme events by introducing the tail event \[\label{eq:rogue} \mathscr{A}_\varepsilon^\infty = \left\{ \sup_{\boldsymbol{x} \in \mathbb{R}^d} |u(t, \boldsymbol{x})| > z_0 {\varepsilon^{-\frac{1}{2}}} \right\},\tag{13}\] where \(z_0 > 0\) is a fixed threshold. This event describes the occurrence of a wave of amplitude at least \(z_0 \varepsilon^{-1/2}\) at a fixed time \(t>0\).
The main objective of this paper is to describe the asymptotic behavior of the probability \(\mathbb{P}(\mathscr{A}_\varepsilon^\infty)\) as \(\varepsilon \to 0^+\).
Our main result is a nonlinear large deviations principle in the subcritical time regime (1), together with a linear large deviations principle (2) and a well-posedness result (3).
Theorem 1 (Large Deviations Principle). Consider randomized quasi-periodic Cauchy problem [eq:nls]–[eq:initialdata] together with 1, and the rogue wave event \(\mathscr{A}_\varepsilon^\infty\) defined in 13 . If the observation time satisfies \[t =\mathcal{O}\left(\varepsilon^{-\beta}\right), \quad\beta=\alpha-1-\frac{3}{2}\eta-\frac{\mu_3}{2}.\] Then we have \[\lim_{\varepsilon\to0^+} \varepsilon \log \mathbb{P}\big(\mathscr{A}_\varepsilon^\infty\big) = -\frac{z_0^2}{\sum_{\boldsymbol{n}\in\mathbb{Z}^\nu} |c(\boldsymbol{n})|^2}.\]
Remark 3 (Function Space). To the best of our knowledge, this is the first large deviations result on randomization for the spatially quasi-periodic data Cauchy problem in the framework of dispersive equations.
In contrast to decaying or periodic functions, quasi-periodic functions exhibit persistent oscillations at infinity, which makes the analysis considerably more delicate than in the classical settings.
Remark 4 (Time Scale). Clearly, \(\alpha - 1 = \beta+\frac{3}{2}\eta+\frac{\mu_3}{2}>\beta\), which corresponds to the so-called subcritical regime in the sense of [1], namely the regime preceding the onset of the first nonlinear effect; see [27].
Remark 5 (Decay Assumption). In [1], Garrido et al. established a large deviations principle for the periodic cubic NLS in one spatial dimension under an exponential decay assumption in Fourier space, or rather, on the determining component \(\{c_n\}_{n\in\mathbb{Z}}\) of Fourier coefficients. They further raised the question of identifying the optimal decay condition under which the LDP holds.
Recently, this assumption has been relaxed. In [22], Liang and Wang improved the result to \(\ell^1(\mathbb{Z})\) decay, while Fan and Ye [23] further extended it to \(c_{n\in\mathbb{Z}} = \langle n \rangle^{-(\frac{1}{2}+\theta)}\) for any \(\theta>0\).
In the present paper, we consider the randomized quasi-periodic Cauchy problem [eq:nls]–[eq:initialdata] under a polynomial decay assumption of the form 6 based on [28]. It is unclear whether the corresponding exponent is optimal, and determining the sharp decay threshold ensuring the validity of the large deviations principle in the quasi-periodic setting remains open.
Remark 6 (Proof). (a) The proof of 1 consists of three main steps.
We first establish a linear LDP for the associated linear problem. In this step, we analyze the explicit law of the linear solution, in particular its modulus squared. The lower bound follows from pointwise estimates, while the upper bound is obtained via global control in polar coordinates, combined with analytic and probabilistic arguments; see 2 and 1.
We then carry out a well-posedness analysis and derive Duhamel estimates on the relevant time scale. To this end, we employ a combinatorial argument (see [29]–[31]) to identify the relevant time scale and derive decay properties of the solution; see 3.
Finally, we establish the nonlinear LDP by combining Duhamel’s formula with the above estimates, thereby completing the proof; see 4.
(b) In the one-dimensional periodic setting, that is, on the torus \(\mathbb{T}\), we compare our method with that of [1].
In the linear setting, the analysis of [1] is based on the Gärtner–Ellis theorem, reducing the derivation of the large deviation principle (LDP) to the computation of the cumulant generating function. The associated rate function is then obtained through the Fenchel–Legendre transform, yielding the linear LDP. Moreover, their argument relies on specific number-theoretic properties of \(\mathbb{Z}\), whose analogues in the quasi-periodic setting seem unavailable.
In contrast, we directly exploit the explicit distribution of the squared modulus of the linear solution, combined with truncation and probabilistic arguments, which is new and simplifies the original argument. Independently, Liang and Wang also provided a new proof, avoiding combinatorial arguments and instead relying on analytic methods, for the linear LDP within the framework of the Gärtner-Ellis theorem; see [22].
In 2, we establish a linear LDP for the associated linear problem. In 3, we prove a well-posedness result, including the relevant time scale and Fourier decay properties. In 4, we derive the Duhamel estimates and establish the nonlinear LDP. In 5, we collect some basic facts on probability theory related to complex Gaussian random variables.
We summarize some notation used throughout the paper below.
| Notation | Meaning |
|---|---|
| \({}^\top\) | transpose (vectors are written in row form) |
| \(\langle \boldsymbol{y},\boldsymbol{z}\rangle\) | \(\boldsymbol{y}\boldsymbol{z}^{\top}=\sum_{j=1}^N y_j z_j\) |
| \(A=\mathcal{O}(B)\;(\text{or } A\lesssim B)\) | \(|A|\le CB\) for some \(C>0\) independent of parameters |
| \(A=\mathcal{O}_D(B)\;(\text{or } A\lesssim_D B)\) | \(|A|\le C_D B\), where \(C_D\) may depend on \(D\) |
| \(a+\) | \(a+\mu\), where \(0<\mu\ll1\) arbitrarily small |
| \(a-\) | \(a-\mu\), where \(0<\mu\ll1\) arbitrarily small |
Consider the linear problem obtained by setting \(\varepsilon = 0\) in [eq:nls], namely \[\begin{cases} \mathrm{i}\partial_t u + \Delta u = 0,\\ u(0,\boldsymbol{x}) = \displaystyle \sum_{\boldsymbol{n} \in \mathbb{Z}^\nu} c(\boldsymbol{n})\, g_{\boldsymbol{n}}\, e^{\mathrm{i}\langle \boldsymbol{n}\boldsymbol{\Omega}, \boldsymbol{x} \rangle}. \end{cases}\] The solution is given explicitly by random Fourier series \[\label{eq:linearsolution} u_{\mathrm{linear}}(t,\boldsymbol{x}) = \sum_{\boldsymbol{n} \in \mathbb{Z}^\nu} c(\boldsymbol{n})\, g_{\boldsymbol{n}}\, e^{-\mathrm{i} t Q(\boldsymbol{n})}\, e^{\mathrm{i}\langle \boldsymbol{n}\boldsymbol{\Omega}, \boldsymbol{x} \rangle},\tag{14}\] where the dispersion relation (neglecting the minus) is \[Q(\boldsymbol{n}) = \sum_{j=1}^d \big\langle n_j,\omega_j \big\rangle^2, \qquad \boldsymbol{n}=(n_1,\cdots,n_d)\in\mathbb{Z}^{\nu_1}\times\cdots\times\mathbb{Z}^{\nu_d}.\]
Theorem 2 (Linear LDP). Replacing \(u\) in 13 by linear evolution 14 , we obtain the following large deviations principle: \[\begin{align} \label{eq:linearldp} \lim_{\varepsilon \to 0^+} \varepsilon \log \mathbb{P}(\mathscr{A}_\varepsilon^\infty) = -\frac{z_0^2}{\sum_{\boldsymbol{n} \in \mathbb{Z}^\nu} |c(\boldsymbol{n})|^2}. \end{align}\tag{15}\]
In this subsection, we study the distribution of the linear solution and its modulus squared.
Since \(\{g_{\boldsymbol{n}}\}_{\boldsymbol{n} \in \mathbb{Z}^\nu}\) are i.i.d.standard complex Gaussian random variables, the linear evolution preserves Gaussianity. In particular, \(u_{\mathrm{linear}}(t,\boldsymbol{x})\) is a complex Gaussian random variable for every \((t,\boldsymbol{x}) \in \mathbb{R} \times \mathbb{R}^d\). Moreover, it is centered, i.e., \[\mathbb{E}\big[u_{\mathrm{linear}}(t,\boldsymbol{x})\big]=0.\] In addition, it follows from \(\mathbb{E}[g_{\boldsymbol{n}} g_{\boldsymbol{m}}] = 0\) and \(\mathbb{E}[g_{\boldsymbol{n}} \overline{g_{\boldsymbol{m}}}] = \delta_{\boldsymbol{n},\boldsymbol{m}}\) that \[\begin{align} \mathbb{E}\big[ (u_{\mathrm{linear}}(t, \boldsymbol{x}))^2 \big] &= \mathbb{E} \Bigg[ \sum_{\boldsymbol{n}, \boldsymbol{m} \in \mathbb{Z}^\nu} c(\boldsymbol{n}) c(\boldsymbol{m}) \, g_{\boldsymbol{n}} g_{\boldsymbol{m}} \, e^{-\mathrm{i} t (Q(\boldsymbol{n})+Q(\boldsymbol{m}))} e^{\mathrm{i} \langle (\boldsymbol{n}+\boldsymbol{m})\boldsymbol{\Omega}, \boldsymbol{x} \rangle} \Bigg] \nonumber \\ &= \sum_{\boldsymbol{n}, \boldsymbol{m} \in \mathbb{Z}^\nu} c(\boldsymbol{n}) c(\boldsymbol{m}) \, \mathbb{E}\big[g_{\boldsymbol{n}} g_{\boldsymbol{m}}\big] \, e^{-\mathrm{i} t (Q(\boldsymbol{n})+Q(\boldsymbol{m}))} e^{\mathrm{i} \langle (\boldsymbol{n}+\boldsymbol{m})\boldsymbol{\Omega}, \boldsymbol{x} \rangle} \nonumber \\ &=0, \end{align}\] and \[\begin{align} \mathbb{E}\big[ |u_{\mathrm{linear}}(t, \boldsymbol{x})|^2 \big] &= \mathbb{E} \Bigg[ \sum_{\boldsymbol{n}, \boldsymbol{m} \in \mathbb{Z}^\nu} c(\boldsymbol{n}) \overline{c(\boldsymbol{m})} \, g_{\boldsymbol{n}} \overline{g_{\boldsymbol{m}}} \, e^{-\mathrm{i} t (Q(\boldsymbol{n})-Q(\boldsymbol{m}))} e^{\mathrm{i} \langle (\boldsymbol{n}-\boldsymbol{m})\boldsymbol{\Omega}, \boldsymbol{x} \rangle} \Bigg] \nonumber \\ &= \sum_{\boldsymbol{n}, \boldsymbol{m} \in \mathbb{Z}^\nu} c(\boldsymbol{n}) \overline{c(\boldsymbol{m})} \, \mathbb{E}\big[g_{\boldsymbol{n}} \overline{g_{\boldsymbol{m}}}\big] \, e^{-\mathrm{i} t (Q(\boldsymbol{n})-Q(\boldsymbol{m}))} e^{\mathrm{i} \langle (\boldsymbol{n}-\boldsymbol{m})\boldsymbol{\Omega}, \boldsymbol{x} \rangle} \nonumber \\ &= \sum_{\boldsymbol{n} \in \mathbb{Z}^\nu} |c(\boldsymbol{n})|^2\nonumber\\ &\triangleq 2\sigma^2. \end{align}\]
Consequently, for each \((t,\boldsymbol{x}) \in \mathbb{R} \times \mathbb{R}^d\), the linear solution \(u_{\mathrm{linear}}(t,\boldsymbol{x})\) is distributed as a centered complex Gaussian \[u_{\mathrm{linear}}(t,\boldsymbol{x}) \sim \mathcal{N}_{\mathbb{C}}(0,2\sigma^2),\] with variance independent of both time and space.
By the properties of complex Gaussian random variables, the normalized modulus squared follows a chi-squared distribution with two degrees of freedom, which is equivalent to an exponential distribution with rate parameter \(1/2\). That is, \[\sigma^{-2} |u_{\mathrm{linear}}(t,\boldsymbol{x})|^2\sim\chi_2^2=\mathscr E\left(\frac{1}{2}\right).\] Equivalently, \(|u(t,\boldsymbol{x})|^2\) is exponentially distributed with rate parameter \(\frac{1}{2\sigma^2}\), i.e. \[|u_{\mathrm{linear}}(t,\boldsymbol{x})|^2 \sim \mathscr{E}\!\left(\frac{1}{2\sigma^2}\right),\] with probability density function \[f(x) = \frac{1}{2\sigma^2}\, e^{-\frac{x}{2\sigma^2}} \boldsymbol{1}_{[0,\infty)}(x).\]
Recall that \(u_{\mathrm{linear}}(t,\boldsymbol{x})\) is a centered complex Gaussian random field with variance \(2\sigma^2\), invariant in both time and space. Consequently, for every fixed \((t,\boldsymbol{x})\in\mathbb{R}\times\mathbb{R}^d\), the random variable \(|u_{\mathrm{linear}}(t,\boldsymbol{x})|^2\) follows an exponential distribution. This enables us to obtain a lower bound through the analysis of its pointwise tail probability, which we carry out in this subsection.
To this end, for any fixed \((t,\boldsymbol{x}) \in \mathbb{R} \times \mathbb{R}^d\), we define the event \[\mathscr{A}^{(0)}_\varepsilon = \left\{ |u_{\mathrm{linear}}(t,\boldsymbol{x})| > z_0\,\varepsilon^{-1/2} \right\}.\]
Since the supremum over the spatial domain dominates any pointwise value, we have \(\mathscr{A}^{(0)}_\varepsilon \subset \mathscr{A}_\varepsilon^\infty\). Consequently, \[\begin{align} \mathbb{P}(\mathscr{A}_\varepsilon^\infty) &\geq \mathbb{P}(\mathscr{A}^{(0)}_\varepsilon) \\ &= \mathbb{P}\big(|u_{\mathrm{linear}}(t,\boldsymbol{x})|^2 \geq z_0^2 \varepsilon^{-1}\big) \\ &= \int_{z_0^2 \varepsilon^{-1}}^{\infty} \frac{1}{2\sigma^2} e^{-\frac{x}{2\sigma^2}}\, dx \\ &= \exp\!\left(-\frac{z_0^2}{2\sigma^2}\varepsilon^{-1}\right). \end{align}\]
Taking logarithms and multiplying by \(\varepsilon\), we obtain \[\varepsilon \log \mathbb{P}(\mathscr{A}_\varepsilon^\infty) \geq -\frac{z_0^2}{2\sigma^2}.\] Passing to the limit \(\varepsilon \to 0^+\) yields \[\liminf_{\varepsilon \to 0^+} \varepsilon \log \mathbb{P}(\mathscr{A}_\varepsilon^\infty) \geq -\frac{z_0^2}{2\sigma^2}.\]
Recalling that \(2\sigma^2 = \sum_{\boldsymbol{n} \in \mathbb{Z}^\nu} |c(\boldsymbol{n})|^2\), we obtain \[\label{eq:lower} \liminf_{\varepsilon \to 0^+} \varepsilon \log \mathbb{P}(\mathscr{A}_\varepsilon^\infty) \geq -\frac{z_0^2}{\sum_{\boldsymbol{n} \in \mathbb{Z}^\nu} |c(\boldsymbol{n})|^2}.\tag{16}\]
In this subsection, we establish an upper bound for the probability of the rogue wave event \(\mathscr{A}_\varepsilon^\infty\). In contrast to the pointwise argument used for the lower bound, the proof of the upper bound relies on a global control of the random field, obtained via polar coordinates, truncation , Hölder’s inequality, and standard probabilistic estimates including Chernoff bound.
Let \(g_{\boldsymbol{n}} = r_{\boldsymbol{n}} e^{\mathrm{i}\theta_{\boldsymbol{n}}}\), where \(r_{\boldsymbol{n}} \ge 0\) and \(\theta_{\boldsymbol{n}}\in[0,2\pi)\) are independent, with \(2r_{\boldsymbol{n}}^2 \overset{\mathrm{iid}}{\sim} \mathscr{E}(1/2)\) and \(\theta_{\boldsymbol{n}} \overset{\mathrm{iid}}{\sim} \mathscr{U}[0,2\pi]\). Substituting this polar representation into 14 , we obtain \[\begin{align} \label{usolution} u_{\mathrm{linear}}(t,\boldsymbol{x}) = \sum_{\boldsymbol{n}\in\mathbb{Z}^\nu} c(\boldsymbol{n})\, r_{\boldsymbol{n}}\, e^{\mathrm{i}\phi_{\boldsymbol{n}}(t,\boldsymbol{x})}, \end{align}\tag{17}\] where the phase \[\phi_{\boldsymbol{n}}(t,\boldsymbol{x}) = \theta_{\boldsymbol{n}} - Q(\boldsymbol{n})t + \langle \boldsymbol{n}\boldsymbol{\Omega}, \boldsymbol{x} \rangle.\]
To control the supremum of the field, we observe that \[\begin{align} \sup_{\boldsymbol{x}\in\mathbb{R}^d} |u_{\mathrm{linear}}(t,\boldsymbol{x})|^2 &= \sup_{\boldsymbol{x}\in\mathbb{R}^d} \operatorname{Re}\!\left(u_{\mathrm{linear}}(t,\boldsymbol{x})\,\overline{u_{\mathrm{linear}}(t,\boldsymbol{x})}\right) \nonumber\\ &= \sup_{\boldsymbol{x}\in\mathbb{R}^d} \sum_{\boldsymbol{n},\boldsymbol{m}\in\mathbb{Z}^\nu} r_{\boldsymbol{n}} r_{\boldsymbol{m}} \operatorname{Re}\!\left( c(\boldsymbol{n})\,\overline{c(\boldsymbol{m})}\, e^{\mathrm{i}(\phi_{\boldsymbol{n}}(t,\boldsymbol{x})-\phi_{\boldsymbol{m}}(t,\boldsymbol{x}))} \right) \nonumber\\ &\le \sum_{\boldsymbol{n},\boldsymbol{m}\in\mathbb{Z}^\nu} r_{\boldsymbol{n}} r_{\boldsymbol{m}}\, |c(\boldsymbol{n})|\, |c(\boldsymbol{m})| \nonumber\\ &= \left(\sum_{\boldsymbol{n}\in\mathbb{Z}^\nu} |c(\boldsymbol{n})|\, r_{\boldsymbol{n}}\right)^2. \end{align}\]
Consequently, we obtain the uniform bound \[\begin{align} \label{eq:cr} \sup_{\boldsymbol{x}\in\mathbb{R}^d} |u_{\mathrm{linear}}(t,\boldsymbol{x})| \le \sum_{\boldsymbol{n}\in\mathbb{Z}^\nu} |c(\boldsymbol{n})|\, r_{\boldsymbol{n}}, \end{align}\tag{18}\] in which we have gotten rid of the supremum on \(\mathbb{R}^d\).
Remark 8. Set \[\mathscr A^{(1)}_{\varepsilon}=\left\{\sum_{\boldsymbol{n}\in\mathbb{Z}^\nu} |c(\boldsymbol{n})|\, r_{\boldsymbol{n}}>z_0\varepsilon^{-1/2}\right\}.\] Clearly, \[\begin{align} \label{eq:i1} \mathscr{A}_\varepsilon^\infty\subset\mathscr A_\varepsilon^{(1)}. \end{align}\tag{19}\]
To study the probability of \(\mathscr A_\varepsilon^{(1)}\), we introduce the truncation parameter \[\begin{align} \label{eq:N} N=\varepsilon^{-\frac{1}{\nu}+\mu_1}, \qquad 0<\mu_1\ll\frac{1}{\nu}. \end{align}\tag{20}\]
We define the truncated lattice \(\Lambda_N\subset\mathbb{Z}^\nu\) by \[\begin{align} \Lambda_N &=\left\{\boldsymbol{n}\in\mathbb{Z}^\nu:\;|\boldsymbol{n}|\le N\right\} \\ &=\left\{ (n_1,\ldots,n_d)\in\prod_{j=1}^{d}\mathbb{Z}^{\nu_j} :\;|n_j|\le N,\quad j=1,\ldots,d \right\} \\ &=\left\{ (n_{1,1},\ldots,n_{1,\nu_1};\,\ldots;\,n_{d,1},\ldots,n_{d,\nu_d}) \in\prod_{j=1}^{d}\mathbb{Z}^{\nu_j} :\;|n_{j,j'}|\le N,\; j'=1,\ldots,\nu_j,\; j=1,\ldots,d \right\}. \end{align}\]
Using the truncation set \(\Lambda_N\), we decompose the right-hand side of 18 into a principal part \(\mathscr S_N\) and a remainder term \(\mathscr R_N\): \[\label{eq:sr} \sum_{\boldsymbol{n}\in\mathbb{Z}^\nu} |c(\boldsymbol{n})|\,r_{\boldsymbol{n}} = \mathscr S_N+\mathscr R_N,\tag{21}\] where \[\mathscr S_N = \sum_{\boldsymbol{n}\in\Lambda_N} |c(\boldsymbol{n})|\,r_{\boldsymbol{n}}, \qquad \mathscr R_N = \sum_{\boldsymbol{n}\notin\Lambda_N} |c(\boldsymbol{n})|\,r_{\boldsymbol{n}}.\]
Remark 9. With the above notation, the event \(\mathscr A_{\varepsilon}^{(1)}\) can be expressed as \[\begin{align} \label{eq:a1} \mathscr A_{\varepsilon}^{(1)} = \left\{ \mathscr S_N+\mathscr R_N > z_0\,\varepsilon^{-1/2} \right\}. \end{align}\tag{22}\]
Lemma 1 (Estimates of \(\mathscr R_N\)). Let \(N\sim\varepsilon^{-\frac{1}{\nu}+\mu_1}\) be defined by 20 , and let \(\delta\) be as in 1. Then, for every \(\omega\notin\Omega_\delta\), \[\begin{align} \label{rne} \mathscr R_N = \mathcal{O}\!\left( \varepsilon^{-\frac{1}{2}+\frac{\mu_2}{2}} \right). \end{align}\tag{23}\]
Proof. Fix \(\omega\notin\Omega_\delta\). By Corollary 1, \[|c(\boldsymbol{n})|\,r_{\boldsymbol{n}}(\omega) = |c(\boldsymbol{n})g_{\boldsymbol{n}}(\omega)| \le \delta^{-1/2} \tnorm{\boldsymbol{n}}_{-(\rho-\kappa)}.\] Hence \[\begin{align} \mathscr R_N &= \sum_{|\boldsymbol{n}|>N} |c(\boldsymbol{n})|\,r_{\boldsymbol{n}}(\omega) \\ &\le \delta^{-1/2} \sum_{|\boldsymbol{n}|>N} \tnorm{\boldsymbol{n}}_{-(\rho-\kappa)} \\ &\lesssim \delta^{-1/2} \prod_{j=1}^{d}\prod_{j'=1}^{\nu_j} \sum_{|n_{j,j'}|>N} (1+|n_{j,j'}|)^{-(\rho_{j,j'}-\kappa_{j,j'})}. \end{align}\]
Since \(\rho_{j,j'}-\kappa_{j,j'}>2\), \[\sum_{|n_{j,j'}|>N} (1+|n_{j,j'}|)^{-(\rho_{j,j'}-\kappa_{j,j'})} \lesssim N^{-(\rho_{j,j'}-\kappa_{j,j'}-1)}.\] Therefore \[\mathscr R_N \lesssim \delta^{-1/2} N^{-\eta_2},\] where \(\eta_2\) is defined by 10 .
Using \(N\sim\varepsilon^{-\frac{1}{\nu}+\mu_1}\), we obtain \[\mathscr R_N \lesssim \delta^{-1/2} \varepsilon^{\left(\frac{1}{\nu}-\mu_1\right) \eta_2}.\]
By assumption 9 on \(\delta\), \[\delta^{-1/2} \le \varepsilon^{-\frac{1}{2} -\left(\frac{1}{\nu}-\mu_1\right) \eta_2 +\frac{\mu_2}{2}}.\] Combining the above estimates, we conclude that \[\mathscr R_N \lesssim \varepsilon^{-\frac{1}{2}+\frac{\mu_2}{2}},\] which proves the claim. ◻
Remark 10. By Lemma 1, there exists a constant \(C>0\) such that \[\mathscr R_N \le C\,\varepsilon^{-\frac{1}{2}+\frac{\mu_2}{2}}.\] Define \[\widetilde{\Omega} = \left\{ \mathscr R_N \le C\,\varepsilon^{-\frac{1}{2}+\frac{\mu_2}{2}} \right\},\] and \[\mathscr A^{(2)}_\varepsilon = \left\{ \mathscr S_N > \bigl(z_0-C\varepsilon^{\frac{\mu_2}{2}}\bigr)\varepsilon^{-\frac{1}{2}} \right\}.\] Then \[\begin{align} \label{eq:i2} \mathscr A_{\varepsilon}^{(1)} \cap \widetilde{\Omega} \subset \mathscr A_\varepsilon^{(2)}, \end{align}\tag{24}\] and \[\begin{align} \label{eq:io} \Omega_\delta^{\,c} \subset \widetilde{\Omega}. \end{align}\tag{25}\]
Here we analyze the principal part \(\mathscr{S}_N\), which provides the dominant contribution to the upper bound. The analysis reduces to studying a sum of independent random variables after separating deterministic and random components.
Clearly, \[\mathscr A^{(2)}_\varepsilon=\left\{\mathscr S_N^2>(z_0-C\varepsilon^{\frac{\mu_2}{2}})^2 \varepsilon^{-1}\right\}.\] Applying Hölder’s inequality, we obtain \[\begin{align} \mathscr S_N^2 &\leq \sum_{\boldsymbol{n}\in\Lambda_N}\frac{1}{2}|c(\boldsymbol{n})|^2 \cdot \sum_{\boldsymbol{n}\in\Lambda_N}2r_{\boldsymbol{n}}^2 \nonumber\\ &\leq \left(\sum_{\boldsymbol{n}\in\mathbb{Z}^\nu}\frac{1}{2}|c(\boldsymbol{n})|^2\right)\cdot \left(\sum_{\boldsymbol{n}\in\Lambda_N}2r_{\boldsymbol{n}}^2\right). \end{align}\]
Remark 11. Set \[\begin{align} \mathscr A_\varepsilon^{(3)}&= \left\{\xi_N>\mathfrak{I}_\varepsilon\varepsilon^{-1}\right\}, \end{align}\] where \[\begin{align} \xi_N&=\sum_{\boldsymbol{n}\in\Lambda_N}2r_{\boldsymbol{n}}^2,\nonumber\\ \mathfrak{I}_\varepsilon &= \frac{\big(z_0 - C\varepsilon^{\frac{\mu_2}{2}}\big)^2}{\frac{1}{2}\sum_{\boldsymbol{n}\in\mathbb{Z}^\nu} |c(\boldsymbol{n})|^2}.\label{rate} \end{align}\tag{26}\] Clearly, \[\begin{align} \label{eq:i3} \mathscr A^{(2)}_\varepsilon\subset\mathscr A^{(3)}_{\varepsilon}. \end{align}\tag{27}\]
Lemma 2 (LDP for \(\mathscr A^{(3)}_{\varepsilon}\)). Let \(N \sim \varepsilon^{-\frac{1}{\nu}+\mu_1}\) be defined in 20 , we have \[\lim_{\varepsilon\to 0^+}\varepsilon\log\mathbb{P}(\mathscr A^{(3)}_{\varepsilon})=-\frac{z_0^2}{\sum_{\boldsymbol{n}\in\mathbb{Z}^\nu} |c(\boldsymbol{n})|^2}.\]
Proof. Step 1. Lower bound via pointwise control
Clearly, \(\mathscr A_\varepsilon^{(3)}\supset\{2r_{\boldsymbol{n}}^2 > \mathfrak{I}_\varepsilon \varepsilon^{-1}\}\), where \(2r_{\boldsymbol{n}}^2\sim\mathscr{E}(1/2)\). Hence we have \[\begin{align} \mathbb{P}(\mathscr A_\varepsilon^{(3)}) &\geq \mathbb{P}(2r_{\boldsymbol{n}}^2 > \mathfrak{I}_\varepsilon \varepsilon^{-1}) \\ &= \int_{\mathfrak{I}_\varepsilon \varepsilon^{-1}}^{\infty} \frac{1}{2} e^{-\frac{1}{2}x} \, dx \\ &= e^{-\frac{1}{2} \mathfrak{I}_\varepsilon \varepsilon^{-1}}. \end{align}\] Taking logarithms, multiplying by \(\varepsilon\), and then taking the lower limit with respect to \(\varepsilon\to0^+\), we obtain that \[\liminf_{\varepsilon\to0^+}\varepsilon\log\mathbb{P}(\mathscr A_\varepsilon^{(3)})\geq\liminf_{\varepsilon\to0^+}\left(-\frac{1}{2}\mathfrak{I}_\varepsilon\right) =-\frac{z_0^2}{\sum_{\boldsymbol{n}\in\mathbb{Z}^\nu}|c(\boldsymbol{n})|^2}.\]
Step 2. Upper bound via Chernoff bound
Since \(2 r_{\boldsymbol{n}}^2 \sim \chi^2_2 = \mathscr E(1/2)\), its moment generating function is given by \[\mathbb{E}\bigl[e^{\lambda \cdot 2 r_{\boldsymbol{n}}^2}\bigr] = (1 - 2\lambda)^{-1}, \qquad \lambda \in (0, \tfrac12).\] By independence and \(|\Lambda_N| = (2N+1)^\nu\), it follows that \[\mathbb{E}\bigl[e^{\lambda \xi_N}\bigr] = (1 - 2\lambda)^{-(2N+1)^\nu}.\]
Applying Lemma 5, we obtain \[\mathbb{P}(\mathscr A_\varepsilon^{(3)}) \le (1 - 2\lambda)^{-(2N+1)^\nu} e^{-\lambda \mathfrak{I}_\varepsilon\varepsilon^{-1}}.\]
To optimize over \(\lambda\in(0,1/2)\), set \[\log h(\lambda) = -(2N+1)^\nu \log(1 - 2\lambda) - \lambda \mathfrak{I}_\varepsilon\varepsilon^{-1}.\] Then \[\frac{d}{d\lambda} \log h(\lambda) = \frac{2(2N+1)^\nu}{1 - 2\lambda} - \mathfrak{I}_\varepsilon\varepsilon^{-1}.\] Solving \(\frac{d}{d\lambda} \log h(\lambda)=0\) yields \[\lambda^* = \frac{1}{2}\left(1 - \frac{2(2N+1)^\nu}{\mathfrak{I}_\varepsilon\varepsilon^{-1}}\right),\] which lies in \((0,1/2)\) provided that \(\mathfrak{I}_\varepsilon\varepsilon^{-1} > 2(2N+1)^\nu\). This can be achieved by choosing an appropriate scaling of \(N\) with respect to \(\varepsilon\); see 20 .
Thus we have \[\mathbb{P}(\mathscr A_\varepsilon^{(3)}) \le (1 - 2\lambda^*)^{-(2N+1)^\nu} e^{-\lambda^* \mathfrak{I}_\varepsilon\varepsilon^{-1}}.\] Taking logarithms and multiplying both sides by \(\varepsilon\) gives \[\varepsilon \log \mathbb{P}(\mathscr A_\varepsilon^{(3)}) \le -\lambda^*\mathfrak{I}_\varepsilon - \varepsilon(2N+1)^\nu \log(1-2\lambda^*).\] Substituting \(\lambda^*\) and simplifying, \[\varepsilon \log \mathbb{P}(\mathscr A_\varepsilon^{(3)}) \le -\frac{1}{2}\mathfrak{I}_\varepsilon + \varepsilon(2N+1)^\nu - \varepsilon(2N+1)^\nu \log\!\left(\frac{2 \varepsilon (2N+1)^\nu}{\mathfrak{I}_\varepsilon}\right).\]
As \(\varepsilon \to 0^+\), recalling that \(N \sim \varepsilon^{-\frac{1}{\nu}+\mu}\) with \(0<\mu<1/\nu\) as in 20 , we have \[\mathfrak{I}_\varepsilon \to \frac{z_0^2}{\frac{1}{2}\sum_{\boldsymbol{n}\in\mathbb{Z}^\nu}|c(\boldsymbol{n})|^2}, \quad \varepsilon(2N+1)^\nu\to 0, \quad \text{and} \quad \varepsilon (2N+1)^\nu\log\!\left(\frac{2 \varepsilon (2N+1)^\nu }{\mathfrak{I}_\varepsilon}\right)\to 0.\] These imply that \[\limsup_{\varepsilon\to0^+} \varepsilon \log \mathbb{P}(\mathscr A_\varepsilon^{(3)}) \le -\frac{z_0^2}{\sum_{\boldsymbol{n}\in\mathbb{Z}^\nu}|c(\boldsymbol{n})|^2}.\] Proof of 2 is complete. ◻
It follows from 19 , 22 , 24 , 25 , 27 and the law of total probability that \[\begin{align} \mathbb{P}(\mathscr A_\varepsilon^\infty)\le\mathbb{P}(\mathscr A_\varepsilon^{(1)}) &=\mathbb{P}(\mathscr A_\varepsilon^{(1)}\cap\widetilde{\Omega})+ \mathbb{P}(\mathscr A_\varepsilon^{(1)}\cap\widetilde{\Omega}^c)\\ &\le\mathbb{P}(\mathscr A_\varepsilon^{(2)})+\mathbb{P}(\widetilde{\Omega}^c)\\ &\le\mathbb{P}(\mathscr A_\varepsilon^{(3)})+\mathbb{P}(\Omega_\delta)\\ &=\mathbb{P}(\mathscr A_\varepsilon^{(3)})\left(1+\frac{\mathbb{P}(\Omega_\delta)}{\mathbb{P}(\mathscr A_\varepsilon^{(3)})}\right). \end{align}\]
Taking logarithms and multiplying by \(\varepsilon\), we have \[\begin{align} \label{ee} \varepsilon\log\mathbb{P}(\mathscr A_\varepsilon^\infty)\leq\varepsilon\log\mathbb{P}(\mathscr A^{(3)}_\varepsilon)+\varepsilon\log\left(1+\frac{\mathbb{P}(\Omega_\delta)}{\mathbb{P}(\mathscr A_\varepsilon^{(3)})}\right). \end{align}\tag{28}\] The first term satisfies a large deviation principle (see 2), namely \[\mathbb{P}(\mathscr A_\varepsilon^{(3)}) = \exp\!\left(-\frac{\mathfrak I_0}{2\varepsilon} + o(\varepsilon^{-1})\right).\] The second term is controlled by 2, namely \[\mathbb{P}(\Omega_\delta) \lesssim \exp(-\delta^{-1}),\quad \delta\sim\varepsilon^{1+\eta}<\varepsilon,\] where \(\eta\) is specified in 1. Thus, \[\lim_{\varepsilon \to 0^+} \varepsilon \log\!\left(1 + \frac{\mathbb{P}(\Omega_\delta)}{\mathbb{P}(\mathscr A_\varepsilon^{(3)})}\right) = 0.\]
Taking the upper limit with respect to \(\varepsilon\to0^+\) on both sides of 28 , we have \[\begin{align} \limsup_{\varepsilon \to 0^+} \varepsilon \log \mathbb{P}(\mathscr A_\varepsilon^\infty) &\le -\frac{z_0^2}{\sum_{\boldsymbol{n}\in\mathbb{Z}^\nu} |c(\boldsymbol{n})|^2}, \end{align}\] which agrees with the right-hand side of 16 . This completes the proof of 2.
In 2, we established the linear LDP, which serves as a baseline for the nonlinear analysis. Before turning to the nonlinear case, we first derive in 3 a suitable well-posedness result that captures the relevant time scale, along with the necessary estimates.
Theorem 3 (Existence, Uniqueness and Fourier Decay).
Consider the randomized quasi-periodic Cauchy problem [eq:nls]–[eq:initialdata] together with 1. Then there exists a positive time \[{ {T_\varepsilon = \mathcal{O}\left(\varepsilon^{-\alpha+1+\eta}\right)}, }\] where \(\eta\) is chosen in accordance with 9 , and a randomized spatially quasi-periodic solution of the form \[\begin{align} \label{eq:solution2} {u}(t,\boldsymbol{x}) = \sum_{\boldsymbol{n}\in\mathbb{Z}^\nu} {c}(t,\boldsymbol{n}) e^{\mathrm{i}\langle \boldsymbol{n}\boldsymbol{\Omega}, \boldsymbol{x}\rangle} \end{align}\tag{29}\] defined on \([0,T_\varepsilon]\), with Fourier coefficients satisfying the uniform-in-time decay estimates \[\begin{align} \label{eq:decay2} \sup_{t\in[0,T_\varepsilon]} |{c}(t,\boldsymbol{n})| \lesssim {\varepsilon^{-\frac{1}{2}-\frac{\eta}{2}}}\tnorm{\boldsymbol{n}}_{-\frac{\rho-\kappa}{2}}. \end{align}\tag{30}\]
The proof of 3 is based on the strategy developed in [29]–[31]. We outline the main steps, emphasizing the key modifications.
It follows from the orthogonality of \(\{e^{\mathrm{i}\langle\boldsymbol{n}\Omega,\boldsymbol{x}\rangle}\}_{\boldsymbol{n}\in\mathbb{Z}^\nu}\) (see [31]) that, in Fourier space, [eq:nls] reads \[\begin{align} \label{eq:fourier95nls} \partial_t{c}(t,\boldsymbol{n}) = -\mathrm{i} Q(\boldsymbol{n}){c}(t,\boldsymbol{n}) + \mathrm{i}{\varepsilon^\alpha} \sum_{\substack{\boldsymbol{n}^j \in \mathbb{Z}^\nu,\, j=1,2,3\\ \boldsymbol{n}^1 - \boldsymbol{n}^2 + \boldsymbol{n}^3 = \boldsymbol{n}}} \prod_{j=1}^3 \bigl\{{c}(t,\boldsymbol{n}^j)\bigr\}^{\ast^{[j-1]}}. \end{align}\tag{31}\]
The initial condition in Fourier space is given by \[\begin{align} \label{eq:initialdata95fourier} {c}(0,\boldsymbol{n})={c}(\boldsymbol{n})g_{\boldsymbol{n}}. \end{align}\tag{32}\]
For each \(\boldsymbol{n} \in \mathbb{Z}^\nu\), the Duhamel formulation reads \[\begin{align} \label{eq:duhamel95c} {c}(t,\boldsymbol{n}) = e^{-\mathrm{i}Q(\boldsymbol{n})t}{c}(\boldsymbol{n})g_{\boldsymbol{n}} + \mathrm{i}{\varepsilon^\alpha} \int_0^t e^{-\mathrm{i}Q(\boldsymbol{n})(t-s)} \sum_{\substack{\boldsymbol{n}^j \in \mathbb{Z}^\nu,\, j=1,2,3\\ \boldsymbol{n}^1 - \boldsymbol{n}^2 + \boldsymbol{n}^3 = \boldsymbol{n}}} \prod_{j=1}^{3} \bigl\{{c}(s,\boldsymbol{n}^j)\bigr\}^{\ast^{[j-1]}} \,\mathrm{d}s. \end{align}\tag{33}\]
Define the Picard iteration \(\{{c}_k(t,\boldsymbol{n})\}_{k\ge0}\) by \[\begin{align} \label{eq:initialguess} {c}_0(t,\boldsymbol{n}) = e^{-\mathrm{i}Q(\boldsymbol{n})t}{c}(\boldsymbol{n})g_{\boldsymbol{n}}, \end{align}\tag{34}\] and for \(k\ge1\), \[\begin{align} \label{eq:picard} {c}_k(t,\boldsymbol{n}) = {c}_0(t,\boldsymbol{n}) + \mathrm{i}{\varepsilon^\alpha} \int_0^t e^{-\mathrm{i}Q(\boldsymbol{n})(t-s)} \sum_{\substack{\boldsymbol{n}^j \in \mathbb{Z}^\nu,\, j=1,2,3\\ \boldsymbol{n}^1 - \boldsymbol{n}^2 + \boldsymbol{n}^3 = \boldsymbol{n}}} \prod_{j=1}^{3} \bigl\{{c}_{k-1}(s,\boldsymbol{n}^j)\bigr\}^{\ast^{[j-1]}} \,\mathrm{d}s. \end{align}\tag{35}\]
In this subsection, we represent the Picard iteration using tree expansions, in order to systematically encode the combinatorial structure of the Duhamel terms. This representation allows us to track the multilinear interactions arising from the nonlinear evolution; see 6.
We begin by illustrating the decomposition of the linear and nonlinear contributions (see 2 and 3). We then briefly recall the combinatorial formalism developed in [30]–[32], which provides the basis for the tree representation.
Definition 1. The branch set \(\Gamma^{(k)}\) is defined recursively by \[\begin{align} \Gamma^{(k)} = \begin{cases} \{0,1\}, & k = 1; \\ \{0\} \cup \bigl(\Gamma^{(k-1)}\bigr)^3, & k \ge 2, \end{cases} \end{align}\] where \(0\) corresponds to a linear node and \(1\) corresponds to a cubic interaction node.
Definition 2. The first counting function \(\sigma\) on \(\Gamma^{(k)}\) is defined by \[\begin{align} \sigma(\gamma^{(k)})= \begin{cases} \frac{1}{2}, & \gamma^{(k)}=0, \;k\ge1; \\ \frac{3}{2}, & \gamma^{(1)}=1; \\ \sum_{j=1}^{3}\sigma(\gamma_j^{(k-1)}), & \gamma^{(k)}=(\gamma_j^{(k-1)})_{j=1}^3,\;k\ge2. \end{cases} \end{align}\]
Definition 3. The second counting function \(\ell\) is defined by \[\begin{align} \ell(\gamma^{(k)})= \begin{cases} 0, & \gamma^{(k)}=0, \;k\ge1; \\ 1, & \gamma^{(1)}=1; \\ 1+\sum_{j=1}^{3}\ell(\gamma_j^{(k-1)}), & \gamma^{(k)}=(\gamma_j^{(k-1)})_{j=1}^3,\;k\ge2. \end{cases} \end{align}\]
Proposition 3. For all \(k\ge1\), we have \[\begin{align} \sigma(\gamma^{(k)})=\ell(\gamma^{(k)})+\frac{1}{2}. \end{align}\]
Definition 4 (Domain). Define recursively \[\begin{align} {\mathrm D}^{(k,\gamma^{(k)})}= \begin{cases} (\mathbb{Z}^{\nu}) , & \gamma^{(k)}=0;\\ (\mathbb{Z}^{\nu})^3, & \gamma^{(1)}=1;\\ \prod_{j=1}^3 {\mathrm D}^{(k-1,\gamma_j^{(k-1)})}, & \gamma^{(k)}\in(\Gamma^{(k-1)})^3. \end{cases} \end{align}\]
Proposition 4. Let \(\dim_{\mathbb{Z}^\nu}\mathrm D^{(k,\gamma^{(k)})}\) denote the number of \(\mathbb{Z}^\nu\)-components of \(\mathrm D^{(k,\gamma^{(k)})}\). Then \[\dim_{\mathbb{Z}^{\nu}} \mathrm D^{(k,\gamma^{(k)})} = 2\sigma(\gamma^{(k)}).\] That is, \[\begin{align} \mathrm D^{(k,\gamma^{(k)})} \cong (\mathbb{Z}^{\nu})^{2\sigma(\gamma^{(k)})}. \end{align}\]
Proposition 5. The domain \({\mathrm D}^{(k,\gamma^{(k)})}\) admits a canonical identification with \[\begin{align} {\mathscr D}^{(k,\gamma^{(k)})}= \begin{cases} \prod_{j=1}^d \mathbb{Z}^{\nu_j}, & \gamma^{(k)}=0,\;k\ge1;\\ \prod_{j=1}^d (\mathbb{Z}^{\nu_j})^3, & \gamma^{(1)}=1;\\ \prod_{j=1}^3 {\mathscr D}^{(k-1,\gamma_j^{(k-1)})}, & \gamma^{(k)}\in(\Gamma^{(k-1)})^3,\;k\ge2. \end{cases} \end{align}\] This identification is induced by coordinate-wise projection.
More precisely, for each \(j=1,\dots,d\), we define the projection \[\iota_{x_j}:\left(\prod_{j=1}^d \mathbb{Z}^{\nu_j}\right)^s \to (\mathbb{Z}^{\nu_j})^s, \quad \iota_{x_j}\big((q_1^1,\dots,q_d^1),\dots,(q_1^s,\dots,q_d^s)\big) = (q_j^1,\dots,q_j^s).\] The full map \(\iota=(\iota_{x_1},\dots,\iota_{x_d})\) induces the above identification.
Definition 5 (Alternating sum). We define the map \[\lambda = \mathfrak{c}\mathfrak{a}\mathfrak{s}\circ \iota = (\mathfrak{c}\mathfrak{a}\mathfrak{s}_{x_1}\circ \iota_{x_1}, \dots, \mathfrak{c}\mathfrak{a}\mathfrak{s}_{x_d}\circ \iota_{x_d}),\] where for each \(j=1,\dots,d\), the map \[\mathfrak{c}\mathfrak{a}\mathfrak{s}_{x_j}: \bigcup_{s=1}^\infty (\mathbb{Z}^{\nu_j})^s \to \mathbb{Z}^{\nu_j}\] is defined by \[\mathfrak{c}\mathfrak{a}\mathfrak{s}_{x_j}(q_1,\dots,q_s) := \sum_{m=1}^s (-1)^{m-1} q_m.\]
Proposition 6 (Tree representation of \({c}_{k}(t,\boldsymbol{n})\)).
For all \(k\geq1\), we have \[{c}_k(t,\boldsymbol{n}) = \sum_{\gamma^{(k)}\in\Gamma^{(k)}} \sum_{\substack{\spadesuit^{(k)}\in{\mathrm D}^{(k,\gamma^{(k)})}\\ \lambda(\spadesuit^{(k)})=\boldsymbol{n}}} \mathfrak C^{(k,\gamma^{(k)})}(\spadesuit^{(k)}) \mathfrak I^{(k,\gamma^{(k)})}(t,\spadesuit^{(k)}) \mathfrak F^{(k,\gamma^{(k)})}(\spadesuit^{(k)}),\] where for \(0\in\Gamma^{(k)}\), \[\begin{align} \mathfrak C^{(k,0)}(\spadesuit^{(k)}) &={c}(\lambda(\spadesuit^{(k)}))g_{\lambda(\spadesuit^{(k)})},\\ \mathfrak I^{(k,0)}(t,\spadesuit^{(k)}) &=e^{-\mathrm{i}Q(\lambda(\spadesuit^{(k)})) t},\\ \mathfrak F^{(k,0)}(\spadesuit^{(k)}) &=1; \end{align}\] for \(1\in\Gamma^{(1)}\), \[\begin{align} \mathfrak C^{(1,1)}(\spadesuit^{(1)}) &=\prod_{j=1}^3\{{c}(\boldsymbol{n}^j)g_{\boldsymbol{n}^j}\}^{\ast^{[j-1]}},\\ \mathfrak I^{(1,1)}(t,\spadesuit^{(1)}) &=\int_0^t e^{-\mathrm{i} Q(\lambda(\spadesuit^{(1)}))(t-s)} \prod_{j=1}^{3} \left\{e^{-\mathrm{i} Q(\boldsymbol{n}^j) s}\right\}^{\ast^{[j-1]}} \,\mathrm{d}s,\\ \mathfrak F^{(1,1)}(\spadesuit^{(1)}) &=\mathrm{i}{\varepsilon^\alpha} ; \end{align}\] and for \(\gamma^{(k)}=(\gamma_j^{(k-1)})_{j=1}^3\), \[\begin{align} \mathfrak C^{(k,\gamma^{(k)})}(\spadesuit^{(k)}) &=\prod_{j=1}^3 \left\{\mathfrak C^{(k-1,\gamma_j^{(k-1)})}(\spadesuit_j^{(k-1)})\right\}^{\ast^{[j-1]}},\\ \mathfrak I^{(k,\gamma^{(k)})}(t,\spadesuit^{(k)}) &=\int_0^t e^{-\mathrm{i}Q(\lambda(\spadesuit^{(k)}))(t-s)} \prod_{j=1}^3 \left\{\mathfrak I^{(k-1,\gamma_j^{(k-1)})} (s,\spadesuit_j^{(k-1)})\right\}^{\ast^{[j-1]}} \,\mathrm{d}s,\\ \mathfrak F^{(k,\gamma^{(k)})}(\spadesuit^{(k)}) &=\mathrm{i}{\varepsilon^\alpha} \prod_{j=1}^3 \left\{\mathfrak F^{(k-1,\gamma_j^{(k-1)})}(\spadesuit_j^{(k-1)})\right\}^{\ast^{[j-1]}}. \end{align}\]
Let \(\boldsymbol{n}\in\mathbb{Z}^\nu\) be fixed. We also expand it as \[(n_1,\cdots,n_d)\in\mathbb{Z}^{\nu_1}\times\cdots\times\mathbb{Z}^{\nu_d},\] and further write \[(n_{1,1},\cdots,n_{1,\nu_1};\cdots;n_{d,1},\cdots,n_{d,\nu_d})\in\underbrace{\mathbb{Z}\times\cdots\times\mathbb{Z}}_{\nu_1}\times\cdots\times\underbrace{\mathbb{Z}\times\cdots\times\mathbb{Z}}_{\nu_d}.\]
For \(\gamma^{(k)}=0\in\Gamma^{(k)}\) with \(k\geq1\) (see 2), we set \(\boldsymbol{m}^1=\boldsymbol{n}\), that is, \((m^1_1,\cdots,m^1_d)=(n_1,\cdots,n_d)\), or equivalently, \[\left(m^1_{1,1},\cdots,m^1_{1,\nu_1}; \cdots; m^1_{d,1},\cdots, m^1_{d,\nu_d}\right) = \left(n_{1,1},\cdots,n_{1,\nu_1}; \cdots; n_{d,1},\cdots, n_{d,\nu_d}\right).\]
For \(\gamma^{(k)}=(\gamma_1^{(k-1)},\gamma_2^{(k-1)},\gamma_3^{(k-1)})\in\Gamma^{(k)}\setminus\{0\}\), we write \(\boldsymbol{n}=\boldsymbol{n}^1-\boldsymbol{n}^2+\boldsymbol{n}^3\in\mathbb{Z}^\nu\) (see 3), and set \[\begin{align} \mathbb{Z}^\nu\ni\boldsymbol{m}^j= \begin{cases} \boldsymbol{n}^{1,j}, & 1\leq j\leq 2\sigma(\gamma_1^{(k-1)});\\[4pt] \boldsymbol{n}^{2,\,j-2\sigma(\gamma_1^{(k-1)})}, & 2\sigma(\gamma_1^{(k-1)})+1\le j\le 2\sigma(\gamma_1^{(k-1)})+2\sigma(\gamma_2^{(k-1)});\\[4pt] \boldsymbol{n}^{3,\,j-2\sigma(\gamma_1^{(k-1)})-2\sigma(\gamma_2^{(k-1)})}, & 2\sigma(\gamma_1^{(k-1)})+2\sigma(\gamma_2^{(k-1)})+1\le j\le 2\sigma(\gamma^{(k)}). \end{cases} \end{align}\] Equivalently, for each component \(s=1,\dots,d\), we have \[\begin{align} \mathbb{Z}^{\nu_{s}}\ni m_{s}^j= \begin{cases} n_{s}^{1,j}, & 1\leq j\leq 2\sigma(\gamma_1^{(k-1)});\\[4pt] n_{s}^{2,\,j-2\sigma(\gamma_1^{(k-1)})}, & 2\sigma(\gamma_1^{(k-1)})+1\le j\le 2\sigma(\gamma_1^{(k-1)})+2\sigma(\gamma_2^{(k-1)});\\[4pt] n_{s}^{3,\,j-2\sigma(\gamma_1^{(k-1)})-2\sigma(\gamma_2^{(k-1)})}, & 2\sigma(\gamma_1^{(k-1)})+2\sigma(\gamma_2^{(k-1)})+1\le j\le 2\sigma(\gamma^{(k)}), \end{cases} \end{align}\] where \(m_{s}^j=\left(m^j_{s,1},\cdots,m^j_{s,\nu_{s}}\right)\in\mathbb{Z}^{\nu_{s}}\).
Thus \[\widetilde{\spadesuit}^{(k)}_{x_s}=(m_s^j)_{1\le j\le2\sigma(\gamma^{(k)})},\qquad 1\le s\le d,\quad k\geq1.\]
Proposition 7 (Estimates of \(\mathfrak C,\mathfrak I\) and \(\mathfrak F\)). For all \(k\geq1\), the following estimates hold:
\[\begin{align} \label{ce} \mathfrak C^{(k,\gamma^{(k)})}(\spadesuit^{(k)}) =\prod_{j=1}^{2\sigma(\gamma^{(k)})}\left\{c(\boldsymbol{m}^j)g_{\boldsymbol{m}^j}\right\}^{\ast^{[j-1]}}. \end{align}\tag{36}\] Moreover, under the decay property 11 , it holds that \[\begin{align} \label{ece} |\mathfrak C^{(k,\gamma^{(k)})}(\spadesuit^{(k)})|&\leq {\varepsilon^{-(1+\eta)\sigma(\gamma^{(k)})}}\prod_{j=1}^{2\sigma(\gamma^{(k)})}\tnorm{\boldsymbol{m}^j}_{-(\rho-\kappa)}\nonumber\\ &={\varepsilon^{-(1+\eta)\sigma(\gamma^{(k)})}} \prod_{j=1}^{2\sigma(\gamma^{(k)})}\prod_{s=1}^{d}\prod_{q=1}^{\nu_s}\left(1+|m^{j}_{s,q}|\right)^{-(\rho_{s,q}-\kappa_{s,q})}; \end{align}\tag{37}\]
\[\begin{align} \label{ie} |\mathfrak I^{(k,\gamma^{(k)})}(t,\spadesuit^{(k)})|\leq\frac{t^{\ell(\gamma^{(k)})}}{\mathfrak D(\gamma^{(k)})}, \end{align}\tag{38}\] where \[\begin{align} \mathfrak D(\gamma^{(k)}) = \begin{cases} 1, & \gamma^{(k)}=0\in\Gamma^{(k)},\;k\geq1,\\[4pt] 1, & \gamma^{(1)}\in\Gamma^{(1)},\\[4pt] \displaystyle \ell(\gamma^{(k)})\prod_{j=1}^3\mathfrak D(\gamma_j^{(k-1)}), & \gamma^{(k)}=(\gamma_j^{(k-1)})\in(\Gamma^{(k-1)})^3; \end{cases} \end{align}\]
\[\begin{align} \label{fe} |\mathfrak F^{(k,\gamma^{(k)})}(\spadesuit^{(k)})|\leq {\varepsilon^{\alpha\ell(\gamma^{(k)})}}. \end{align}\tag{39}\]
Proof. These estimates are proved by induction, following the argument in [31]. ◻
Lemma 3 (Uniform-in-time polynomial decay of \({c}_k(t,\boldsymbol{n})\)). On the complement of \(\Omega_\delta\) with \(\delta\sim\varepsilon^{1+\eta}\), for all \(k\geq1\), it holds that \[\begin{align} \label{eq:decayt} \sup_{t\in[0,T_\varepsilon]}|{c}_k(t,\boldsymbol{n})|\lesssim {\varepsilon^{-\frac{1}{2}-\frac{\eta}{2}}}\tnorm{\boldsymbol{n}}_{-\frac{\rho-\kappa}{2}} \end{align}\tag{40}\]
Proof. It follows from 6 and 7 that
\[\begin{align} |c_k(t,\boldsymbol{n})|\leq {\varepsilon^{-\frac{1}{2}-\frac{\eta}{2}}} \sum_{\gamma^{(k)}\in\Gamma^{(k)}}\frac{({\varepsilon^{\alpha-1-\eta}} t)^{\ell(\gamma^{(k)})}}{\mathfrak D(\gamma^{(k)})} \sum_{\substack{\lambda(\spadesuit^{(k)})=\boldsymbol{n}\\\spadesuit^{(k)}\in\mathrm D^{(k,\gamma^{(k)})}}}\prod_{j=1}^{2\sigma(\gamma^{(k)})} \prod_{s=1}^{d}\prod_{q=1}^{\nu_s}\left(1+|m^{j}_{s,q}|\right)^{-(\rho_{s,q}-\kappa_{s,q})}. \end{align}\]
We split the weight into two factors, each carrying exponent \(-\frac{\rho_{s,q}-\kappa_{s,q}}{2}\).
On the one hand, we estimate one factor using the Bernoulli’s inequality and obtain the following result: \[\begin{align} \prod_{j=1}^{2\sigma(\gamma^{(k)})} \prod_{s=1}^{d}\prod_{q=1}^{\nu_s}\left(1+|m^{j}_{s,q}|\right)^{-(\rho_{s,q}-\kappa_{s,q})} &=\prod_{s=1}^{d}\prod_{q=1}^{\nu_s} \left\{\prod_{j=1}^{2\sigma(\gamma^{(k)})} \left(1+|m^{j}_{s,q}|\right)\right\}^{-\frac{\rho_{s,q}-\kappa_{s,q}}{2}}\\ &\leq\prod_{s=1}^{d}\prod_{q=1}^{\nu_s} \left\{1+\sum_{j=1}^{2\sigma(\gamma^{(k)})} |m^{j}_{s,q}|\right\}^{-\frac{\rho_{s,q}-\kappa_{s,q}}{2}}\\ &\leq\prod_{s=1}^{d}\prod_{q=1}^{\nu_s} \left\{1+\left|\sum_{j=1}^{2\sigma(\gamma^{(k)})} (-1)^{j-1}m^{j}_{s,q}\right|\right\}^{-\frac{\rho_{s,q}-\kappa_{s,q}}{2}}\\ &=\prod_{s=1}^{d}\prod_{q=1}^{\nu_s} \left(1+|n_{s,q}|\right)^{-\frac{\rho_{s,q}-\kappa_{s,q}}{2}}\\ &=\tnorm{\boldsymbol{n}}_{-\frac{\rho-\kappa}{2}}. \end{align}\] Here we use the condition \(\lambda(\spadesuit^{(k)})=\boldsymbol{n}\), that is, \[\sum_{j=1}^{2\sigma(\gamma^{(k)})}(-1)^{j-1}m^j_{s,q}=n_{s,q},\qquad 1\leq s\leq d\quad\text{and}\quad1\leq q\leq\nu_s.\]
On the other hand, we have \[\begin{align} \sum_{\substack{\lambda(\spadesuit^{(k)})=\boldsymbol{n}\\\spadesuit^{(k)}\in\mathrm D^{(k,\gamma^{(k)})}}} \prod_{j=1}^{2\sigma(\gamma^{(k)})} \prod_{s=1}^{d}\prod_{q=1}^{\nu_s}\left(1+|m^{j}_{s,q}|\right)^{-\frac{\rho_{s,q}-\kappa_{s,q}}{2}} &\leq \prod_{j=1}^{2\sigma(\gamma^{(k)})} \prod_{s=1}^{d}\prod_{q=1}^{\nu_s}\sum_{m_{s,q}^{j}\in\mathbb{Z}}\left(1+|m^{j}_{s,q}|\right)^{-\frac{\rho_{s,q}-\kappa_{s,q}}{2}}\\ &\leq \prod_{j=1}^{2\sigma(\gamma^{(k)})} \prod_{s=1}^{d}\prod_{q=1}^{\nu_s}\sum_{m_{s,q}^{j}\in\mathbb{Z}}\left(1+|m^{j}_{s,q}|\right)^{-\frac{(\rho-\kappa)_{\min}}{2}}\\ &\leq \prod_{j=1}^{2\sigma(\gamma^{(k)})} \prod_{s=1}^{d}\prod_{q=1}^{\nu_s} \left\{1+2\mathfrak{b}\left(\frac{(\rho-\kappa)_{\min}}{2}\right)\right\}\\ &\leq \mathfrak{B}^{2\sigma(\gamma^{(k)})\nu}\\ &= \mathfrak{B}^{(2\ell(\gamma^{(k)})+1)\nu}\\ &= \mathfrak{B}^{\nu} \cdot \left(\mathfrak{B}^{2\nu}\right)^{\ell(\gamma^{(k)})}, \end{align}\] where \[\begin{align} (\rho-\kappa)_{\text{min}}&=\min_{\substack{1\leq j\leq d\\1\le j^\prime\le\nu_j}}\rho_{j,j^\prime}-\kappa_{j,j^\prime};\\ \mathfrak{b}(s) &= \frac{s}{s-1} \ge \zeta(s) = \sum_{n=1}^{\infty} \frac{1}{n^s}, \quad s > 1;\\ \mathfrak{B} &= 1 + 2\mathfrak{b}\left(\frac{(\rho-\kappa)_{\min}}{2}\right). \end{align}\]
Thus \[\big|c_k(t,\boldsymbol{n})\big| \lesssim {\varepsilon^{-\frac{1}{2}-\frac{\eta}{2}}} \tnorm{\boldsymbol{n}}_{-\frac{\rho-\kappa}{2}} \sum_{\gamma^{(k)}\in\Gamma^{(k)}} \frac{({{\varepsilon^{\alpha-1-\eta}}\mathfrak B^{2\nu}} t)^{\ell(\gamma^{(k)})}}{\mathfrak D(\gamma^{(k)})}.\]
It follows from [32] that the above series over \(\Gamma^{(k)}\) is uniformly bounded provided that \[0<t\leq\frac{4}{27}\mathfrak B^{-2\nu}{\varepsilon^{-\alpha+1+\eta}}\triangleq T_{\varepsilon}\lesssim{\varepsilon^{-\alpha+1+\eta}}.\]
Consequently, 40 holds. ◻
In this section, we combine the linear LDP (1) with the well-posedness result (3) to establish the nonlinear LDP (1). The main idea is to transfer the large deviations estimates from the linearized equation to the full nonlinear dynamics via a suitable perturbative argument. In particular, the well-posedness theory guarantees that the nonlinear solution can be constructed globally in the regime under consideration, which allows us to compare it with the corresponding linear evolution and control the nonlinear error terms; see 4.
Lemma 4 (Duhamel estimates). Let \(u\) be the nonlinear solution to the Cauchy problem [eq:nls]–[eq:initialdata], and let \(u_{\mathrm{linear}}\) denote the corresponding linearized solution (see 14 ). On the complement of \(\Omega_\delta\) with \(\delta\sim\varepsilon^{1+\eta}\), we have \[\|u(t)-u_{\mathrm{linear}}(t)\|_{L_{\boldsymbol{x}}^\infty(\mathbb{R}^d)} =\mathcal{O}\left(\varepsilon^{-\frac{1}{2}+\frac{\mu_3}{2}}\right), \quad t\sim\varepsilon^{-\alpha+1+\frac{3}{2}\eta+\frac{\mu_3}{2}}.\]
Proof. We begin by writing \[\begin{align} \| u(t)- u_{\mathrm{linear}}(t)\|_{L_x^\infty(\mathbb{R}^d)} &\leq \sum_{\boldsymbol{n}\in\mathbb{Z}^\nu} |(c - c_0)(t,\boldsymbol{n})|. \end{align}\] Using the Duhamel expansion for the nonlinear evolution, we obtain \[\begin{align} \| u(t)- u_{\mathrm{linear}}(t)\|_{L_x^\infty(\mathbb{R}^d)} &\leq {\varepsilon^{\alpha}} \int_0^t \sum_{\boldsymbol{n}\in\mathbb{Z}^\nu} \sum_{\substack{\boldsymbol{n}^j\in\mathbb{Z}^\nu,\; j=1,2,3\\ \boldsymbol{n}^1-\boldsymbol{n}^2+\boldsymbol{n}^3=\boldsymbol{n}}} \prod_{j=1}^3 |c(s,\boldsymbol{n}^j)|\, \mathrm{d}s. \end{align}\] We then apply the decay structure 40 of the coefficients and obtain \[\begin{align} \| u(t)- u_{\mathrm{linear}}(t)\|_{L_x^\infty(\mathbb{R}^d)} &\lesssim {\varepsilon^{\alpha-\frac{3}{2}-\frac{3}{2}\eta}}\, \varepsilon^{-\alpha+1+\frac{3}{2}\eta+\frac{\mu_3}{2}} \prod_{j=1}^3 \sum_{\boldsymbol{n}^j\in\mathbb{Z}^\nu} \tnorm{\boldsymbol{n}^j}_{-\frac{\rho-\kappa}{2}}\\ &\lesssim {\varepsilon^{-\frac{1}{2}+\frac{\mu_3}{2}}}\, \prod_{j=1}^3 \sum_{\boldsymbol{n}^j\in\mathbb{Z}^\nu} \tnorm{\boldsymbol{n}^j}_{-\frac{\rho-\kappa}{2}}. \end{align}\] Since the weighted sums are finite under the condition \[\rho_{j,j'} - \kappa_{j,j'}>2, \quad j^\prime=1,\cdots,\nu_j, j=1,\cdots,d,\] we conclude that \[\| u(t)- u_{\mathrm{linear}}(t)\|_{L_x^\infty(\mathbb{R}^d)} \lesssim \varepsilon^{-\frac{1}{2}+\frac{\mu_3}{2}}.\] ◻
Remark 12. By 4, there exists a positive constant \(C\) such that \[\| u(t)- u_{\mathrm{linear}}(t)\|_{L_x^\infty(\mathbb{R}^d)} \le C\varepsilon^{-\frac{1}{2}+\frac{\mu_3}{2}}.\] Set \[\widehat{\Omega}=\left\{\| u(t)- u_{\mathrm{linear}}(t)\|_{L_x^\infty(\mathbb{R}^d)} \leq C\varepsilon^{-\frac{1}{2}+\frac{\mu_3}{2}}\right\}.\] Clearly, \[\begin{align} \label{eq:hatomega} \Omega_\delta^c\subset\widehat{\Omega}. \end{align}\tag{41}\]
Set \[\begin{align} \mathscr A_\varepsilon^- &= \left\{ \| u_{\text{linear}} \|_{L^\infty_x(\mathbb{R}^d)} - \| u - u_{\text{linear}} \|_{L^\infty_x(\mathbb{R}^d)} > z_0 \, \varepsilon^{-\frac{1}{2}} \right\},\\ \mathscr B_\varepsilon^- &= \left\{ \| u_{\text{linear}} \|_{L^\infty_x(\mathbb{R}^d)} > \left( z_0 + C \, \varepsilon^{\frac{\mu_3}{2}} \right) \varepsilon^{-\frac{1}{2}} \right\}. \end{align}\] Clearly, we have the following inclusion relation: \[\mathscr B_\varepsilon^-\cap\widehat{\Omega} \subset \mathscr A_\varepsilon^- \subset \mathscr A_\varepsilon^\infty.\]
By the law of total probability and 41 , we have \[\begin{align} \mathbb{P}(\mathscr A_\varepsilon^\infty) &\ge \mathbb{P}(\mathscr B_\varepsilon^-\cap\widehat{\Omega})\\ &= \mathbb{P}(\mathscr B_\varepsilon^-) -\mathbb{P}(\mathscr B_\varepsilon^-\cap\widehat{\Omega}^{\,c})\\ &\ge \mathbb{P}(\mathscr B_\varepsilon^-)-\mathbb{P}(\widehat{\Omega}^{\,c})\\ &\ge \mathbb{P}(\mathscr B_\varepsilon^-)-\mathbb{P}(\Omega_\delta)\\ &=\mathbb{P}(\mathscr B_\varepsilon^-)\left(1-\frac{\mathbb{P}(\Omega_\delta)}{\mathbb{P}(\mathscr B_\varepsilon^-)}\right). \end{align}\] Taking logarithms and multiplying by \(\varepsilon\), we have \[\begin{align} \label{w2} \varepsilon\log\mathbb{P}(\mathscr A_\varepsilon^\infty)\geq \varepsilon\log\mathbb{P}(\mathscr B_\varepsilon^-) +\varepsilon\log\left(1-\frac{\mathbb{P}(\Omega_\delta)}{\mathbb{P}(\mathscr B_\varepsilon^-)}\right). \end{align}\tag{42}\]
Applying 2, we deduce that \[\begin{align} \lim_{\varepsilon \to 0^+} \varepsilon \log\mathbb{P} \left(\mathscr B_\varepsilon^- \right) &= -\frac{z_0^2}{\sum_{\boldsymbol{n}\in\mathbb{Z}^\nu}| c(\boldsymbol{n})|^2}. \end{align}\]
Similar to the analysis in 2.3.5, \[\lim_{\varepsilon\to0}\varepsilon\log\left(1-\frac{\mathbb{P}(\Omega_\delta)}{\mathbb{P}(\mathscr B_\varepsilon^-)}\right)=0.\]
Therefore \[\begin{align} \label{eqlower} \liminf_{\varepsilon\to0^+}\varepsilon\log\mathbb{P}(\mathscr A_\varepsilon^\infty) \ge-\frac{z_0^2}{\sum_{\boldsymbol{n}\in\mathbb{Z}^\nu}| c(\boldsymbol{n})|^2}. \end{align}\tag{43}\]
Set \[\begin{align} \mathscr A_\varepsilon^+ &= \left\{ \| u_{\text{linear}} \|_{L^\infty_x(\mathbb{R}^d)} + \| u - u_{\text{linear}} \|_{L^\infty_x(\mathbb{R}^d)} > z_0 \, \varepsilon^{-\frac{1}{2}} \right\},\\ \mathscr B_\varepsilon^+ &= \left\{ \| u_{\text{linear}} \|_{L^\infty_x(\mathbb{R}^d)} > \left( z_0 - C \, \varepsilon^{\frac{\mu_3}{2}} \right) \varepsilon^{-\frac{1}{2}} \right\}. \end{align}\] Clearly, we have the following inclusion relation: \[\mathscr A_\varepsilon^\infty \subset \mathscr A_\varepsilon^+,\quad \mathscr A_\varepsilon^+\cap\widehat{\Omega} \subset \mathscr B_\varepsilon^+.\]
It follows from the law of total probability and 41 that \[\begin{align} \mathbb{P}(\mathscr A_\varepsilon^\infty) &\le \mathbb{P}(\mathscr A_\varepsilon^+)\\ &=\mathbb{P}(\mathscr A_\varepsilon^+\cap\widehat{\Omega}) +\mathbb{P}(\mathscr A_\varepsilon^+\cap\widehat{\Omega}^c)\\ &\le \mathbb{P}(\mathscr B_\varepsilon^+)+\mathbb{P}(\widehat{\Omega}^c)\\ &\le \mathbb{P}(\mathscr B_\varepsilon^+)+\mathbb{P}(\Omega_\delta)\\ &=\mathbb{P}(\mathscr B_\varepsilon^+)\left(1+\frac{\mathbb{P}(\Omega_\delta)}{\mathbb{P}(\mathscr B_\varepsilon^+)}\right). \end{align}\] Taking logarithms and multiplying by \(\varepsilon\), we have \[\begin{align} \label{w3} \varepsilon\log\mathbb{P}(\mathscr A_\varepsilon^\infty)\le \varepsilon\log\mathbb{P}(\mathscr B_\varepsilon^+) +\varepsilon\log\left(1+\frac{\mathbb{P}(\Omega_\delta)}{\mathbb{P}(\mathscr B_\varepsilon^+)}\right). \end{align}\tag{44}\]
Applying 2, we deduce that \[\begin{align} \lim_{\varepsilon \to 0^+} \varepsilon \log \mathbb{P}\left(\mathscr B_\varepsilon^+ \right) &= -\frac{z_0^2}{\sum_{\boldsymbol{n}\in\mathbb{Z}^\nu}| c(\boldsymbol{n})|^2}. \end{align}\]
Analogously, \[\lim_{\varepsilon\to0}\varepsilon\log\left(1+\frac{\mathbb{P}(\Omega_\delta)}{\mathbb{P}(\mathscr B_\varepsilon^+)}\right)=0.\]
Hence we obtain \[\begin{align} \label{equpper} \limsup_{\varepsilon\to0^+}\varepsilon\log\mathbb{P}(\mathscr A_\varepsilon^\infty) \le-\frac{z_0^2}{\sum_{\boldsymbol{n}\in\mathbb{Z}^\nu}| c(\boldsymbol{n})|^2}. \end{align}\tag{45}\]
Combining 43 and 45 yields that \[\begin{align} \lim_{\varepsilon\to0^+}\varepsilon\log\mathbb{P}\Big( \|u(t)\|_{L_{\boldsymbol{x}}^\infty(\mathbb{R}^d)} > z_0 \varepsilon^{-\frac{1}{2}} \Big) =-\frac{z_0^2}{\sum_{\boldsymbol{n}\in\mathbb{Z}^\nu}|c(\boldsymbol{n})|^2},\quad t\sim\varepsilon^{-\alpha+1+\frac{3}{2}\eta+\frac{\mu_3}{2}}. \end{align}\] This completes the proof of 1.
In this section, we review some fundamental notation and concepts from probability and statistics.
Let the triple \((\Omega, \mathcal{F}, \mathbb{P})\) be a probability space. The set \(\Omega\) denotes the sample space, comprising all possible outcomes, and \(\omega \in \Omega\) represents an individual outcome (sample point). The collection \(\mathcal{F}\) is a \(\sigma\)-algebra on \(\Omega\), referred to as the event space, whose elements are subsets of \(\Omega\); these events are typically denoted by uppercase letters such as \(A\), \(B\), \(C\). The function \(\mathbb{P} : \mathcal{F} \to [0,1]\) is a probability measure, assigning to each event \(A \in \mathcal{F}\) its probability \(\mathbb{P}(A)\).
Let \(\mathbb{K} \in \{\mathbb{R}, \mathbb{C}\}\) be a field. A \(\mathbb{K}\)-valued random variable \(X\) is a measurable function from the sample space to \(\mathbb{K}\), that is, \(X : \Omega \to \mathbb{K}\), \(\omega \mapsto X(\omega)\).
We now focus on real-valued random variables, i.e., the case \(\mathbb{K} = \mathbb{R}\). Events of the form \(\{\omega \in \Omega : X(\omega) \in D \subset \mathbb{R}\} := A\) are commonly abbreviated as \(\{X \in D\}\). Furthermore, \(\mathbb{P}(A) = \mathbb{P}(\{X \in D\}) \triangleq \mathbb{P}(X \in D)\).
The cumulative distribution function (CDF) of a random variable \(X\), denoted by \(F_X\), is defined as \(F_X(x) = \mathbb{P}(X \leq x)\) for all \(x \in \mathbb{R}\).
Under standard regularity conditions, such as the differentiability of \(F_X\) with respect to \(x\), the probability density function (PDF) of \(X\), denoted \(f_X\), is defined as the derivative \[f_X(x) = \frac{{\rm d}}{{\rm d}x} F_X(x).\]
The expectation, or mean, of a random variable \(X\), denoted \(\mathbb{E}[X]\), is the weighted average of \(X(\omega)\) (or \(X\)) using the probability measure \({\rm d}\mathbb{P}(\omega)\) (or PDF \(f_X\)) as weights, that is, \[\begin{align} \mathbb{E}[X] &= \int_\Omega X(\omega)\,{\rm d}\mathbb{P}(\omega) \quad (\text{using the probability measure in abstract space}), \\ &= \int_{\mathbb{R}} x\,{\rm d}F_X(x) \quad (\text{using the CDF via Stieltjes integral}), \\ &= \int_{\mathbb{R}} x f_X(x)\,{\rm d}x \quad (\text{using the PDF in the absolutely continuous case}). \end{align}\] We also denote this by \(\mu_X \triangleq \mu\).
The expectation satisfies linearity, that is, \[\mathbb{E}[a_1X_1 + a_2X_2] = a_1\mathbb{E}[X_1] + a_2\mathbb{E}[X_2], \quad \forall a_1,a_2 \in \mathbb{R}.\]
The variance of a random variable \(X\), denoted \(\mathbb{V}(X)\), is defined as the expected value of \((X-\mu_X)^2\), that is, \[\mathbb{V}(X) = \mathbb{E}[(X-\mu_X)^2].\] We also denote this by \(\sigma_X^2 \triangleq \sigma^2\). By linearity of expectation, \[\mathbb{V}(X) = \mathbb{E}[X^2] - \mu_X^2.\] Furthermore, \[\mathbb{V}(aX + b) = a^2 \mathbb{V}(X), \quad \forall a \in \mathbb{R}.\]
The characteristic function (CF) of \(X\), denoted \(\phi_X\), is defined by \[\phi_X(t) = \mathbb{E}[e^{{\rm i} t X}], \quad t \in \mathbb{R}.\] It should be noted that characteristic functions uniquely determine the distribution: \(\phi_X = \phi_Y\) if and only if \(X\) and \(Y\) have the same distribution. This follows from the Lévy continuity theorem; see [33].
In this subsection, we briefly introduce several important classical distributions relevant to studying rogue wave mechanisms from a probabilistic perspective. These distributions are repeatedly applied in the calculation of LDP for rogue waves.
We say that a random variable \(X\) follows the Gaussian/normal distribution with mean \(\mu\) and variance \(\sigma^2\) if its PDF is \[f_X(x;\mu,\sigma) = \frac{1}{\sqrt{2\pi}\sigma} e^{-\frac{(x-\mu)^2}{2\sigma^2}}, \quad x \in \mathbb{R}.\] The corresponding mean, variance and characteristic function are, respectively, \[\mathbb{E}[X] = \mu, \quad \mathbb{V}(X) = \sigma^2, \quad \phi_X(t) = e^{{\rm i}\mu t - \frac{1}{2}\sigma^2 t^2}.\]
Here and throughout, we write \(X \sim \mathscr{N}_\mathbb{R}(\mu,\sigma^2)\) to denote a real-valued random variable following this distribution.
If \(X \sim \mathscr{N}_\mathbb{R}(\mu,\sigma^2)\), then the normalization \(Y = \frac{X-\mu}{\sigma} \sim \mathscr{N}_\mathbb{R}(0,1)\). In this case, we say that \(Y\) follows the standard Gaussian/normal distribution, with PDF \(f_Y(y) = \frac{1}{\sqrt{2\pi}}e^{-\frac{1}{2}y^2}\) and CF \(\phi_Y(t) = e^{-\frac{1}{2}t^2}\).
We say that a random variable \(X\) follows the exponential distribution with rate parameter \(\lambda > 0\) if its PDF is \[f_X(x;\lambda) = \lambda e^{-\lambda x} \boldsymbol{1}_{[0,\infty)}(x).\]
Throughout, we write \(X \sim \mathscr{E}(\lambda)\) to denote that \(X\) follows this distribution.
If \(X \sim \mathscr{E}(\lambda)\), then \[\mathbb{E}[X] = \frac{1}{\lambda}, \quad \mathbb{V}(X) = \frac{1}{\lambda^2},\] and \(aX \sim \mathscr{E}\bigl(\frac{\lambda}{a}\bigr)\) for \(a > 0\).
We say that a random variable \(X\) follows the chi-squared distribution with \(n\) degrees of freedom if its PDF is \[f_X(x;n) = \frac{1}{\Gamma\left(\frac{n}{2}\right)2^{n/2}} x^{\frac{n}{2}-1} e^{-x/2} \boldsymbol{1}_{[0,\infty)}(x).\]
Throughout, we write \(X \sim \chi_n^2\) to denote this distribution.
If \(X \sim \chi_n^2\), then \[\mathbb{E}[X] = n, \quad \mathbb{V}(X) = 2n.\]
In addition, we have
\(\chi_2^2 = \mathscr{E}\bigl(\frac{1}{2}\bigr)\).
If \(X_1 \sim \chi_{n_1}^2\) and \(X_2 \sim \chi_{n_2}^2\) are independent, then \(X_1 + X_2 \sim \chi_{n_1+n_2}^2\).
If \(\{X_j\}_{1\leq j\leq n}\) are i.i.d. \(\mathscr{E}(\lambda)\), then \(2\lambda\sum_{j=1}^n X_j \sim \chi_{2n}^2\).
If \(\{X_j\}_{1\leq j\leq n}\) are i.i.d. \(\mathscr{N}_\mathbb{R}(0,1)\), then \(\sum_{j=1}^n X_j^2 \sim \chi_n^2\).
To introduce complex-valued random variables, we first review the concepts of joint distributions (see 5.3) and random independence (see 5.4), initially for two random variables. These concepts naturally generalize to any finite collection of random variables.
Let \(X\) and \(Y\) be random variables defined on a common probability space \((\Omega, \mathcal{F}, \mathbb{P})\). The joint CDF of \((X,Y)\), denoted \(F_{X,Y}\), is defined by \[F_{X,Y}(x,y) = \mathbb{P}(X \leq x, Y \leq y), \quad (x,y) \in \mathbb{R}^2.\]
Under standard regularity conditions (such as the existence of mixed partial derivatives), the joint PDF, denoted \(f_{X,Y}\), is given by \[f_{X,Y}(x,y) = \frac{\partial^2 F_{X,Y}(x,y)}{\partial x \partial y}.\]
Let \(f_X\) and \(f_Y\) denote the marginal PDFs of \(X\) and \(Y\). Then \[\begin{align} f_X(x) &= \int_{\mathbb{R}} f_{X,Y}(x,y)\,{\rm d}y = \frac{\partial F_{X,Y}(x,y)}{\partial x}\Big|_{y\to +\infty}, \\ f_Y(y) &= \int_{\mathbb{R}} f_{X,Y}(x,y)\,{\rm d}x = \frac{\partial F_{X,Y}(x,y)}{\partial y}\Big|_{x\to +\infty}. \end{align}\] Here, \(f_X\) and \(f_Y\) are called the marginal distributions.
The joint CF for the joint distribution of \((X,Y)\), denoted by \(\phi_{X,Y}\), is defined by \[\phi_{X,Y}(t_1,t_2) = \mathbb{E}[e^{{\rm i}(t_1X + t_2Y)}], \quad (t_1,t_2) \in \mathbb{R}^2.\]
Let \(A\) and \(B\) be events in the probability space \((\Omega, \mathcal{F}, \mathbb{P})\) with \(\mathbb{P}(B) > 0\). The conditional probability of \(A\) given \(B\), denoted \(\mathbb{P}(A|B)\), is defined by \[\mathbb{P}(A|B) = \frac{\mathbb{P}(A \cap B)}{\mathbb{P}(B)}.\] This yields the multiplication rule: \[\mathbb{P}(A \cap B) = \mathbb{P}(A|B) \mathbb{P}(B).\]
Events \(A\) and \(B\) are independent if \[\mathbb{P}(A|B) = \mathbb{P}(A),\] or equivalently, \[\mathbb{P}(A \cap B) = \mathbb{P}(A) \mathbb{P}(B).\] In this case, \(\mathbb{P}(B|A) = \mathbb{P}(B)\) also holds.
Let \(X\) and \(Y\) be random variables defined on a common probability space \((\Omega, \mathcal{F}, \mathbb{P})\). We say \(X\) and \(Y\) are independent if the events \(\{X \leq x\}\) and \(\{Y \leq y\}\) are independent for all \(x,y \in \mathbb{R}\). Specifically, \[\mathbb{P}(X \leq x, Y \leq y) = \mathbb{P}(X \leq x) \mathbb{P}(Y \leq y).\] That is, \[\begin{align} F_{X,Y}(x,y) &= F_X(x) F_Y(y) \quad (\text{using the CDF}), \\ f_{X,Y}(x,y) &= f_X(x) f_Y(y) \quad (\text{using the PDF}), \\ \phi_{X,Y}(t_1,t_2) &= \phi_X(t_1) \phi_Y(t_2) \quad (\text{using the CF}). \end{align}\]
From now on, we focus on complex-valued random variables, i.e., \(\mathbb{K} = \mathbb{C}\), particularly those following complex Gaussian distributions.
Let \(Z\) be a complex random variable on the probability space \((\Omega, \mathcal{F}, \mathbb{P})\), that is, \(Z : \Omega \to \mathbb{C}\), \(\omega \mapsto Z(\omega)\).
The distribution of \(Z\) admits two natural representations. The first, based on the Cartesian coordinate system, identifies it with the joint distribution of the real and imaginary parts \(\operatorname{Re}Z\) and \(\operatorname{Im}Z\). The second, based on the polar coordinate system, identifies it with the joint distribution of the modulus \(|Z|\) and argument \(\arg Z\).
In the Cartesian approach, write \(Z(\omega) = X(\omega) + {\rm i} Y(\omega)\) for \(\omega \in \Omega\), where \(X,Y\) are real-valued random variables on \((\Omega, \mathcal{F}, \mathbb{P})\). The distribution of \(Z\) is the joint distribution of \((X,Y)\) under the partial order \(\prec\) defined by \[z_1 \prec z_2 \Longleftrightarrow \operatorname{Re} z_1 \leq \operatorname{Re} z_2 \;\text{and}\; \operatorname{Im} z_1 \leq \operatorname{Im} z_2.\] For \(z = x + {\rm i} y \in \mathbb{C}\), the CDF is \[F_Z(z) = \mathbb{P}(Z \prec z) = \mathbb{P}(X \leq x, Y \leq y) = F_{X,Y}(x,y).\] Hence the PDF is given by \[f_Z(z)=f_{X,Y}(x, y)=\frac{\partial^2F_{X,Y}(x,y)}{\partial x\partial y};\] and for \(t=t_1+{\rm i}t_2\), where \(t_1,t_2\in\mathbb{R}\), the CF is given by3 \[\phi_Z(t) = \phi_{X,Y}(t_1, t_2) = \mathbb{E} \left[ e^{\mathrm{i}(t_1 X + t_2 Y)} \right], \quad (t_1, t_2) \in \mathbb{R}^2.\]
For the expectation of \(Z = X + {\rm i} Y\), we have \[\mathbb{E}[Z] = \mathbb{E}[X] + {\rm i} \mathbb{E}[Y].\] Whenever \(\mathbb{E}[Z]\) exists, expectation and complex conjugation commute, that is, \[\mathbb{E}[\overline{Z}] = \overline{\mathbb{E}[Z]}.\]
For the variance of \(Z = X + {\rm i} Y\), it is defined by \[\mathbb{V}(Z) = \mathbb{E}[|Z - \mu_Z|^2].\] Using \(|z|^2 = z\overline{z}\), we have \[\begin{align} \mathbb{V}(Z) &= \mathbb{E}\bigl[(Z-\mu_Z)(\overline{Z}-\overline{\mu_Z})\bigr] \\ &= \mathbb{E}[|Z|^2] - |\mu_Z|^2. \end{align}\] Using \(|z|^2 = x^2 + y^2\) with \(Z-\mu_Z = (X-\mu_X) + {\rm i}(Y-\mu_Y)\), we obtain \[\begin{align} \mathbb{V}(Z) &= \mathbb{E}[|X-\mu_X|^2 + |Y-\mu_Y|^2] \\ &= \mathbb{V}(X) + \mathbb{V}(Y). \end{align}\]
In the polar approach, write \(Z(\omega) = R(\omega) e^{{\rm i} \Theta(\omega)}\) for \(\omega \in \Omega\), where \(R(\omega) > 0\), \(0 \leq \Theta(\omega) < 2\pi\), and \(R, \Theta\) are real-valued random variables on \((\Omega, \mathcal{F}, \mathbb{P})\). The distribution of \(Z\) is then identified with the joint distribution of \((R, \Theta)\).
In fact, these two approaches are equivalent via the transformation \[X = R \cos \Theta, \quad Y = R \sin \Theta.\]
Here we focus on centered complex-valued Gaussian random variables, i.e., those with zero mean. Write \(Z = X + {\rm i}Y,\) where \(X \sim \mathscr{N}_\mathbb{R}(0,\sigma_1^2)\) and \(Y \sim \mathscr{N}_\mathbb{R}(0,\sigma_2^2)\) are independent. Then \(\mathbb{V}(Z) = \sigma_1^2 + \sigma_2^2,\) so \(Z \sim \mathscr{N}_\mathbb{C}(0,\sigma_1^2 + \sigma_2^2)\).
We pay special attention to the equi-variance case \(\sigma_1 = \sigma_2 = \sigma\), where \(Z \sim \mathscr{N}_\mathbb{C}(0,2\sigma^2)\). In particular, when \(2\sigma^2 = 1\), \(Z\) is called a standard complex Gaussian random variable; see [34].
Let \(Z = X + {\rm i}Y\), where \(X\) and \(Y\) are independent centered Gaussians with variance \(\sigma^2\), i.e., \(X,Y \sim \mathscr{N}_\mathbb{R}(0,\sigma^2)\). After normalization, \(X/\sigma, Y/\sigma \sim \mathscr{N}_\mathbb{R}(0,1)\), so \[\sigma^{-2}|Z|^2 = \sigma^{-2}(X^2 + Y^2) = \left(\frac{X}{\sigma}\right)^2 + \left(\frac{Y}{\sigma}\right)^2 \sim \chi_2^2 = \mathscr{E}\left(\frac{1}{2}\right),\] both having PDF \[f_{\sigma^{-2}|Z|^2}(x) = \frac{1}{2} e^{-x/2} \boldsymbol{1}_{[0,\infty)}(x).\]
Furthermore (see 5.2.2), \(|Z|^2\) follows the exponential distribution with rate parameter \(\frac{1}{2\sigma^2}\), that is, \[\label{square} |Z|^2 \sim \mathscr{E}\left(\frac{1}{2\sigma^2}\right).\tag{46}\]
Similarly, for the distribution of \(|Z|\), we have \[f_{|Z|}(x) = \frac{x}{\sigma^2} e^{-x^2/(2\sigma^2)} \boldsymbol{1}_{[0,\infty)}(x).\] This is the PDF of the so-called Rayleigh distribution with parameter \(\sigma\), denoted \(|Z| \sim \mathscr{R}(\sigma)\); see [35].
Let \(Z = R e^{{\rm i} \Theta}\). From the preceding discussion, the radius \(R\) follows the Rayleigh distribution with parameter \(\sigma\), i.e., \(R \sim \mathscr{R}(\sigma)\), and \(R^2\) follows the exponential distribution with rate parameter \(\frac{1}{2\sigma^2}\), i.e., \[R^2 \sim \mathscr{E}\left(\frac{1}{2\sigma^2}\right).\] Furthermore, the phase \(\Theta\) is uniformly distributed on \([0,2\pi)\), i.e., \(\Theta \sim \mathscr{U}[0,2\pi)\); see [1].
Lemma 5 (Chernoff bound). Let \(X\) be a nonnegative random variable. Then, for any \(a \ge 0\) and \(\lambda > 0\), \[\mathbb{P}(X \ge a) \le e^{-\lambda a}\,\mathbb{E}\bigl[e^{\lambda X}\bigr].\]
Proof. Fix any \(\lambda > 0\). Since \(x \mapsto e^{\lambda x}\) is strictly increasing, we have \[\{X \ge a\} = \{e^{\lambda X} \ge e^{\lambda a}\}.\] Applying Markov’s inequality to the nonnegative random variable \(e^{\lambda X}\), we obtain \[\begin{align} \mathbb{P}(X \ge a) &= \mathbb{P}(e^{\lambda X} \ge e^{\lambda a}) \\ &\le \frac{\mathbb{E}[e^{\lambda X}]}{e^{\lambda a}} \\ &= e^{-\lambda a}\,\mathbb{E}[e^{\lambda X}]. \end{align}\] This completes the proof of 5. ◻
The first author (F. Xu) was supported in part by the National Natural Science Foundation of China (Grant No. 12501235), the China Postdoctoral Science Foundation (Grant Nos. BX20240138 and 2025M783104), and the Jilin Provincial Postdoctoral Science Foundation (2025). He would like to thank David Damanik (Rice University) for sharing several papers on Wave Turbulence Theory, Gigliola Staffilani (MIT) for discussions on higher-order normal form, fundamental insights into the nature of the criticality (onset of the first nonlinear effect), and valuable comments on earlier versions of the paper, Yvain Bruned (Université de Lorraine), Yuzhao Wang (University of Birmingham) and Xinyu Zhao (New Jersey Institute of Technology) for related discussions.↩︎
The second author (Y. Li) is the corresponding author. He was supported in part by the National Basic Research Program of China (Grant No. 2013CB834100), and the National Natural Science Foundation of China (Grant Nos. 12471183 and 12531009).↩︎
In studies of complex random variables, the CF is often defined as \(\phi_Z(t) = \mathbb{E}\bigl[e^{{\rm i} \operatorname{Re}(\overline{t} Z)}\bigr]\) for \(t \in \mathbb{C}\). For \(t = t_1 + {\rm i} t_2\) and \(Z = X + {\rm i} Y\), this equals \(\mathbb{E}\bigl[e^{{\rm i}(t_1X + t_2Y)}\bigr]\), the joint CF of the real and imaginary parts.↩︎