May 28, 2026
We study the problem of assigning a scalar score to a short trajectory window that reflects its position on an ordered continuum of predictability regimes, spanning structured deterministic dynamics to unstructured stochastic noise. Existing methods address deterministic-versus-stochastic discrimination within a single system and do not produce scores with a consistent numerical interpretation across systems. We formalize this as ordinal estimation over a five-level predictability ladder and identify a structural source of cross-system ambiguity: ranking supervision alone leaves the score coordinate unfixed up to a monotone reparameterization, which we term the gauge freedom of ordinal scoring. We propose the Gauge-Fixed Ordinal Network (GON), a temporal convolutional model trained with an anchor-and-variance objective that pins level-wise score means to shared target coordinates. GON operates on 2-jet features that expose local trajectory geometry, preserved by smooth flows and disrupted by stochastic surrogate procedures. On five held-out dynamical systems, initializing from a pretrained GON checkpoint consistently outperforms training from scratch across all window budgets, with adaptation depth reflecting geometric proximity to the training family. Zero-shot scores retain ordinal structure at the stochastic boundary, where surrogate procedures most strongly disrupt nonlinear geometry, and pretrained initialization consistently beats scratch across all window budgets. Pairwise discrimination and globally coherent ordinal scoring are distinct properties requiring a stable score coordinate for cross-system transfer, with direct implications for predictability assessment, model selection, and early-warning diagnostics across natural and engineered dynamical systems.
An observed time series can often be viewed as a projection of an underlying dynamical system, though the generative process may be deterministic, stochastic, or a mixture of both. The source of predictive structure shapes which modeling strategies are appropriate. Deterministic chaotic systems are governed by an underlying flow that supports geometry-aware representations, even when long-horizon forecasting is limited by sensitive dependence on initial conditions. Purely stochastic processes are better described through statistical regularities. Both regimes can share low-order moments such as autocorrelation and power spectrum, making the distinction difficult to draw from observations alone. A signal may reflect low-dimensional deterministic dynamics, a stochastic process, or an intermediate regime that retains some structure while losing others.
Classical tools approach this distinction within a single system. Surrogate hypothesis tests [1], [2] ask whether a signal departs from a chosen stochastic null. Lyapunov exponent estimators [3], [4] quantify instability under an assumed deterministic model. Complexity statistics such as permutation entropy [5] and sample entropy [6] compress a signal into a scalar intended to separate deterministic from stochastic behavior. Machine learning approaches have extended work on predictive structure in dynamical systems, including neural differential equation models [7], neural operators [8], [9], and pretrained models for chaotic forecasting [10]. These methods characterize structure within a system or family, and a new signal typically requires new calibration, new surrogates, or a fresh thresholding decision. A detailed comparison is provided in Appendix 7.
This paper addresses a specific problem: given a trajectory window, assign it a position on an ordered continuum of predictability regimes. We use predictability to mean the degree to which the underlying dynamics constrain future trajectory evolution, a property of the governing flow rather than of any particular model or observation process. We instantiate this continuum as a five-level predictability ladder spanning structured deterministic dynamics, weakly chaotic regimes, strongly chaotic regimes, structured stochastic surrogates, and unstructured noise. Standard deterministic-versus-stochastic discrimination corresponds to one boundary in this richer ordinal problem. The central question is whether a learned ordinal score transfers across systems, rather than merely separating classes within a single one.
The ordinal setting introduces a structural ambiguity. If a scalar score correctly orders the ladder levels, any strictly increasing reparameterization preserves the same ordinal decisions after a corresponding threshold adjustment. Correct ordering alone does not determine the numerical coordinate of the score, an identifiability property of ordered regression models [11], leaving scores across systems non-comparable even when both induce identical ordinal predictions. We refer to this as the gauge freedom of ordinal scoring and identify it as the structural barrier to cross-system transfer; post-hoc per-system recalibration cannot substitute, because the intended use case is zero-shot transfer where no target-system data is available to fit a calibration mapping.
To address this, we introduce the Gauge-Fixed Ordinal Network (GON), a temporal convolutional model trained with an anchor-based ordinal objective. The anchor fixes target locations for level-wise score means and penalizes excessive within-level spread, producing a scalar score with a stable coordinate convention. GON operates on 2-jet features (position, velocity, and acceleration), which expose local trajectory geometry [12] preserved by smooth flows and disrupted by surrogate procedures; this representation has been applied to learned physical dynamics in [13]. We train on twelve dissipative chaotic ODE systems, chosen for their bounded attractors and varied geometric structure, and evaluate zero-shot transfer and few-shot adaptation on five fully held-out systems.
The main contributions are as follows.
Problem formulation. We formalize ordinal predictability estimation and introduce the predictability ladder, identifying gauge freedom as the structural source of cross-system score ambiguity.
Method. We propose GON, which combines 2-jet trajectory features with an anchor-based objective to learn a shared score coordinate across systems.
Empirical results. We show that gauge fixing enables cross-system ordinal transfer: pretrained GON outperforms training from scratch across all held-out systems, zero-shot scores retain ordinal coherence at the surrogate boundary, and adaptation depth reflects geometric proximity to the training family.
We formalize ordinal predictability estimation: given a trajectory window, assign it a position on an ordered predictability spectrum that remains meaningful across systems, subsuming binary deterministic-versus-stochastic discrimination as a special case.
Let \(x = (x_t)_{t=1}^T \in \mathbb{R}^{d \times T}\) denote a multivariate time series generated by an unknown process. We define a five-level ordered taxonomy \[\mathcal{R} = \{L_0, L_1, L_2, L_3, L_4\}\] called the predictability ladder, arranged from more predictable to less predictable regimes.
Definition 1 (Predictability Ladder). The levels satisfy \[L_0 \prec L_1 \prec L_2 \prec L_3 \prec L_4,\] with the following interpretation:
Stable deterministic. Fixed points or limit cycles, where the present state strongly constrains the future, e.g.,deterministic flows with non-positive largest Lyapunov exponent \(\lambda_{\max} \le 0\).
Weakly chaotic.Deterministic dynamics with small positive instability and a longer but finite predictability horizon, \(0 < \lambda_{\max}\tau \ll 1\).
Strongly chaotic. Deterministic dynamics with stronger instability and rapid trajectory divergence, \(\lambda_{\max}\tau \gg 0\).
Structured stochastic. Processes that preserve selected linear statistics of a deterministic signal while disrupting nonlinear structure, e.g.,surrogate transformations \(x_t^{(3)} = \mathcal{S}(x_t^{(2)})\) that preserve the power spectrum but destroy deterministic phase-space geometry.
Unstructured stochastic.i.i.d.noise with no predictive structure beyond the marginal distribution, \(x_t \sim \mathcal{P}\) and \(I(x_t ; x_{t+\tau}) = 0\) for \(\tau>0\).
This ladder reflects two axes of limited predictability. Along \(L_0\)–\(L_2\), regimes remain deterministic, with predictability degrading as Lyapunov instability shortens the forecast horizon. The \(L_2\!\to\!L_3\) transition instead destroys deterministic geometry: surrogate procedures preserve linear statistics while removing nonlinear flow structure. Level \(L_4\) removes all temporal structure beyond the marginal distribution. Finer granularity can be introduced within adjacent levels by progressively increasing instability (\(L_0\)–\(L_2\)), disrupting deterministic geometry (\(L_2\)–\(L_3\)), or removing residual temporal dependence (\(L_3\)–\(L_4\)). This ordering parallels predictability-horizon arguments in nonlinear dynamical systems, where increasing trajectory divergence reduces forecastability.
Definition 2 (Ordinal Detectability). A scoring function \(E_\theta:\mathbb{R}^{d \times T}\rightarrow\mathbb{R}\) is ordinally detectable* under a distribution \(\mathcal{D}\) over \((x,y)\in\mathcal{X}\times\{0,1,2,3,4\}\) if \[\mathbb{E}[E_\theta(x)\mid y=i] < \mathbb{E}[E_\theta(x)\mid y=j] \qquad \forall\, i<j.\]*
Let \(\mu_k = \mathbb{E}[E_\theta(x)\mid y=k]\). To quantify global ordinal structure beyond pairwise separation, we use the monotonicity slope \[\beta = \frac{\sum_k (k-\bar{k})(\mu_k-\bar{\mu})}{\sum_k (k-\bar{k})^2},\] the least-squares slope of level means on ladder indices. A method may achieve a strong AUROC at a single boundary while producing \(\beta \approx 0\) if the full ladder is not coherently ordered. We report \(\beta\) both in-distribution and under transfer to systems withheld entirely from training.
A scalar ordinal score is useful across systems only if its numerical values retain a stable interpretation; ranking alone does not guarantee this. Given a score \(E:\mathcal{X}\to\mathbb{R}\) and thresholds \(\tau_0 < \tau_1 < \tau_2 < \tau_3\), the induced predictor is \(h_E(x) = \sum_{m=0}^{3}\mathbf{1}\{E(x)>\tau_m\}\), mapping \(x\) to one of five ordered levels.
Proposition 1 (Monotone reparameterization invariance). Let \(f:\mathbb{R}\to\mathbb{R}\) be strictly increasing. The predictor induced by \(E\) is unchanged under \(\tilde{E} = f \circ E\) when thresholds are transformed to \(\tilde{\tau}_m = f(\tau_m)\). (Proof in Appendix 8.)**
Ordinal supervision therefore determines score ordering but not its numerical coordinate: the score is identifiable only up to a monotone reparameterization. While this ambiguity is harmless within a single system, it becomes consequential under transfer, where identical score values may correspond to different predictability regimes. Cross-system comparison thus requires a shared coordinate convention, which the anchor objective provides by fixing level-wise score locations globally.
GON has three components: 2-jet preprocessing, a temporal convolutional encoder, and a gauge-fixed ordinal objective.
Raw state coordinates carry system-specific scaling, orientation, and units. Representing each window by its 2-jet exposes local trajectory geometry independent of these system-specific factors. For a smooth trajectory \(x(t)\), the 2-jet is the tuple \[(x(t),\dot{x}(t),\ddot{x}(t)),\] consisting of position, velocity, and acceleration. These quantities summarize local behavior up to second order and provide direct access to geometric features such as tangent direction and curvature. For trajectories generated by smooth ODEs, these components are linked by the governing flow. Surrogate procedures, by contrast, preserve selected linear statistics while disrupting the original nonlinear relation among position, velocity, and acceleration. This makes the 2-jet a natural representation for distinguishing deterministic trajectories from structured stochastic surrogates.
Given a window \(x\in\mathbb{R}^{d\times T}\), we first apply a Savitzky–Golay filter (window length \(7\), polynomial order \(3\); clamped to signal length and kept odd if shorter) to obtain a smoothed trajectory \(\hat{x}\). Central finite differences are then used to estimate derivatives, as they provide symmetric second-order accurate approximations while minimizing phase bias in the local trajectory geometry: \[\begin{align} \dot{x}_t &= \frac{\hat{x}_{t+1}-\hat{x}_{t-1}}{2\Delta t}, \\ \ddot{x}_t &= \frac{\hat{x}_{t+1}-2\hat{x}_t+\hat{x}_{t-1}}{\Delta t^2}. \end{align}\] One timestep is removed from each boundary to avoid edge effects. The resulting representation is \[J^2(x) = [\hat{x},\,\dot{x},\,\ddot{x}] \in \mathbb{R}^{3d \times (T-2)},\] which serves as the network input.
The encoder is a temporal convolutional network (TCN) of dilated residual blocks that maps the 2-jet input to a hidden sequence \(h \in \mathbb{R}^{H\times T'}\). After normalization, multi-scale average pooling aggregates \(h\) into a vector \(z \in \mathbb{R}^{4H}\), which a two-layer MLP maps to a scalar score: \[E_\theta(x) = \mathbf{W}_2\,\mathrm{GELU}\!\left(\mathrm{LN}(\mathbf{W}_1 z + b_1)\right) + b_2.\] Outputs are smoothly clipped via \(E_\theta \leftarrow c\,\tanh(E_\theta/c)\), \(c=5\), to keep scores aligned with the anchor range. Full details appear in Appendix 9.
Ranking-based supervision does not pin a unique score coordinate (Section 2.3). We fix one by assigning target locations \[\{t_k\}_{k=0}^4 = \{-4,-2,0,+2,+4\}\] and training the model so that each level concentrates around its anchor. These numeric targets are arbitrarily ordered anchors chosen to fix a shared coordinate; any monotonic rescaling yields an equivalent ordinal representation. Equal spacing is chosen for simplicity; the ordinal structure is invariant to any monotone rescaling of the targets, so the specific values affect only the absolute score scale and not relative ordering. The objective is \[\mathcal{L}(\theta) = \lambda_a \sum_{k=0}^{4} (\mu_k - t_k)^2 + \lambda_v \sum_{k=0}^{4} [\sigma_k - \sigma^*]_+^2, \label{eq:loss}\tag{1}\] where \(\mu_k\) and \(\sigma_k\) are the mini-batch mean and standard deviation of scores for level \(k\), \([\cdot]_+=\max(0,\cdot)\), and \(\sigma^*=0.3\). The anchor term fixes level-wise score locations, selecting a canonical representative from the class of monotone-equivalent solutions (Appendix 11); the variance term limits within-level dispersion, together defining a shared affine coordinate across systems.
A natural alternative is margin loss \(\mathcal{L}_{\text{margin}} = \sum_{k=0}^{3}\max(0,\,\gamma-(\mu_{k+1}-\mu_k))\), which enforces adjacent-level separation but leaves the absolute score axis unfixed. Scores therefore remain non-comparable across systems even when levels are well-separated; Section 4.4 tests this directly.
We set \(\lambda_a=1.0\) with linear warmup over the first five epochs and \(\lambda_v=0.01\). Optimization uses AdamW with learning rate \(2\times10^{-3}\), weight decay \(10^{-4}\), cosine decay over \(30\) epochs, gradient clipping at \(5.0\), and EMA decay \(0.995\). Training uses a batch size of \(64\). Weighted sampling balances ladder levels. The full experimental suite required approximately 40–50 GPU-hours on a single A100.
Augmentation is applied to raw trajectory windows \(x \in \mathbb{R}^{d \times T}\) before 2-jet preprocessing and includes additive Gaussian noise (\(\sigma_{\text{rel}}=0.05\)), multiplicative scaling (\(\times[0.9,1.1]\)), temporal masking (\(5\%\)), linear drift (\(\pm 0.05\)), and causal time shifts (\(\leq 5\) steps). Each operation is applied independently with probability \(0.5\); the full pipeline activates with probability \(0.7\). Augmentation precedes smoothing, so the filter attenuates discontinuities introduced by masking or time shifts before finite differences are computed.
The experiments address four questions:
Does the gauge-fixed coordinate remain calibrated on the source distribution, and does it transfer in zero-shot to held-out systems?
Does pretraining provide a better initialization than scratch for adapting to unseen systems under limited labeled windows?
Which objective component is responsible for cross-system coordinate stability?
Does coherent 2-jet structure contribute beyond dimensionality-matched controls?
Training uses 12 three-dimensional dissipative chaotic ODE systems spanning a range of attractor geometries: Chen, Chua, Duffing, Finance, Genesio–Tesi, Hastings–Powell, Lorenz-63, Lorenz-84, Rössler, Rucklidge, Shimizu–Morioka, and Halvorsen. Each system is simulated to produce trajectories of length \(N=4096\), with system-specific timesteps chosen to resolve dominant dynamical timescales; \(\Delta t\) is determined from the observed sampling rate and requires no system-level knowledge.
Ladder levels follow Definition 1: deterministic regimes (\(L_0\)–\(L_2\)) are assigned using \(\lambda_1\), \(L_3\) is generated by applying surrogate procedures (AAFT, IAAFT, WLS, phase-shuffle) to \(L_2\) trajectories, and \(L_4\) is i.i.d.noise. Full simulation and labeling details appear in Appendix 10.
Cross-system transfer is evaluated on five held-out systems: Thomas, Tigan, Newton–Leipnik, Rabinovich–Fabrikant, and the forced pendulum.
Trajectories are segmented into windows of length \(T=256\) with stride \(128\). Inputs are normalized via per-trajectory, per-channel z-score standardization, requiring no source-distribution statistics and applying equally to any held-out system. After 2-jet preprocessing, each example is a \(9\times 254\) tensor. Windows are the sole input unit; system-level information is unknown throughout. The adaptation experiment measures how many windows suffice to recalibrate a pretrained model to a new system under limited data.
The monotonicity slope \(\beta\) (Definition 2) measures global ordinal coherence. Cross-method zero-shot comparisons use the scale-normalized \(\beta_{\text{norm}}\); raw \(\beta\) is used within the GON family, where the score scale is fixed by construction. Adjacent-pair AUROC at each boundary \((L_0,L_1)\) through \((L_3,L_4)\) serves as a diagnostic, since a method can separate one transition well while failing global coherence. For in-distribution calibration, we also report anchor drift \[\delta_{\text{anchor}}=\frac{1}{5}\sum_{k=0}^4 |\mu_k-t_k|,\] measuring how closely the learned level means match the anchor targets.
Models are trained on the source systems and applied directly to the held-out systems.
Before assessing transfer, it is useful to verify that the gauge-fixed score remains well aligned with its intended coordinate system on held-out source-domain data. Table 1 reports level-wise score statistics on the test split of the training systems. GON closely matches the anchor targets, with mean anchor drift \(\delta_{\text{anchor}}=0.0097\). The within-level standard deviations (\(0.08\)–\(0.32\)) remain well below the 2-unit spacing between adjacent anchors.
| \(L_0\) | \(L_1\) | \(L_2\) | \(L_3\) | \(L_4\) | |
|---|---|---|---|---|---|
| Target | \(-4.0\) | \(-2.0\) | \(0.0\) | \(+2.0\) | \(+4.0\) |
| GON | \(-3.99{\pm}0.32\) | \(-1.99{\pm}0.21\) | \(-0.02{\pm}0.11\) | \(+2.00{\pm}0.23\) | \(+4.01{\pm}0.08\) |
Figure 2 compares GON with neural baselines on the held-out systems. All supervised baselines cluster tightly at \(\beta_{\text{norm}} \approx 0.46\)–\(0.48\) despite using different objectives. GON achieves the highest value, \(\beta_{\text{norm}}=0.559\). Full per-boundary AUROC and per-system neural baseline results are reported in Appendix 13. The transfer pattern is selective rather than uniform. The strongest zero-shot separation occurs at the stochastic boundary (\(L_2 \to L_3\) and \(L_3 \to L_4\)), while deterministic-side boundaries remain weaker: \(\text{AUROC}_{01}\) is near chance and \(\text{AUROC}_{12}\) is only moderate (Appendix 13). The gauge-fixed score thus retains meaningful ordinal structure under distribution shift, with transfer concentrated where surrogate procedures most strongly disrupt nonlinear geometry.
We evaluate whether GON pretraining provides a better initialization than training from scratch across \(k \in \{5,10,\ldots,100\}\) labeled windows per held-out system, averaged over five random seeds. Three strategies are compared: scratch (random initialization), pre_head (frozen encoder, readout fine-tuned), and pre_all (full model fine-tuned from the pretrained checkpoint). Table 2 reports averages over all five systems; Figure 3 and Appendix 14 show per-system results and fine-tuning details.
At \(k=5\), pre_all leads scratch by \(+0.776\) in \(\beta\) (\(1.319\) vs.\(0.543\)) and \(+0.191\) in \(\text{AUROC}_{23}\) (\(0.758\) vs.\(0.567\)), and this gap does not close at \(k=100\) (Table 2), indicating that the pretrained encoder encodes geometric structure that random initialization cannot recover within this window budget. pre_head reliably exceeds scratch on \(\text{AUROC}_{23}\) but captures little of pre_all’s gain in \(\beta\), confirming that effective adaptation requires updating the encoder.
| \(\text{AUROC}_{23}\) | \(\beta\) | |||||
|---|---|---|---|---|---|---|
| 2-4(lr)5-7 \(k\) | scratch | pre_head | pre_all | scratch | pre_head | pre_all |
| \(5\) | \(0.567\) \({\scriptstyle\,\pm\,0.14}\) | \(0.717\) \({\scriptstyle\,\pm\,0.25}\) | \(\mathbf{0.758}\) \({\scriptstyle\,\pm\,0.21}\) | \(0.543\) \({\scriptstyle\,\pm\,0.48}\) | \(1.122\) \({\scriptstyle\,\pm\,0.56}\) | \(\mathbf{1.319}\) \({\scriptstyle\,\pm\,0.46}\) |
| \(10\) | \(0.566\) \({\scriptstyle\,\pm\,0.13}\) | \(0.714\) \({\scriptstyle\,\pm\,0.25}\) | \(\mathbf{0.762}\) \({\scriptstyle\,\pm\,0.21}\) | \(0.664\) \({\scriptstyle\,\pm\,0.51}\) | \(1.135\) \({\scriptstyle\,\pm\,0.55}\) | \(\mathbf{1.341}\) \({\scriptstyle\,\pm\,0.48}\) |
| \(25\) | \(0.567\) \({\scriptstyle\,\pm\,0.13}\) | \(0.718\) \({\scriptstyle\,\pm\,0.25}\) | \(\mathbf{0.765}\) \({\scriptstyle\,\pm\,0.19}\) | \(0.958\) \({\scriptstyle\,\pm\,0.62}\) | \(1.149\) \({\scriptstyle\,\pm\,0.52}\) | \(\mathbf{1.385}\) \({\scriptstyle\,\pm\,0.40}\) |
| \(50\) | \(0.580\) \({\scriptstyle\,\pm\,0.14}\) | \(0.723\) \({\scriptstyle\,\pm\,0.26}\) | \(\mathbf{0.795}\) \({\scriptstyle\,\pm\,0.17}\) | \(1.055\) \({\scriptstyle\,\pm\,0.62}\) | \(1.177\) \({\scriptstyle\,\pm\,0.52}\) | \(\mathbf{1.551}\) \({\scriptstyle\,\pm\,0.34}\) |
| \(100\) | \(0.562\) \({\scriptstyle\,\pm\,0.15}\) | \(0.723\) \({\scriptstyle\,\pm\,0.25}\) | \(\mathbf{0.782}\) \({\scriptstyle\,\pm\,0.18}\) | \(1.062\) \({\scriptstyle\,\pm\,0.72}\) | \(1.177\) \({\scriptstyle\,\pm\,0.54}\) | \(\mathbf{1.489}\) \({\scriptstyle\,\pm\,0.44}\) |
Systems with polynomial or exponential vector fields represented in the source distribution (Rabinovich–Fabrikant, Newton–Leipnik, Tigan) adapt immediately or are solved at zero-shot. Rabinovich–Fabrikant is the clearest case: scratch yields \(\beta\approx0.01\)–\(0.54\) across all \(k\), never recovering ordinal structure, whereas pre_all exceeds zero-shot from \(k=10\). Thomas and the forced pendulum are harder: both contain state-dependent sinusoidal vector-field terms absent from training (e.g.,\(\sin\theta\), \(\sin y\)), producing unseen oscillatory curvature in the 2-jet. In both cases, pre_all recovers global ordinal structure but fails to resolve the \(L_2\to L_3\) boundary, identifying sinusoidal vector-field geometry as the primary gap in source-distribution coverage.
The ablations ask which parts of the method matter for transfer.
Table 3 shows that coordinate stability, not ordinal coherence alone, is critical for transfer. Replacing the anchor with a margin objective leaves in-distribution AUROC unchanged by at most \(0.010\) across adjacent pairs and produces comparable scale-normalized transfer (\(\beta_\text{norm}=0.554\) vs.\(0.559\)), yet anchor drift rises from \(0.005\) to \(2.556\). A model can therefore retain ordinal coherence while losing coordinate stability, confirming that relative level separation is insufficient for cross-system comparability: without a fixed coordinate, scores on unseen systems carry no consistent numerical meaning. The variance term plays a secondary but consistent role: removing it increases within-level spread from \(0.195\) to \(0.230\) and lowers transfer \(\beta\) from \(1.118\) to \(1.055\).
| Objective | \(\delta_\text{anchor}\) (indist) | \(\beta\) (indist) | \(\beta\) (Transfer) | \(\bar{\sigma}\) (indist) |
|---|---|---|---|---|
| Anchor \(+\) Variance (ours) | 0.005 | 2.002 | 1.118 | 0.195 |
| Margin \(+\) Variance | 2.556 | 1.075 | 0.596 | 0.244 |
| Anchor only | 0.014 | 2.006 | 1.055 | 0.230 |
Table 4 shows that coherent derivative structure improves transfer beyond raw states. Adding first derivatives improves \(\beta\) from \(1.132\) to \(1.336\); the full 2-jet improves further to \(\beta=1.504\) and \(\text{AUROC}_{23}=0.846\). Shuffling derivatives across time drops \(\beta\) to \(0.983\), confirming that temporal coherence among position, velocity, and acceleration drives the gain. A control replacing derivatives with per-pass Gaussian noise remains near the raw-state baseline, ruling out channel count.
| Input representation | Transfer \(\beta\) | Transfer \(\text{AUROC}_{23}\) |
|---|---|---|
| Raw state (\(x\) only) | 1.132 | 0.791 |
| 1-jet (\(x\), \(\dot{x}\)) | 1.336 | 0.789 |
| 2-jet (\(x\), \(\dot{x}\), \(\ddot{x}\)) | 1.504 | 0.846 |
| Shuffled 2-jet\(^\dagger\) (temporal control) | 0.983 | 0.737 |
| Random channels\(^\dagger\) (dim.control) | 1.174 | 0.734 |
GON remains robust under moderate corruption, with \(\text{AUROC}_{23}\) staying above \(0.89\) at \({\geq}\,20\) dB SNR and \(30\%\) temporal dropout. Performance degrades gracefully in \(\beta\), indicating reduced transfer strength before loss of discriminability. Full results appear in Appendix 15.
Replacing the anchor term with a margin loss leaves in-distribution discrimination nearly unchanged but increases anchor drift and reduces transfer \(\beta\) by \(47\%\) (Table 3). Source-domain separability alone, therefore, does not ensure that a scalar score retains a stable meaning on unseen systems. Gauge fixing instead learns a shared ordinal coordinate, enabling predictability comparisons that are transferred across heterogeneous dynamics.
Zero-shot transfer is strongest at the stochastic boundary (\(L_2 \to L_3\) and \(L_3 \to L_4\)); deterministic-side boundaries remain weaker. The same asymmetry holds under perturbation: \(\text{AUROC}_{23}\) is more robust than \(\text{AUROC}_{01}\) to noise and dropout. GON captures geometric structure disrupted by surrogate generation, whereas finer distinctions within deterministic chaos are more system-specific.
The 2-jet representation improves transfer by exposing local trajectory geometry. The benefit depends on coherent relationships among position, velocity, and acceleration; disrupting this alignment reduces the transfer even at a matched dimensionality. This supports the view that transferable predictability cues arise from shared geometric structure rather than system-specific statistics.
Systems whose vector fields resemble those seen in training adapt immediately or are solved at zero-shot. In contrast, systems with unseen structure (e.g., sinusoidal forcing in Thomas and the forced pendulum) require more data and may fail to resolve finer boundaries. Performance degrades smoothly with geometric distance from the source distribution.
The benchmark covers 12 source and 5 target systems, all low-dimensional and dissipative; extensions to higher-dimensional systems, conservative dynamics, or real-world signals, remain unestablished. This work takes one of the earliest steps toward a general theory of transferable predictability estimation, establishing gauge-fixing as a necessary condition for cross-system ordinal comparability. Predictive structure in real systems is often limited by chaos, stochastic forcing, or their interaction, arising across weather and climate dynamics, turbulent flows, biological oscillations, financial processes, and engineered systems. By learning a transferable ordinal predictability coordinate, the framework enables cross-system comparison without system-specific calibration, informing model selection, early-warning diagnostics, and the design of resilient coupled human–natural systems. The method operates exclusively on synthetic ODE signals, involves no personal data, and has no foreseeable negative societal impacts.
This paper identifies gauge-fixing as the missing ingredient for transferable ordinal predictability scores. The central challenge is not ranking regimes within a single system, but assigning scores whose numerical values retain a consistent interpretation across systems. The loss ablation makes the point cleanly: replacing the anchor with a margin objective leaves the source-domain discrimination nearly unchanged while reducing transfer \(\beta\) by \(47\%\). Ordinal coherence is robust to 4-bit quantization and Gaussian noise down to 20 dB SNR, confirming the learned coordinate survives common observational degradations. Pairwise discrimination and globally coherent ordinal scoring are distinct properties that can dissociate, and a stable coordinate is necessary to bridge them under transfer. The benchmark is limited to low-dimensional dissipative systems; this bound does not invalidate the core finding, since the ablation isolates coordinate convention as the operative variable within a fixed system class. Extension to higher-dimensional attractors, partial observability, and real-world signals in weather, climate, and physiology remain the most consequential direction for future work.
Geometric diagnostics [14], [15] and prediction-based tests study rejection of the stochastic null hypothesis within a single system. Permutation entropy [5] and sample entropy [6] provide scalar complexity measures widely used as binary indicators of dynamical complexity. The present work places this binary separation within a broader ordinal hierarchy, where pairwise discriminability at a single boundary and global ordinal coherence across levels need not coincide.
The classical toolbox for the binary problem is surrogate hypothesis testing [1], [2], which constructs a null distribution from the signal and tests whether its nonlinear statistics exceed those of matched stochastic alternatives. Complementary approaches estimate the maximal Lyapunov exponent directly [3], [4]. Both are designed for single-system analysis: each new signal requires new surrogates or freshly tuned estimators, and both address a binary question. The present work considers an ordinal formulation in which predictability is represented on a continuum, with scores intended to remain comparable across systems without per-signal calibration.
Critical slowing down theory [16] and spectral precursor methods [17] study statistical signatures of approaching bifurcations. The predictability ladder subsumes early-warning detection as a special case: \(L_1\) is precisely the transition regime where such signatures emerge. The present benchmark does not directly evaluate early-warning performance. Transfer is strongest at the stochastic boundary rather than the deterministic-side transition, and calibrated cross-system early-warning detection remains a natural extension.
Cumulative link models [18] and neural extensions such as CORAL [19] enforce a consistent rank structure across ordered labels. Ordinal contrastive methods [20] encode ordering in representation space. Proxy-based metric learning [21] uses class-level anchor prototypes to stabilize embedding geometry. These formulations study ordering within a single system or dataset, while the setting considered here involves scores that are intended to remain comparable across heterogeneous systems.
Neural differential equation models [7] and Lagrangian networks [13] encode geometric and physics-aware inductive biases in learned representations. Reservoir computing has been applied to chaotic forecasting [22]. The 2-jet representation used here is grounded in differential geometry: position, velocity, and acceleration jointly encode local trajectory curvature in a way that is invariant to system-specific scaling and is disrupted by surrogate generation procedures.
Gradient-based meta-learning [23] has been applied to dynamical systems forecasting [24]. Neural operators [8], [9] learn solution operators across parameter families. [10] trained a transformer on a large synthetic corpus of chaotic systems and reported zero-shot forecasting on unseen ODEs and PDEs. These approaches study predictive generalization across systems under supervised forecasting objectives. The setting considered here focuses on learning representations of predictability regimes, where the goal is to assign comparable ordinal scores rather than to predict future states.
Let \(f:\mathbb{R}\to\mathbb{R}\) be strictly increasing. For any threshold \(\tau_m\) and input \(x\), \[E(x) > \tau_m \iff f(E(x)) > f(\tau_m),\] since strictly increasing functions preserve the direction of all inequalities. The induced ordinal predictor \(h_E(x) = \sum_{m=0}^{3}\mathbf{1}\{E(x) > \tau_m\}\) depends only on these threshold comparisons. Under the transformed score \(\tilde{E} = f \circ E\) with transformed thresholds \(\tilde{\tau}_m = f(\tau_m)\), each comparison \(\tilde{E}(x) > \tilde{\tau}_m\) is equivalent to \(E(x) > \tau_m\). Therefore, \(h_{\tilde{E}}(x) = h_E(x)\) for all \(x\), and the induced ordinal prediction is unchanged. 0◻
The encoder is a temporal convolutional network (TCN) composed of \(B = 6\) residual blocks with exponentially increasing dilation factors \([1, 2, 4, 8, 16, 32]\). Each block consists of two 1D dilated convolutions with kernel size \(k = 3\), followed by Group Normalization and GELU activation. Residual connections are applied at the block level.
The hidden dimension is fixed at \(H = 128\) across all layers.
The effective receptive field of the network is \[\mathcal{R} = 1 + 2(k - 1)\sum_{i=0}^{B-1} 2^i = 253,\] which covers nearly the full temporal extent of the input after 2-jet boundary trimming (\(T' = T - 2\)). This allows the model to capture long-range temporal dependencies relevant for distinguishing predictability regimes.
The encoder produces a feature map \(h \in \mathbb{R}^{H \times T'}\). RMS normalization is applied before aggregation. Average pooling is performed at scales \(s \in \{4, 16, 64,
256\}\) to capture structure at multiple temporal resolutions; for \(s > T'\) the kernel is clamped to \(k = \min(s, T')\), so at scale \(256\) this reduces to a global average over all \(254\) timesteps. Pooling uses no padding; non-integer-multiple tails are truncated by F.avg_pool1d with default
padding=0. The pooled representations are concatenated to form a feature vector \(z \in \mathbb{R}^{4H}\).
The pooled feature vector \(z \in \mathbb{R}^{4H}\) (dimension \(512\)) is mapped to a scalar score via a two-layer MLP with hidden dimension \(H=128\): \[E_\theta(x) = \mathbf{W}_2\,\mathrm{GELU}\!\left(\mathrm{LN}(\mathbf{W}_1 z + b_1)\right) + b_2,\] where \(\mathbf{W}_1 \in \mathbb{R}^{128 \times 512}\) and \(\mathbf{W}_2 \in \mathbb{R}^{1 \times 128}\). To prevent unbounded growth and maintain consistency with the fixed anchor targets, the output is smoothly clipped: \[E_\theta \leftarrow c \tanh(E_\theta / c), \quad c = 5.\] The full model contains \(616{,}233\) trainable parameters.
Each system is simulated to produce \(N = 4096\) uniformly sampled time steps. The timestep \(\Delta t\) is selected to resolve the dominant dynamical timescale of each system. We estimate a characteristic timescale \(\tau\) from the mean inter-peak interval of the trajectory, and set \(\Delta t\) such that approximately 64 dominant cycles are covered, with \(\Delta t\) clamped to \([10^{-4}, 10^{-1}]\).
Integration is performed using either a fixed-step fourth-order Runge–Kutta method or the adaptive Dormand–Prince (DOPRI5) solver, depending on system stiffness.
Deterministic regimes follow Definition 1, with \(\lambda_1\) estimated via [4]. \(L_0\) corresponds to \(\lambda_1 < 0\); \(L_1\) and \(L_2\) are distinguished by the relative magnitude of \(\lambda_1\), cross-checked by visual inspection of attractor phase portraits. Because attractor geometries differ substantially across systems, no single numerical threshold is applied uniformly; regime boundaries are confirmed by qualitative separation in both the exponent estimate and the phase portrait. These boundaries are fixed at data generation time and held constant throughout all experiments; the qualitative inspection step affects only this initial label assignment and does not enter any learned component of the method. \(L_3\) is generated from \(L_2\) trajectories using AAFT, IAAFT, WLS, and phase randomization [1], [2], [25], [26]. \(L_4\) is i.i.d.Gaussian noise matched in length and dimensionality.
Train/validation/test splits are performed at the trajectory level using an 80/10/10 partition. All normalization statistics are computed on the training split and held fixed for validation and testing.
Ordinal regression determines only the ordering of latent scores. For any strictly increasing \(f\), the transformed score \(\tilde{E}(x)=f(E(x))\) induces identical ordinal decisions under the corresponding threshold transformation, so \(E\) is identifiable only up to a monotone reparameterization [11]. Within a single system, this ambiguity is harmless: thresholds are learned jointly with the score, and ranking suffices. Under cross-system transfer, it becomes consequential because identical ordinal predictions on two systems need not assign the same numeric value to the same predictability regime.
The anchor objective removes this degree of freedom by fixing the level-wise score means, \[\mathbb{E}[E(x)\mid y=k] = t_k,\] standardizing the macroscopic location and scale of the score distribution, breaking the unbounded gauge freedom of ordinal regression. The variance term further restricts within-level dispersion. Together, they establish a shared coordinate convention across systems without claiming strict uniqueness: infinitely many nonlinear strictly increasing functions could, in principle, satisfy these moment constraints, but the anchor fixes the operationally relevant degrees of freedom for cross-system comparison.
This differs from post-hoc calibration, which rescales scores after training on a per-system basis and, therefore, cannot produce a universal numeric coordinate. Gauge fixing constrains the representation during learning, so the coordinate generalizes by construction.
For each system, we report the governing equations, fixed parameter values, the single parameter varied to traverse the predictability ladder, and the parameter range assigned to each regime. Regime boundaries were determined by the protocol described in Appendix 10: the sign and relative magnitude of the leading Lyapunov exponent \(\lambda_1\) cross-checked by visual inspection of attractor phase portraits. Initial conditions were sampled from multiple distributions per system (Gaussian perturbations about a base point, uniform on a sphere, and system-specific defaults); all distributions produced qualitatively consistent regime behavior.
The stable range corresponds to L0, the transition range to L1 (weakly chaotic), and the chaotic range to L2 (strongly chaotic).
Unless noted otherwise, initial conditions were sampled from three distributions: Gaussian perturbations about a base point, uniform on a sphere, and system-specific defaults. All distributions produced qualitatively consistent regime behavior. Two exceptions are noted explicitly: Hastings–Powell, which requires ecologically structured non-negative initial conditions, and Chua, which additionally uses double-scroll sampling centered on the two attractor scrolls.
\[\begin{align} \dot{x} &= a(y - x) \\ \dot{y} &= (c - a)x - xz + cy \\ \dot{z} &= xy - bz \end{align}\]
Fixed parameters: \(a = 35\), \(b = 3\). Varied parameter: \(c\).
| Regime | \(c\) range |
|---|---|
| L0 (stable) | \([15.0,\;23.0)\) |
| L1 (transition) | \([23.0,\;28.0)\) |
| L2 (chaotic) | \([28.0,\;33.0)\) |
\[\begin{align} \dot{x} &= \alpha\bigl(y - x - f(x)\bigr) \\ \dot{y} &= x - y + z \\ \dot{z} &= -\beta y \end{align}\]
where the piecewise-linear Chua diode is \[f(x) = m_1 x + \tfrac{1}{2}(m_0 - m_1)\bigl(|x+1| - |x-1|\bigr).\]
Fixed parameters: \(\beta = 28.0\), \(m_0 = -1.143\), \(m_1 = -0.714\). Varied parameter: \(\alpha\).
| Regime | \(\alpha\) range |
|---|---|
| L0 (stable) | \([6.0,\;10.0)\) |
| L1 (transition) | \([10.0,\;13.0)\) |
| L2 (chaotic) | \([13.0,\;16.0)\) |
Chua additionally uses double-scroll sampling (\(\sigma = 0.3\) about \((\pm 1.5,\;0,\;0)\)) and Gaussian perturbations (\(\sigma = 0.01\)) about the base point \((0.7,\;0,\;0)\).
The Duffing oscillator is non-autonomous; the cosine forcing term renders the effective phase space three-dimensional.
\[\begin{align} \dot{x} &= y \\ \dot{y} &= -\delta y - \alpha x - \beta x^3 + \gamma\cos(\omega t) \end{align}\]
Fixed parameters: \(\delta = 0.1\), \(\alpha = -1.0\), \(\beta = 1.0\), \(\omega = 1.0\). Varied parameter: \(\gamma\) (forcing amplitude).
| Regime | \(\gamma\) range |
|---|---|
| L0 (stable) | \([0.05,\;0.15)\) |
| L1 (transition) | \([0.25,\;0.35)\) |
| L2 (chaotic) | \([0.35,\;0.60)\) |
\[\begin{align} \dot{x} &= \left(\tfrac{1}{b} - a\right)x + z + xy \\ \dot{y} &= -by - x^2 \\ \dot{z} &= -x - cz \end{align}\]
Fixed parameters: \(b = 0.2\), \(c = 1.1\). Varied parameter: \(a\).
| Regime | \(a\) range |
|---|---|
| L0 (stable) | \((4.0,\;8.0]\) |
| L1 (transition) | \([3.0,\;4.0)\) |
| L2 (chaotic) | \([0.001,\;3.0)\) |
\[\begin{align} \dot{x} &= y \\ \dot{y} &= z \\ \dot{z} &= -cx - by - az + x^2 \end{align}\]
Fixed parameters: \(b = 1.1\), \(c = 1.0\). Varied parameter: \(a\).
| Regime | \(a\) range |
|---|---|
| L0 (stable) | \([0.65,\;1.20)\) |
| L1 (transition) | \([0.50,\;0.65)\) |
| L2 (chaotic) | \([0.20,\;0.50)\) |
\[\begin{align} \dot{x} &= -ax - by - bz - y^2 \\ \dot{y} &= -ay - bz - bx - z^2 \\ \dot{z} &= -az - bx - by - x^2 \end{align}\]
Fixed parameter: \(b = 4.0\). Varied parameter: \(a\).
| Regime | \(a\) range |
|---|---|
| L0 (stable) | \([1.86,\;4.0)\) |
| L1 (transition) | \([1.80,\;1.86)\) |
| L2 (chaotic) | \([1.20,\;1.80)\) |
\[\begin{align} \dot{x} &= x(1-x) - \frac{a_1 xy}{1 + b_1 x} \\ \dot{y} &= \frac{a_1 xy}{1 + b_1 x} - \frac{a_2 yz}{1 + b_2 y} - d_1 y \\ \dot{z} &= \frac{a_2 yz}{1 + b_2 y} - d_2 z \end{align}\]
Fixed parameters: \(a_2 = 0.1\), \(b_1 = 3.0\), \(b_2 = 2.0\), \(d_1 = 0.4\), \(d_2 = 0.01\). Varied parameter: \(a_1\) (prey–predator interaction rate).
| Regime | \(a_1\) range |
|---|---|
| L0 (stable) | \([2.5,\;3.5)\) |
| L1 (transition) | \([3.5,\;4.8)\) |
| L2 (chaotic) | \([4.8,\;5.2)\) |
Initial conditions were absolute-valued to ensure non-negative populations, with ecological structured sampling (prey \(\in [0.5,1.0]\), predator \(\in [0.1,0.4]\), superpredator \(\in [0.2,0.8]\)) used in place of sphere sampling. Integration used an adaptive timescale with an \(8\times\) ecological multiplier applied to the estimated dominant period \(\tau\), reflecting the longer characteristic timescales of trophic oscillations.
\[\begin{align} \dot{x} &= \sigma(y - x) \\ \dot{y} &= x(\rho - z) - y \\ \dot{z} &= xy - \beta z \end{align}\]
Fixed parameters: \(\sigma = 10.0\), \(\beta = 8/3\). Varied parameter: \(\rho\).
| Regime | \(\rho\) range |
|---|---|
| L0 (stable) | \([10.0,\;20.0)\) |
| L1 (transition) | \([20.0,\;28.0)\) |
| L2 (chaotic) | \([28.0,\;40.0)\) |
\[\begin{align} \dot{x} &= -ax - y^2 - z^2 + aF \\ \dot{y} &= -y + xy - bxz + G \\ \dot{z} &= -z + bxy + xz \end{align}\]
Fixed parameters: \(a = 0.25\), \(b = 4.0\), \(G = 1.0\). Varied parameter: \(F\) (forcing).
| Regime | \(F\) range |
|---|---|
| L0 (stable) | \([1.5,\;4.0)\) |
| L1 (transition) | \([4.0,\;6.0)\) |
| L2 (chaotic) | \([6.0,\;10.0)\) |
\[\begin{align} \dot{x} &= -y - z \\ \dot{y} &= x + ay \\ \dot{z} &= b + z(x - c) \end{align}\]
Fixed parameters: \(a = 0.2\), \(b = 0.2\). Varied parameter: \(c\).
| Regime | \(c\) range |
|---|---|
| L0 (stable) | \([2.0,\;2.83)\) |
| L1 (transition) | \([2.83,\;4.2)\) |
| L2 (chaotic) | \([4.2,\;6.0)\) |
\[\begin{align} \dot{x} &= -ax + by - yz \\ \dot{y} &= x \\ \dot{z} &= -z + y^2 \end{align}\]
Fixed parameter: \(a = 2.0\). Varied parameter: \(b\).
| Regime | \(b\) range |
|---|---|
| L0 (stable) | \([1.0,\;3.0)\) |
| L1 (transition) | \([3.0,\;5.5)\) |
| L2 (chaotic) | \([5.5,\;8.0)\) |
\[\begin{align} \dot{x} &= y \\ \dot{y} &= x - ay - xz \\ \dot{z} &= -bz + x^2 \end{align}\]
Fixed parameter: \(b = 0.5\). Varied parameter: \(a\).
| Regime | \(a\) range |
|---|---|
| L0 (stable) | \([1.30,\;1.60)\) |
| L1 (transition) | \([1.20,\;1.30)\) |
| L2 (chaotic) | \([0.10,\;1.20)\) |
Non-autonomous; cosine forcing renders the effective phase space three-dimensional.
\[\begin{align} \dot{\theta} &= v \\ \dot{v} &= -\gamma v - \sin\theta + A\cos(\omega t) \end{align}\]
Fixed parameters: \(\gamma = 0.2\), \(\omega = 2/3\). Varied parameter: \(A\) (forcing amplitude).
| Regime | \(A\) range |
|---|---|
| L0 (stable) | \([0.30,\;0.60)\) |
| L1 (transition) | \([0.60,\;0.85)\) |
| L2 (chaotic) | \([0.90,\;1.40)\) |
\[\begin{align} \dot{x} &= -ax + y + 10yz \\ \dot{y} &= -x - 0.4y + 5xz \\ \dot{z} &= bz - 5xy \end{align}\]
Fixed parameter: \(a = 0.4\). Varied parameter: \(b\) (z-axis damping).
| Regime | \(b\) range |
|---|---|
| L0 (stable) | \([0.030,\;0.100)\) |
| L1 (transition) | \([0.100,\;0.150)\) |
| L2 (chaotic) | \([0.150,\;0.200)\) |
\[\begin{align} \dot{x} &= y(z - 1 + x^2) + \gamma x \\ \dot{y} &= x(3z + 1 - x^2) + \gamma y \\ \dot{z} &= -2z(\alpha + xy) \end{align}\]
Fixed parameter: \(\alpha = 0.1\). Varied parameter: \(\gamma\).
| Regime | \(\gamma\) range |
|---|---|
| L0 (stable) | \([0.05,\;0.12)\) |
| L1 (transition) | \([0.12,\;0.26)\) |
| L2 (chaotic) | \([0.26,\;0.29)\) |
Trajectories exceeding \(\|s\|_\infty = 25\) were truncated to the bounded portion before divergence; trajectories with fewer than 300 retained points were discarded.
\[\begin{align} \dot{x} &= \sin y - bx \\ \dot{y} &= \sin z - by \\ \dot{z} &= \sin x - bz \end{align}\]
No fixed parameters. Varied parameter: \(b\) (dissipation).
| Regime | \(b\) range |
|---|---|
| L0 (stable) | \([0.25,\;0.35)\) |
| L1 (transition) | \([0.15,\;0.25)\) |
| L2 (chaotic) | \([0.05,\;0.15)\) |
\[\begin{align} \dot{x} &= a(y - x) \\ \dot{y} &= (c - a)x - axz \\ \dot{z} &= -bz + xy \end{align}\]
Fixed parameters: \(a = 2.1\), \(b = 0.6\). Varied parameter: \(c\).
| Regime | \(c\) range |
|---|---|
| L0 (stable) | \([5.0,\;12.0)\) |
| L1 (transition) | \([12.0,\;22.0)\) |
| L2 (chaotic) | \([22.0,\;32.0)\) |
We compare GON with three neural baselines trained on the same 12 source systems using identical data, preprocessing, encoders, and optimization protocols. The only change is the training objective.
Regression minimizes mean squared error to integer labels \(\{0,1,2,3,4\}\). Classification minimizes cross-entropy over five classes. CORAL [19] uses cumulative link constraints for ordinal supervision.
Each method is converted to a scalar score for evaluation: the raw output for regression, the expected label \(\sum_k k\,p_k\) for classification, and the cumulative-logit score for CORAL.
Table 5 reports test performance on the 12 source systems. Regression, classification, CORAL, and GON all achieve near-perfect adjacent-pair discrimination in-distribution. To make cross-method comparisons meaningful despite different score scales, we report the scale-normalized monotonicity slope \(\beta_{\text{norm}}\) rather than raw \(\beta\). Under this normalization, regression, classification, CORAL, and GON all show nearly perfect global ordering on the source distribution. This pattern indicates that source-domain performance alone is insufficient to reveal whether a learned score will transfer with a stable numerical meaning.
| Method | \(\beta_{\text{norm}}\) | \(\text{AUROC}_{01}\) | \(\text{AUROC}_{12}\) | \(\text{AUROC}_{23}\) | \(\text{AUROC}_{34}\) | \(\bar{\sigma}\) |
|---|---|---|---|---|---|---|
| Regression | 0.995 | 0.998 | 1.000 | 0.996 | 1.000 | 0.096 |
| Classification | 0.993 | 0.998 | 0.998 | 0.997 | 1.000 | 0.087 |
| CORAL | 0.994 | 0.999 | 0.999 | 0.996 | 1.000 | 0.091 |
| GON | 0.999 | 0.998 | 1.000 | 0.998 | 1.000 | 0.191 |
Table 6 reports zero-shot performance on the 5 held-out target systems. All supervised baselines cluster at similar values of \(\beta_{\text{norm}} \approx 0.46\)–\(0.48\) despite their different objectives, suggesting a shared limitation: none learns a score with a stable cross-system coordinate.
GON achieves the highest scale-normalized slope, \(\beta_{\text{norm}}=0.559\), while also remaining competitive on the strongest transferable boundary \(L_2 \rightarrow L_3\). This is the relevant comparison for cross-system ordinal structure, since raw \(\beta\) is not directly comparable across methods with different score scales.
| Method | \(\beta_{\text{norm}}\) | \(\text{AUROC}_{01}\) | \(\text{AUROC}_{12}\) | \(\text{AUROC}_{23}\) | \(\text{AUROC}_{34}\) | |
|---|---|---|---|---|---|---|
| Regression | 0.479 | 0.501 | 0.636 | 0.643 | 0.980 | |
| Classification | 0.463 | 0.640 | 0.541 | 0.617 | 0.984 | |
| CORAL | 0.461 | 0.648 | 0.606 | 0.724 | 0.992 | |
| GON | 0.559 | 0.494 | 0.583 | 0.715 | 0.973 |
Tables 7–16 report monotonicity slope \(\beta\) and all four adjacent-pair AUROC values for each of the five held-out systems, across \(k \in \{5,10,\ldots,100\}\) labeled windows. All entries are mean \(\pm\) std over 5 random seeds. Zero-shot (ZS) performance is listed in the caption of each system block for reference. Columns correspond to the three adaptation strategies: scratch (random initialization), pre_head (frozen encoder, readout fine-tuned), and pre_all (full model fine-tuned from the pretrained checkpoint).
All three adaptation strategies use AdamW with learning rate \(2\times10^{-3}\), weight decay \(10^{-4}\), cosine annealing over 30 epochs, and gradient clipping at 5.0; identical to pretraining. Data augmentation and EMA are disabled during fine-tuning. For pre_head, all encoder parameters are frozen including GroupNorm weight/bias; only the readout, normalization layer, and scale/bias embeddings are updated. The trajectory-level train/val/test split is fixed across all runs (seed 42). For each \((k, \text{seed})\) pair, \(k\) windows are drawn without replacement from the fixed training pool, with at least one window per ladder level before filling the remaining budget uniformly at random.
Zero-shot: \(\beta=0.737\), \(\text{AUROC}_{01}=0.539\), \(\text{AUROC}_{12}=0.514\), \(\text{AUROC}_{23}=0.856\), \(\text{AUROC}_{34}=1.000\).
| \(k\) | scratch | pre_head | pre_all |
|---|---|---|---|
| 5 | 0.730\(\pm\)0.270 | 0.852\(\pm\)0.067 | 1.328\(\pm\)0.263 |
| 10 | 0.650\(\pm\)0.300 | 0.892\(\pm\)0.056 | 1.443\(\pm\)0.223 |
| 15 | 0.954\(\pm\)0.180 | 0.938\(\pm\)0.065 | 1.541\(\pm\)0.198 |
| 20 | 0.942\(\pm\)0.157 | 0.950\(\pm\)0.086 | 1.496\(\pm\)0.167 |
| 25 | 0.785\(\pm\)0.303 | 0.917\(\pm\)0.038 | 1.415\(\pm\)0.285 |
| 30 | 0.765\(\pm\)0.400 | 0.959\(\pm\)0.026 | 1.615\(\pm\)0.163 |
| 35 | 0.830\(\pm\)0.268 | 0.959\(\pm\)0.055 | 1.592\(\pm\)0.142 |
| 40 | 1.157\(\pm\)0.225 | 0.935\(\pm\)0.083 | 1.588\(\pm\)0.146 |
| 45 | 0.903\(\pm\)0.294 | 0.910\(\pm\)0.131 | 1.649\(\pm\)0.219 |
| 50 | 1.032\(\pm\)0.253 | 0.996\(\pm\)0.020 | 1.636\(\pm\)0.129 |
| 55 | 0.611\(\pm\)0.232 | 0.908\(\pm\)0.084 | 1.782\(\pm\)0.084 |
| 60 | 0.956\(\pm\)0.327 | 0.971\(\pm\)0.066 | 1.568\(\pm\)0.221 |
| 65 | 0.904\(\pm\)0.441 | 0.940\(\pm\)0.073 | 1.666\(\pm\)0.180 |
| 70 | 1.109\(\pm\)0.296 | 1.018\(\pm\)0.023 | 1.597\(\pm\)0.141 |
| 75 | 1.127\(\pm\)0.376 | 0.977\(\pm\)0.021 | 1.658\(\pm\)0.077 |
| 80 | 0.919\(\pm\)0.299 | 0.932\(\pm\)0.072 | 1.664\(\pm\)0.148 |
| 85 | 0.965\(\pm\)0.213 | 0.943\(\pm\)0.047 | 1.675\(\pm\)0.134 |
| 90 | 1.011\(\pm\)0.248 | 0.965\(\pm\)0.035 | 1.703\(\pm\)0.135 |
| 95 | 0.857\(\pm\)0.357 | 0.974\(\pm\)0.063 | 1.654\(\pm\)0.087 |
| 100 | 1.133\(\pm\)0.148 | 0.974\(\pm\)0.029 | 1.620\(\pm\)0.104 |
| \(\text{AUROC}_{01}\) | \(\text{AUROC}_{12}\) | \(\text{AUROC}_{23}\) | \(\text{AUROC}_{34}\) | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2-4(lr)5-7(lr)8-10(lr)11-13 \(k\) | scr | ph | pa | scr | ph | pa | scr | ph | pa | scr | ph | pa |
| 5 | .620 | .581 | .619 | .609 | .502 | .480 | .486 | .903 | .919 | .911 | 1.00 | .997 |
| 10 | .668 | .589 | .636 | .596 | .506 | .536 | .467 | .901 | .927 | .887 | 1.00 | .985 |
| 15 | .656 | .594 | .643 | .630 | .506 | .526 | .465 | .907 | .900 | .960 | 1.00 | .997 |
| 20 | .749 | .591 | .639 | .641 | .508 | .548 | .435 | .920 | .941 | .995 | 1.00 | .983 |
| 25 | .714 | .603 | .604 | .630 | .506 | .526 | .451 | .919 | .907 | .989 | 1.00 | .973 |
| 30 | .732 | .599 | .663 | .630 | .510 | .547 | .461 | .930 | .921 | .998 | 1.00 | .974 |
| 35 | .748 | .601 | .641 | .647 | .507 | .525 | .458 | .929 | .961 | .998 | 1.00 | .994 |
| 40 | .739 | .607 | .677 | .646 | .506 | .520 | .443 | .931 | .893 | .998 | 1.00 | .978 |
| 45 | .739 | .594 | .677 | .654 | .506 | .563 | .432 | .922 | .921 | .995 | 1.00 | .975 |
| 50 | .725 | .604 | .669 | .638 | .510 | .550 | .455 | .933 | .923 | .990 | 1.00 | .980 |
| 55 | .716 | .608 | .678 | .643 | .503 | .536 | .452 | .922 | .888 | .974 | 1.00 | .957 |
| 60 | .756 | .603 | .681 | .663 | .512 | .548 | .439 | .923 | .901 | .998 | 1.00 | .977 |
| 65 | .737 | .601 | .672 | .636 | .505 | .527 | .463 | .921 | .898 | .997 | 1.00 | .980 |
| 70 | .747 | .600 | .677 | .647 | .509 | .571 | .457 | .931 | .938 | .999 | 1.00 | .978 |
| 75 | .756 | .598 | .670 | .663 | .512 | .584 | .445 | .925 | .918 | .998 | 1.00 | .976 |
| 80 | .734 | .603 | .687 | .648 | .505 | .566 | .449 | .920 | .925 | .993 | 1.00 | .976 |
| 85 | .749 | .594 | .688 | .652 | .510 | .616 | .450 | .918 | .892 | .998 | 1.00 | .954 |
| 90 | .726 | .606 | .701 | .633 | .512 | .558 | .463 | .927 | .904 | .997 | 1.00 | .974 |
| 95 | .735 | .606 | .675 | .653 | .509 | .598 | .447 | .931 | .897 | .998 | 1.00 | .986 |
| 100 | .738 | .605 | .690 | .632 | .511 | .538 | .459 | .929 | .932 | .997 | 1.00 | .984 |
Zero-shot: \(\beta=0.489\), \(\text{AUROC}_{01}=0.295\), \(\text{AUROC}_{12}=0.390\), \(\text{AUROC}_{23}=0.604\), \(\text{AUROC}_{34}=1.000\).
| \(k\) | scratch | pre_head | pre_all |
|---|---|---|---|
| 5 | 0.110\(\pm\)0.072 | 0.609\(\pm\)0.035 | 0.701\(\pm\)0.067 |
| 10 | 0.213\(\pm\)0.190 | 0.606\(\pm\)0.023 | 0.780\(\pm\)0.064 |
| 15 | 0.338\(\pm\)0.368 | 0.599\(\pm\)0.060 | 0.775\(\pm\)0.169 |
| 20 | 0.374\(\pm\)0.451 | 0.596\(\pm\)0.040 | 0.733\(\pm\)0.157 |
| 25 | 0.606\(\pm\)0.505 | 0.633\(\pm\)0.028 | 0.834\(\pm\)0.117 |
| 30 | 0.747\(\pm\)0.392 | 0.634\(\pm\)0.020 | 0.722\(\pm\)0.191 |
| 35 | 0.599\(\pm\)0.422 | 0.616\(\pm\)0.044 | 0.981\(\pm\)0.116 |
| 40 | 0.522\(\pm\)0.188 | 0.624\(\pm\)0.056 | 0.973\(\pm\)0.189 |
| 45 | 0.522\(\pm\)0.251 | 0.646\(\pm\)0.039 | 0.880\(\pm\)0.113 |
| 50 | 0.378\(\pm\)0.350 | 0.649\(\pm\)0.034 | 1.084\(\pm\)0.150 |
| 55 | 0.730\(\pm\)0.389 | 0.654\(\pm\)0.041 | 1.039\(\pm\)0.237 |
| 60 | 0.546\(\pm\)0.440 | 0.645\(\pm\)0.016 | 0.876\(\pm\)0.272 |
| 65 | 0.662\(\pm\)0.480 | 0.646\(\pm\)0.034 | 1.107\(\pm\)0.185 |
| 70 | 0.553\(\pm\)0.552 | 0.625\(\pm\)0.033 | 1.050\(\pm\)0.163 |
| 75 | 0.686\(\pm\)0.615 | 0.656\(\pm\)0.030 | 0.937\(\pm\)0.209 |
| 80 | 0.746\(\pm\)0.345 | 0.658\(\pm\)0.014 | 0.880\(\pm\)0.174 |
| 85 | 0.740\(\pm\)0.458 | 0.664\(\pm\)0.017 | 0.960\(\pm\)0.224 |
| 90 | 0.562\(\pm\)0.554 | 0.648\(\pm\)0.028 | 1.035\(\pm\)0.203 |
| 95 | 0.403\(\pm\)0.287 | 0.633\(\pm\)0.032 | 0.921\(\pm\)0.102 |
| 100 | 0.589\(\pm\)0.496 | 0.636\(\pm\)0.026 | 0.835\(\pm\)0.135 |
| \(\text{AUROC}_{01}\) | \(\text{AUROC}_{12}\) | \(\text{AUROC}_{23}\) | \(\text{AUROC}_{34}\) | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2-4(lr)5-7(lr)8-10(lr)11-13 \(k\) | scr | ph | pa | scr | ph | pa | scr | ph | pa | scr | ph | pa |
| 5 | .390 | .337 | .460 | .485 | .448 | .450 | .534 | .593 | .604 | .699 | 1.00 | 1.00 |
| 10 | .434 | .447 | .638 | .492 | .431 | .446 | .507 | .566 | .551 | .741 | 1.00 | 1.00 |
| 15 | .424 | .337 | .566 | .484 | .471 | .480 | .526 | .610 | .600 | .801 | 1.00 | .996 |
| 20 | .593 | .353 | .670 | .456 | .437 | .456 | .512 | .593 | .543 | .703 | 1.00 | .998 |
| 25 | .565 | .464 | .641 | .526 | .501 | .513 | .478 | .550 | .551 | .783 | 1.00 | 1.00 |
| 30 | .654 | .544 | .644 | .509 | .501 | .487 | .465 | .518 | .526 | .703 | 1.00 | .992 |
| 35 | .515 | .405 | .624 | .514 | .531 | .501 | .516 | .577 | .625 | .751 | 1.00 | .998 |
| 40 | .639 | .472 | .728 | .497 | .456 | .486 | .474 | .542 | .569 | .731 | 1.00 | .995 |
| 45 | .622 | .479 | .627 | .513 | .471 | .522 | .493 | .552 | .555 | .720 | 1.00 | 1.00 |
| 50 | .452 | .351 | .644 | .479 | .457 | .517 | .509 | .585 | .626 | .810 | 1.00 | .999 |
| 55 | .605 | .437 | .630 | .530 | .466 | .498 | .483 | .566 | .605 | .740 | 1.00 | .991 |
| 60 | .543 | .476 | .638 | .539 | .482 | .496 | .498 | .552 | .584 | .737 | 1.00 | 1.00 |
| 65 | .542 | .485 | .691 | .558 | .455 | .534 | .493 | .552 | .571 | .675 | 1.00 | 1.00 |
| 70 | .623 | .505 | .661 | .468 | .472 | .495 | .500 | .522 | .609 | .714 | 1.00 | .993 |
| 75 | .658 | .513 | .706 | .528 | .511 | .532 | .492 | .539 | .580 | .713 | 1.00 | .997 |
| 80 | .628 | .398 | .556 | .535 | .514 | .550 | .517 | .579 | .618 | .758 | 1.00 | .995 |
| 85 | .564 | .407 | .647 | .554 | .508 | .523 | .484 | .577 | .613 | .777 | 1.00 | .998 |
| 90 | .618 | .445 | .747 | .517 | .474 | .526 | .506 | .559 | .581 | .695 | 1.00 | .992 |
| 95 | .639 | .506 | .645 | .531 | .488 | .487 | .481 | .542 | .581 | .744 | 1.00 | .994 |
| 100 | .623 | .412 | .637 | .546 | .487 | .493 | .473 | .582 | .595 | .806 | 1.00 | 1.00 |
Zero-shot: \(\beta=1.270\), \(\text{AUROC}_{01}=0.483\), \(\text{AUROC}_{12}=0.576\), \(\text{AUROC}_{23}=0.725\), \(\text{AUROC}_{34}=0.872\).
| \(k\) | scratch | pre_head | pre_all |
|---|---|---|---|
| 5 | 0.009\(\pm\)0.050 | 1.153\(\pm\)0.058 | 1.256\(\pm\)0.143 |
| 10 | 0.218\(\pm\)0.278 | 1.203\(\pm\)0.075 | 1.400\(\pm\)0.250 |
| 15 | 0.254\(\pm\)0.315 | 1.179\(\pm\)0.016 | 1.345\(\pm\)0.211 |
| 20 | 0.244\(\pm\)0.209 | 1.281\(\pm\)0.075 | 1.503\(\pm\)0.110 |
| 25 | 0.280\(\pm\)0.255 | 1.263\(\pm\)0.130 | 1.402\(\pm\)0.255 |
| 30 | 0.277\(\pm\)0.206 | 1.226\(\pm\)0.106 | 1.460\(\pm\)0.239 |
| 35 | 0.128\(\pm\)0.105 | 1.274\(\pm\)0.126 | 1.550\(\pm\)0.175 |
| 40 | 0.180\(\pm\)0.237 | 1.247\(\pm\)0.080 | 1.506\(\pm\)0.192 |
| 45 | 0.468\(\pm\)0.478 | 1.226\(\pm\)0.113 | 1.480\(\pm\)0.207 |
| 50 | 0.536\(\pm\)0.409 | 1.270\(\pm\)0.076 | 1.519\(\pm\)0.184 |
| 55 | 0.430\(\pm\)0.399 | 1.240\(\pm\)0.094 | 1.682\(\pm\)0.128 |
| 60 | 0.155\(\pm\)0.285 | 1.394\(\pm\)0.065 | 1.639\(\pm\)0.114 |
| 65 | 0.410\(\pm\)0.373 | 1.379\(\pm\)0.061 | 1.609\(\pm\)0.197 |
| 70 | 0.537\(\pm\)0.418 | 1.352\(\pm\)0.068 | 1.720\(\pm\)0.080 |
| 75 | 0.425\(\pm\)0.488 | 1.295\(\pm\)0.073 | 1.519\(\pm\)0.200 |
| 80 | 0.601\(\pm\)0.453 | 1.289\(\pm\)0.123 | 1.644\(\pm\)0.202 |
| 85 | 0.421\(\pm\)0.280 | 1.329\(\pm\)0.044 | 1.621\(\pm\)0.192 |
| 90 | 0.836\(\pm\)0.414 | 1.344\(\pm\)0.041 | 1.598\(\pm\)0.158 |
| 95 | 0.379\(\pm\)0.250 | 1.251\(\pm\)0.057 | 1.445\(\pm\)0.187 |
| 100 | 0.105\(\pm\)0.115 | 1.282\(\pm\)0.092 | 1.591\(\pm\)0.175 |
| \(\text{AUROC}_{01}\) | \(\text{AUROC}_{12}\) | \(\text{AUROC}_{23}\) | \(\text{AUROC}_{34}\) | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2-4(lr)5-7(lr)8-10(lr)11-13 \(k\) | scr | ph | pa | scr | ph | pa | scr | ph | pa | scr | ph | pa |
| 5 | .428 | .490 | .519 | .554 | .576 | .557 | .498 | .725 | .766 | .455 | .883 | .895 |
| 10 | .484 | .492 | .541 | .480 | .575 | .568 | .523 | .726 | .780 | .602 | .875 | .900 |
| 15 | .495 | .496 | .519 | .509 | .578 | .573 | .541 | .727 | .757 | .524 | .879 | .905 |
| 20 | .552 | .494 | .514 | .519 | .581 | .579 | .525 | .734 | .812 | .562 | .883 | .919 |
| 25 | .548 | .495 | .523 | .484 | .578 | .570 | .590 | .731 | .769 | .517 | .870 | .885 |
| 30 | .535 | .483 | .536 | .452 | .577 | .562 | .589 | .723 | .792 | .597 | .869 | .906 |
| 35 | .522 | .501 | .520 | .495 | .579 | .571 | .556 | .740 | .819 | .529 | .878 | .905 |
| 40 | .485 | .495 | .538 | .546 | .580 | .580 | .521 | .739 | .795 | .529 | .880 | .914 |
| 45 | .600 | .494 | .501 | .488 | .581 | .590 | .633 | .727 | .791 | .535 | .898 | .919 |
| 50 | .543 | .498 | .529 | .465 | .581 | .581 | .590 | .738 | .798 | .614 | .887 | .903 |
| 55 | .548 | .510 | .558 | .500 | .581 | .584 | .588 | .736 | .825 | .561 | .886 | .916 |
| 60 | .491 | .500 | .524 | .477 | .582 | .587 | .538 | .739 | .807 | .572 | .885 | .909 |
| 65 | .558 | .505 | .583 | .493 | .579 | .581 | .604 | .736 | .816 | .546 | .884 | .907 |
| 70 | .543 | .506 | .577 | .515 | .581 | .584 | .624 | .740 | .828 | .518 | .886 | .916 |
| 75 | .540 | .493 | .523 | .522 | .582 | .594 | .557 | .736 | .778 | .547 | .880 | .921 |
| 80 | .570 | .501 | .570 | .481 | .580 | .567 | .613 | .735 | .822 | .597 | .883 | .908 |
| 85 | .583 | .500 | .560 | .523 | .584 | .586 | .574 | .737 | .799 | .574 | .887 | .911 |
| 90 | .579 | .498 | .534 | .481 | .582 | .596 | .638 | .739 | .799 | .627 | .885 | .924 |
| 95 | .546 | .503 | .542 | .480 | .578 | .576 | .590 | .734 | .781 | .569 | .880 | .902 |
| 100 | .511 | .499 | .555 | .502 | .578 | .586 | .511 | .730 | .789 | .574 | .881 | .908 |
Zero-shot: \(\beta=0.958\), \(\text{AUROC}_{01}=0.508\), \(\text{AUROC}_{12}=0.452\), \(\text{AUROC}_{23}=0.391\), \(\text{AUROC}_{34}=0.995\).
| \(k\) | scratch | pre_head | pre_all |
|---|---|---|---|
| 5 | 0.708\(\pm\)0.318 | 0.937\(\pm\)0.011 | 1.305\(\pm\)0.202 |
| 10 | 0.795\(\pm\)0.502 | 0.940\(\pm\)0.012 | 1.043\(\pm\)0.344 |
| 15 | 0.999\(\pm\)0.472 | 0.938\(\pm\)0.025 | 1.160\(\pm\)0.329 |
| 20 | 1.167\(\pm\)0.432 | 0.944\(\pm\)0.021 | 1.433\(\pm\)0.277 |
| 25 | 1.232\(\pm\)0.491 | 0.944\(\pm\)0.022 | 1.323\(\pm\)0.277 |
| 30 | 1.371\(\pm\)0.240 | 0.927\(\pm\)0.037 | 1.483\(\pm\)0.195 |
| 35 | 1.349\(\pm\)0.506 | 0.945\(\pm\)0.026 | 1.324\(\pm\)0.314 |
| 40 | 1.646\(\pm\)0.192 | 0.934\(\pm\)0.039 | 1.223\(\pm\)0.378 |
| 45 | 1.340\(\pm\)0.361 | 0.952\(\pm\)0.021 | 1.455\(\pm\)0.262 |
| 50 | 1.513\(\pm\)0.361 | 0.957\(\pm\)0.013 | 1.476\(\pm\)0.263 |
| 55 | 1.409\(\pm\)0.391 | 0.954\(\pm\)0.018 | 1.192\(\pm\)0.335 |
| 60 | 1.267\(\pm\)0.427 | 0.958\(\pm\)0.017 | 1.544\(\pm\)0.248 |
| 65 | 1.483\(\pm\)0.456 | 0.965\(\pm\)0.009 | 1.450\(\pm\)0.218 |
| 70 | 1.501\(\pm\)0.260 | 0.962\(\pm\)0.013 | 1.376\(\pm\)0.245 |
| 75 | 1.755\(\pm\)0.139 | 0.968\(\pm\)0.005 | 1.401\(\pm\)0.343 |
| 80 | 1.477\(\pm\)0.414 | 0.956\(\pm\)0.020 | 1.680\(\pm\)0.111 |
| 85 | 1.487\(\pm\)0.234 | 0.953\(\pm\)0.016 | 1.499\(\pm\)0.294 |
| 90 | 1.709\(\pm\)0.153 | 0.963\(\pm\)0.007 | 1.603\(\pm\)0.252 |
| 95 | 1.694\(\pm\)0.225 | 0.951\(\pm\)0.017 | 1.494\(\pm\)0.249 |
| 100 | 1.654\(\pm\)0.152 | 0.949\(\pm\)0.025 | 1.364\(\pm\)0.318 |
| \(\text{AUROC}_{01}\) | \(\text{AUROC}_{12}\) | \(\text{AUROC}_{23}\) | \(\text{AUROC}_{34}\) | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2-4(lr)5-7(lr)8-10(lr)11-13 \(k\) | scr | ph | pa | scr | ph | pa | scr | ph | pa | scr | ph | pa |
| 5 | .607 | .555 | .836 | .623 | .540 | .711 | .494 | .364 | .506 | .763 | .997 | .979 |
| 10 | .659 | .523 | .618 | .595 | .495 | .539 | .546 | .380 | .557 | .833 | .995 | .989 |
| 15 | .787 | .536 | .736 | .634 | .551 | .634 | .524 | .372 | .561 | .844 | .996 | .977 |
| 20 | .753 | .542 | .857 | .806 | .541 | .697 | .523 | .369 | .611 | .819 | .996 | .982 |
| 25 | .799 | .538 | .781 | .712 | .530 | .610 | .527 | .390 | .607 | .859 | .997 | .984 |
| 30 | .861 | .537 | .920 | .747 | .540 | .656 | .530 | .385 | .615 | .902 | .997 | .959 |
| 35 | .839 | .528 | .764 | .815 | .514 | .579 | .532 | .381 | .543 | .884 | .996 | .982 |
| 40 | .817 | .558 | .744 | .825 | .568 | .645 | .533 | .383 | .650 | .893 | .998 | .977 |
| 45 | .803 | .555 | .820 | .754 | .520 | .610 | .539 | .355 | .579 | .893 | .997 | .987 |
| 50 | .875 | .543 | .820 | .819 | .539 | .573 | .530 | .363 | .632 | .904 | .996 | .988 |
| 55 | .818 | .536 | .698 | .707 | .524 | .561 | .529 | .374 | .571 | .862 | .995 | .980 |
| 60 | .931 | .557 | .863 | .778 | .527 | .553 | .507 | .355 | .601 | .850 | .997 | .970 |
| 65 | .812 | .527 | .820 | .700 | .521 | .603 | .592 | .370 | .666 | .897 | .995 | .979 |
| 70 | .868 | .554 | .816 | .707 | .529 | .625 | .557 | .363 | .605 | .884 | .997 | .977 |
| 75 | .911 | .545 | .731 | .802 | .529 | .571 | .519 | .361 | .630 | .876 | .996 | .985 |
| 80 | .882 | .535 | .887 | .781 | .533 | .611 | .535 | .375 | .694 | .884 | .996 | .972 |
| 85 | .874 | .551 | .863 | .741 | .535 | .609 | .549 | .360 | .601 | .883 | .997 | .979 |
| 90 | .943 | .546 | .881 | .786 | .548 | .610 | .525 | .361 | .646 | .890 | .996 | .977 |
| 95 | .966 | .553 | .857 | .828 | .556 | .625 | .518 | .364 | .652 | .907 | .997 | .980 |
| 100 | .920 | .534 | .822 | .786 | .530 | .576 | .540 | .376 | .602 | .904 | .996 | .994 |
Zero-shot: \(\beta=2.105\), \(\text{AUROC}_{01}=0.645\), \(\text{AUROC}_{12}=0.985\), \(\text{AUROC}_{23}=0.997\), \(\text{AUROC}_{34}=1.000\).
| \(k\) | scratch | pre_head | pre_all |
|---|---|---|---|
| 5 | 1.157\(\pm\)0.260 | 2.058\(\pm\)0.083 | 2.003\(\pm\)0.082 |
| 10 | 1.445\(\pm\)0.352 | 2.035\(\pm\)0.068 | 2.041\(\pm\)0.036 |
| 15 | 1.502\(\pm\)0.455 | 2.086\(\pm\)0.051 | 2.009\(\pm\)0.089 |
| 20 | 1.682\(\pm\)0.258 | 1.991\(\pm\)0.086 | 2.028\(\pm\)0.083 |
| 25 | 1.885\(\pm\)0.168 | 1.987\(\pm\)0.059 | 1.950\(\pm\)0.111 |
| 30 | 1.838\(\pm\)0.158 | 2.018\(\pm\)0.090 | 2.043\(\pm\)0.073 |
| 35 | 1.724\(\pm\)0.213 | 2.000\(\pm\)0.049 | 1.984\(\pm\)0.066 |
| 40 | 1.965\(\pm\)0.108 | 2.008\(\pm\)0.045 | 1.976\(\pm\)0.056 |
| 45 | 1.818\(\pm\)0.492 | 2.024\(\pm\)0.057 | 2.038\(\pm\)0.069 |
| 50 | 1.817\(\pm\)0.157 | 2.015\(\pm\)0.053 | 2.038\(\pm\)0.024 |
| 55 | 1.757\(\pm\)0.224 | 2.069\(\pm\)0.055 | 2.033\(\pm\)0.052 |
| 60 | 1.723\(\pm\)0.132 | 2.021\(\pm\)0.087 | 2.011\(\pm\)0.058 |
| 65 | 1.720\(\pm\)0.404 | 2.066\(\pm\)0.023 | 2.048\(\pm\)0.026 |
| 70 | 1.782\(\pm\)0.110 | 2.011\(\pm\)0.025 | 1.970\(\pm\)0.068 |
| 75 | 1.942\(\pm\)0.146 | 2.031\(\pm\)0.046 | 2.025\(\pm\)0.033 |
| 80 | 1.865\(\pm\)0.137 | 2.009\(\pm\)0.047 | 2.017\(\pm\)0.046 |
| 85 | 1.917\(\pm\)0.152 | 2.053\(\pm\)0.049 | 2.017\(\pm\)0.025 |
| 90 | 1.838\(\pm\)0.328 | 2.015\(\pm\)0.060 | 2.056\(\pm\)0.032 |
| 95 | 1.859\(\pm\)0.278 | 2.058\(\pm\)0.035 | 2.070\(\pm\)0.052 |
| 100 | 1.829\(\pm\)0.256 | 2.045\(\pm\)0.059 | 2.034\(\pm\)0.044 |
| \(\text{AUROC}_{01}\) | \(\text{AUROC}_{12}\) | \(\text{AUROC}_{23}\) | \(\text{AUROC}_{34}\) | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2-4(lr)5-7(lr)8-10(lr)11-13 \(k\) | scr | ph | pa | scr | ph | pa | scr | ph | pa | scr | ph | pa |
| 5 | .756 | .744 | .918 | .902 | .987 | .994 | .823 | .998 | .995 | .965 | 1.00 | 1.00 |
| 10 | .796 | .824 | .931 | .936 | .986 | .995 | .785 | .998 | .996 | .960 | 1.00 | 1.00 |
| 15 | .848 | .791 | .936 | .976 | .986 | .994 | .784 | .998 | .995 | .997 | 1.00 | 1.00 |
| 20 | .772 | .850 | .917 | .885 | .988 | .994 | .834 | .998 | .995 | .995 | 1.00 | 1.00 |
| 25 | .871 | .882 | .930 | .953 | .987 | .996 | .788 | .998 | .991 | .971 | 1.00 | .999 |
| 30 | .891 | .853 | .931 | .976 | .984 | .995 | .811 | .998 | .994 | .981 | 1.00 | .997 |
| 35 | .814 | .867 | .972 | .952 | .983 | .990 | .801 | .998 | .997 | .994 | 1.00 | .995 |
| 40 | .947 | .872 | .941 | .991 | .986 | .991 | .779 | .998 | .996 | .978 | 1.00 | .999 |
| 45 | .911 | .904 | .945 | .977 | .984 | .996 | .842 | .998 | .994 | .988 | 1.00 | .993 |
| 50 | .943 | .881 | .964 | .994 | .985 | .992 | .818 | .998 | .995 | .987 | 1.00 | .998 |
| 55 | .994 | .886 | .949 | .998 | .983 | .998 | .774 | .998 | .992 | .991 | 1.00 | .999 |
| 60 | .861 | .868 | .955 | .944 | .985 | .992 | .788 | .998 | .996 | .985 | 1.00 | 1.00 |
| 65 | .925 | .854 | .951 | .999 | .984 | .993 | .834 | .998 | .995 | .995 | 1.00 | 1.00 |
| 70 | .872 | .888 | .939 | .956 | .986 | .992 | .843 | .998 | .995 | .986 | 1.00 | .999 |
| 75 | .987 | .878 | .959 | .998 | .984 | .997 | .800 | .998 | .996 | .993 | 1.00 | .998 |
| 80 | .897 | .890 | .962 | .972 | .986 | .996 | .810 | .998 | .995 | .988 | 1.00 | .998 |
| 85 | .915 | .880 | .973 | .984 | .984 | .997 | .820 | .998 | .995 | .995 | 1.00 | .996 |
| 90 | .892 | .892 | .961 | .978 | .983 | .999 | .795 | .998 | .994 | .994 | 1.00 | .996 |
| 95 | .951 | .880 | .955 | .982 | .986 | .998 | .814 | .998 | .995 | .993 | 1.00 | .998 |
| 100 | .908 | .883 | .951 | .981 | .985 | .998 | .826 | .998 | .994 | .993 | 1.00 | .995 |
Table 17 reports the complete inference-time perturbation sweep on the 12 source systems (in-distribution). No retraining is performed; the pretrained GON checkpoint is applied directly.
All perturbations are applied at inference only under torch.no_grad(). Gaussian noise: SNR is defined on a mean-square power basis, \(\mathrm{SNR}_\mathrm{dB} = 10\log_{10}(P_s/P_n)\) where
\(P_s = \mathrm{mean}(x^2)\) over all channels and timesteps jointly per sample; a single noise power scalar is shared across channels. Temporal dropout: selected timesteps are zero-masked in-place
(sequence length preserved), with the same mask applied across all channels simultaneously; each sample draws an independent mask. Quantization: post-training uniform asymmetric quantization applied per channel per sample (min-max over the
time axis) in pure PyTorch; zero is a representable level (uniform mid-tread).
(1) Ordinal coherence is preserved down to 20 dB SNR and falls below \(1.0\) between 30 and 20 dB; at 5 dB, \(\beta=0.216\), but the score remains positively monotone on average. (2) Degradation is boundary-asymmetric throughout: \(\text{AUROC}_{34}\) holds at \(1.000\) through 30 dB Gaussian noise and all dropout fractions up to 30%, while \(\text{AUROC}_{01}\) falls below \(0.5\) at 20 dB and at 20% dropout; \(\text{AUROC}_{12}\) is intermediate, tracking \(\text{AUROC}_{01}\) in shape but degrading more slowly. (3) Quantization is negligible at all tested precisions: no metric changes by more than \(0.002\) between 16-bit and 4-bit, and \(\beta\) is unchanged to three decimal places down to 6-bit.
| Type | Condition | \(\beta\) | \(\mathrm{AUROC}_{01}\) | \(\mathrm{AUROC}_{12}\) | \(\mathrm{AUROC}_{23}\) | \(\mathrm{AUROC}_{34}\) |
|---|---|---|---|---|---|---|
| clean | 2.065 | 0.965 | 0.993 | 0.994 | 1.000 | |
| Gaussian | 40 dB | 1.946 | 0.932 | 0.944 | 0.997 | 1.000 |
| 30 dB | 1.461 | 0.618 | 0.802 | 0.998 | 1.000 | |
| 20 dB | 0.882 | 0.449 | 0.539 | 0.892 | 0.998 | |
| 10 dB | 0.522 | 0.428 | 0.480 | 0.680 | 0.824 | |
| 5 dB | 0.216 | 0.446 | 0.456 | 0.574 | 0.702 | |
| Quantization | 16-bit | 2.065 | 0.965 | 0.993 | 0.994 | 1.000 |
| 12-bit | 2.065 | 0.965 | 0.993 | 0.994 | 1.000 | |
| 8-bit | 2.065 | 0.965 | 0.993 | 0.994 | 1.000 | |
| 6-bit | 2.065 | 0.965 | 0.993 | 0.994 | 1.000 | |
| 4-bit | 2.064 | 0.966 | 0.993 | 0.993 | 1.000 | |
| Dropout | 5% | 1.702 | 0.727 | 0.881 | 0.993 | 1.000 |
| 10% | 1.405 | 0.590 | 0.763 | 0.979 | 1.000 | |
| 20% | 1.091 | 0.450 | 0.692 | 0.943 | 1.000 | |
| 30% | 0.877 | 0.372 | 0.636 | 0.896 | 0.999 |