May 27, 2026
Federated Conformal RAG (FC-RAG) [1] provides distribution-free coverage for a bandwidth-limited swarm of weak language models, but only at a fixed horizon. We extend it to anytime-valid sequential coverage: validity at every stopping time, preserved under predictable adaptive control (recalibration, per-node bandwidth escalation, distilled-student refresh), at no extra cost in assumptions over fixed-horizon FC-RAG. Naive composition of fixed-horizon FC-RAG with off-the-shelf sequential testing fails because FC-RAG’s marginal coverage bound makes the natural betting e-process a non-supermartingale on adverse calibration draws, and Ville’s inequality cannot be invoked. We give Anytime-FC-RAG, a sequential extension built on a summable per-step calibration-deviation budget that converts the marginal bound into a strict conditional bound on a calibration-good event, paired with a truncated betting e-process that is a nonnegative supermartingale on the entire probability space. From these two ingredients, we obtain four guarantees: time-uniform alarm validity \(\mathbb{P}(\sup_t E_t \ge 1/\delta_e) \le \delta_e + \delta_{\mathrm{cal}}\), a Hoeffding-stitched cumulative-miscoverage envelope at the same total budget, safety under any predictable controller (recalibration, bandwidth escalation, student refresh), and training-side error propagation across an unbounded sequence of Federated Probe-Logit Distillation (FPLD) refreshes via a summable training budget. As a practical consequence, an adaptive controller that escalates retrieval bandwidth only when the e-process crosses a warning threshold matches the alarm rate of a fixed-high-bandwidth schedule at substantially lower communication cost. Synthetic and end-to-end experiments on a GPT-2-small + MiniLM swarm across MMLU, DBpedia, and AG News verify the predicted alarm rate, detection delay, envelope coverage, and \(14\)–\(57\%\) bandwidth savings; the alarm fires when and only when coverage genuinely breaks, not on every drift.
Federated Conformal RAG (FC-RAG) [1] provides a distribution-free coverage guarantee for the answer set produced by a bandwidth-limited swarm of weak language models, but only at a fixed horizon. Real deployments are sequential: operators inspect coverage across the answered-query stream, recalibrate when the rolling buffer freshens, and may escalate per-node retrieval bandwidth or refresh a distilled student in response to observed drift. The fixed-horizon guarantee does not apply once any of these actions is taken. This paper asks how to deliver time-uniform coverage and safe adaptive control on top of FC-RAG without strengthening the i.i.d.-deployment-data assumption that fixed-horizon FC-RAG already relies on, and how to propagate the underlying federated training rate across an unbounded sequence of student refreshes while preserving the sequential guarantee.
As a running example throughout this paper, consider the \(K=4\) topic-specialized GPT-2-small + MiniLM retrieval swarm of [1]: four nodes serving MMLU subjects high_school_statistics, high_school_physics, high_school_biology, and high_school_world_history, each retrieving from its own
subject-specific corpus and uploading bandwidth-limited score summaries to a hub that emits a conformal answer set at level \(1-\alpha = 0.9\). A reasonable operational goal is to monitor coverage across the answered-query
stream and, if the corpus on high_school_biology drifts after an update, fire an alarm and selectively escalate that node’s bandwidth without forfeiting the statistical guarantee already accumulated on the other three subjects.
This operational goal is harder than a naive composition of fixed-horizon FC-RAG and off-the-shelf sequential testing would suggest. The natural betting e-process built on FC-RAG’s marginal coverage bound fails the supermartingale property on adverse calibration draws, so Ville’s inequality cannot be applied off-the-shelf. Adaptive controller actions further change the test centering mid-stream in a way that compounds the difficulty, and the underlying student model may be refreshed multiple times during deployment, so training-side error must propagate cleanly across an unbounded sequence of refreshes. The construction we introduce closes all three gaps via a summable per-step calibration-deviation budget paired with a truncated supermartingale, and the resulting guarantees apply to any sequential conformal protocol with a predictable slack decomposition, beyond the RAG instance considered here (Section 6).
Three lines of work each cover a piece of this question but leave a gap.
Fixed-horizon FC-RAG. The base FC-RAG protocol [1] controls a one-shot marginal coverage guarantee; sequential monitoring with optional stopping or adaptive intervention breaks the guarantee.
Single-site conformal RAG. TRAQ [2] and conformal-RAG-style work [3] apply split-conformal to RAG QA pipelines but assume one model, one corpus, and one calibration set, with no federation, no bandwidth charging, and no sequential validity.
Conformal test martingales. Conformal test martingales [4] give a time-uniform changepoint-detection construction over exchangeable data, but in a single-site setting with no slack decomposition (no \(\Delta_{\mathrm{FL}}\), \(\Delta_{\mathrm{RAG}}\), \(\Delta_{\mathrm{train}}\)) and no model for federated calibration with bandwidth budgets.
Centralized sequential testing tools [5]–[12] and federated conformal methods [13]–[16] cover related ground but not this composition; no prior result gives a time-uniform reliability guarantee for a federated LLM swarm whose retrieval and calibration messages are simultaneously communication-constrained. Table 1 summarizes the joint capability gap.
| Approach | |||||
| uniform | |||||
| charged slack | |||||
| calibration | |||||
| propagation | |||||
| controller | |||||
| FC-RAG [1] |
\(\times\) | ✔ | ✔ | ✔ | \(\times\) |
| Conformal test martingales [4] |
✔ | \(\times\) | \(\times\) | \(\times\) | \(\times\) |
| Online conformal [11], [12] |
✔\(^{*}\) | \(\times\) | \(\times\) | \(\times\) | |
| Federated conformal [13]–[16] |
\(\times\) | partial | ✔ | \(\times\) | \(\times\) |
| Anytime-FC-RAG (ours) | ✔ | ✔ | ✔ | ✔ | ✔ |
\(^{*}\)
Long-run average coverage via adaptive step size; not strictly per-step time-uniform.
We adopt three operational constraints inherited from FC-RAG [1]: (i) no gradient or weight exchange; (ii) no data pooling; (iii) per-uplink budgets \(B_{i,t}\) (per-query inference) and \(B_t^{\mathrm{cal}}\) (per-refresh calibration) are first-class. The goal is to characterize whether a sequential extension can admit a strict conditional bound \(\mathbb{E}[M_t \mid \mathcal{F}_{t-1}] \le b_t = \alpha + 1/(n_{\mathrm{cal},t}+1) + \Delta_{\mathrm{FL},t} + \Delta_{\mathrm{RAG},t} + \Delta_{\mathrm{train},t}\) on a calibration-good event \(G_t\), with the same slack decomposition as fixed-horizon FC-RAG. We further require that this bound be preserved under any predictable controller (recalibration, bandwidth escalation, student refresh). Our aim is to characterize what is provably achievable, not to demonstrate a deployment-ready system.
Anytime-FC-RAG protocol (Section 3): a sequential extension of FC-RAG in which a swarm answers a stream of queries, updates compressed calibration summaries on a rolling buffer, and maintains a betting e-process for alarm-triggered intervention.
Cal-deviation budget and alarm validity (Lemmas 1, 2, Theorem 1): a summable per-step budget \(\{\delta_t^{\mathrm{cal}}\}\) defines an \(\mathcal{F}_{t-1}\)-measurable calibration-good event \(G_t\) on which the per-step miscoverage admits a strict conditional bound; the truncated betting e-process \(\widetilde{E}_t = E_t\,\mathbf{1}_{\bigcap_{s\le t} G_s}\) is then a nonnegative supermartingale on the entire probability space, and Ville plus splitting on \(G_t\) gives \(\mathbb{P}(\sup_t E_t \ge 1/\delta_e) \le \delta_e + \delta_{\mathrm{cal}}\).
Cumulative-miscoverage envelope and safe adaptive control (Theorems 3, 5): a time-uniform Hoeffding boundary \(u_t(\delta)\) controls the empirical miscoverage rate against the predictable slack at probability \(\ge 1-\delta\), and any predictable controller preserves both this envelope and the alarm guarantee.
Training-to-deployment propagation (Theorem 6): an FTC chain inheriting the (B5’) clause of base FC-RAG [1] gives \(\Delta_{\mathrm{train},t} \le f_{\max,t}(\bar K_t + \sqrt{2\bar K_t})\) simultaneously over \(t\) on an event of probability \(\ge 1 - \delta_{\mathrm{train}}\), for a summable budget \(\sum_r \delta_r \le \delta_{\mathrm{train}}\).
This work establishes a time-uniform reliability guarantee for federated LLM swarms with bandwidth-constrained retrieval and calibration, paired with end-to-end empirical validation on a \(K=4\) GPT-2-small + MiniLM swarm. Privacy accounting, delayed-label handling, adversarial-node and architecture-heterogeneous regimes, and deployment-scale benchmarking are natural extensions of the same machinery and are deferred to follow-up work.
Section 2 specifies the sequential deployment model. Section 3 states the Anytime-FC-RAG protocol. Section 4 proves the four theorems (alarm validity, envelope, safe control, training propagation). Section 5 reports synthetic, real-world, and comparative experiments validating the four theorems on a GPT-2-small + MiniLM swarm across three benchmarks. Section 6 situates the contribution against prior work and discusses limitations.
We study sequential discrete-answer prediction in a federated swarm of weak language models constrained by a bandwidth budget. This section fixes the data, communication, filtration, and estimand formalism that the planned theorems will refer to; the concrete protocol sits in Section 3.
Let \(\mathcal{X}\) denote the query or context space and let \(\mathcal{Y}\) be a finite answer space (multiple-choice-style QA, label prediction, or a bounded candidate set extracted from a top-\(p\) truncation of a language model). Let \(\mathcal{V}\) be the token vocabulary of the underlying language model. There are \(K\) nodes; node \(i \in \{1,\dots,K\}\) holds a local retrieval corpus \(C_{i,t}\) and a local retrieval mechanism, and raw node-local data and corpora never leave their node. A global student model \(\widehat P^{(0)}\) is available at deployment start, typically obtained from a federated training stage; we use the Federated Probe-Logit Distillation (FPLD) protocol of [1] as the running training stage, where each node fine-tunes locally and exchanges \(B\)-bit quantized logits on a shared probe set rather than gradients or weights. Theorem 1 of [1] gives an explicit high-probability KL rate for FPLD in \((K,n,m,B,V)\), which we reuse as a black-box rate input for Theorem 6. During deployment, the active student is denoted \(\widehat P_t\) and may be refreshed at selected intervention times.
At each time \(t \in \mathbb{N}\), a query-answer pair \((X_t,Y_t) \in \mathcal{X}\times \mathcal{Y}\) is realized; the operator does not commit to a terminal time in advance, and may inspect, recalibrate, or escalate after every query. We assume immediate label feedback in the main formulation: the true answer \(Y_t\) becomes available before the next query arrives, so revealed labels drive the monitoring process. Arbitrarily delayed labels are deferred to future work. At time \(t\), node \(i\) retrieves a top-\(k_{i,t}\) passage set \(Z_{i,t} = R_{i,t}(X_t, C_{i,t}; k_{i,t}) \subseteq C_{i,t}\); corpora may evolve slowly and are not shared across nodes. Bandwidth is the only resource we charge, and we charge uplink only: at time \(t\), node \(i\) uploads a \(B_{i,t}\)-bit summary of its local scores; at calibration-refresh times \(t \in \mathcal{T}_{\mathrm{cal}}\), node \(i\) also uploads a compressed calibration summary within a total budget \(B_t^{\mathrm{cal}}\); retraining inherits its cost from the underlying training protocol (e.g., FPLD). The per-query inference plus calibration cost is \(\Gamma_t^{\mathrm{comm}} = \sum_{i=1}^K B_{i,t} + B_t^{\mathrm{cal}}\,\mathbf{1}\{t \in \mathcal{T}_{\mathrm{cal}}\}\).
Given query \(X_t\), node \(i\) forms a candidate list \(A_{i,t}(X_t) \subseteq \mathcal{Y}\) and local nonconformity scores \(s_{i,t}(y) = -\log \widehat P_t(y \mid X_t, Z_{i,t})\) for \(y \in A_{i,t}(X_t)\), clipped to \([0, S_{\max}]\). Node \(i\) uploads a compressed message \(U_{i,t} = Q_{B_{i,t}}\big(\{(y,s_{i,t}(y)) : y \in A_{i,t}(X_t)\}\big)\), the hub decodes \(U_{i,t}\) into approximate scores \(\widetilde{s}_{i,t}(y)\), and aggregates them. With \(K_{t,y} = |\{i : y \in A_{i,t}(X_t)\}|\), \[s_t^{\star}(y) \;=\; \frac{1}{K_{t,y}} \sum_{i:\, y \in A_{i,t}(X_t)} s_{i,t}(y), \qquad s_t^{\mathrm{swarm}}(y) \;=\; \frac{1}{K_{t,y}} \sum_{i:\, y \in A_{i,t}(X_t)} \widetilde{s}_{i,t}(y),\] where \(s_t^{\mathrm{swarm}}(y)=+\infty\) if \(K_{t,y}=0\). The first is the oracle uncompressed swarm score; the second is what the hub actually sees. Inheriting candidate-set inclusion (Assumption B2 of [1]), every \(y\) in the test or calibration support lies in \(\bigcap_{i=1}^K A_{i,t}(X_t)\), i.e., \(K_{t,y} = K\) uniformly; the analysis lifts to general \(K_{t,y}\) at the cost of \(y\)-dependent variance constants.
We maintain a rolling labeled calibration buffer \(\mathcal{D}^{\mathrm{cal}}_t = \{(X_s,Y_s) : s \in I_t^{\mathrm{cal}}\}\) with \(n_{\mathrm{cal},t} := |I_t^{\mathrm{cal}}|\), where \(I_t^{\mathrm{cal}}\) is a predictable index set (e.g.a sliding window of the most recent answered queries or a batched refresh buffer). At a refresh time \(t \in \mathcal{T}_{\mathrm{cal}}\), node \(i\) recomputes its local scores on \(\mathcal{D}^{\mathrm{cal}}_t\), compresses the resulting score summary into its allotted share of \(B_t^{\mathrm{cal}}\) bits, and uploads it. Let \(q_t^\star\) denote the oracle conformal threshold from the full uncompressed calibration scores, and \(\widehat q_t\) the threshold reconstructed from compressed node summaries; the swarm outputs the implemented prediction set \(C_t(X_t) = \{y \in \mathcal{Y}: s_t^{\mathrm{swarm}}(y) \le \widehat q_t\}\) with miscoverage indicator \(M_t = \mathbf{1}\{Y_t \notin C_t(X_t)\}\).
Let \(\mathcal{F}_t = \sigma\big(X_{1:t}, Y_{1:t}, \widehat P_{0:t}, \{C_{i,1:t}\}_i, \{B_{i,1:t}\}_i, \{U_{i,1:t}\}_i, \widehat q_{1:t}, \mathcal{A}_{1:t}\big)\) be the observable history, where \(\mathcal{A}_t\) is the intervention action at time \(t\). A stopping time \(\tau\) is any \(\mathcal{F}_t\)-adapted random time. The core modeling restriction is predictability: any bandwidth schedule, recalibration schedule, or retraining trigger used at time \(t\) is \(\mathcal{F}_{t-1}\)-measurable. This is what lets monitoring and intervention coexist without breaking validity.
Three slack terms enter the per-step coverage bound: federated calibration, retrieval bandwidth, and training-side approximation. Each is \(\mathcal{F}_{t-1}\)-measurable by predictability of the associated schedule. We restate each at the level of detail required to follow the analysis of Section 4; full constructions and constants are in [1].
Inheriting Assumption (B3) of [1], each node’s QuantizeScore primitive is a subtractively dithered scalar quantizer, so that the per-node quantized score decomposes additively as \(\widetilde{s}_{i,t}(X_t,y) = s_{i,t}(X_t,y) + \xi_{i,t}(X_t,y)\), with \(\{\xi_{i,t}\}_{i=1}^K\) conditionally independent across nodes given \((\mathcal{F}_{t-1}, X_t)\), mean zero, and bounded conditional second moment \(\mathbb{E}[\xi_{i,t}^2 \mid \mathcal{F}_{t-1}, X_t] \le v(B_{i,t})\) where \(v(B) = O(2^{-2B/b_s})\) for a protocol-specific scale \(b_s\). By the average-aggregation rule, the implemented swarm score satisfies \(s_t^{\mathrm{swarm}}(X_t,y) = s_t^\star(X_t,y) + \bar\xi_t(X_t,y)\) with \(\bar\xi_t := (1/K)\sum_i \xi_{i,t}\), and cross-node independence yields \(\mathbb{E}[\bar\xi_t^2 \mid \mathcal{F}_{t-1}, X_t] \le V_{K,t} := (1/K^2) \sum_i v(B_{i,t})\). Combined with the \(f_{\max,t}\)-Lipschitz score CDF (Assumption 3, (B5’)) and Cauchy–Schwarz on \(\mathbb{E}|\bar\xi_t| \le \sqrt{V_{K,t}}\), the per-step retrieval-bandwidth slack is \[\Delta_{\mathrm{RAG},t} \;=\; f_{\max,t}\,\sqrt{\frac{1}{K^2}\sum_{i=1}^K v(B_{i,t})}.\] The \(1/K^2\) averaging (rather than the worst-case \(1/K\)) is the variance gain from independent cross-node dithering.
The threshold \(\widehat q_t\) deviates from the population \((1-\alpha)\)-quantile \(q_t^{\mathrm{pop}}\) via a deterministic compression piece \(|\widehat q_t - q_t^\star| \le \phi(B_t^{\mathrm{cal}}) = O(2^{-B_t^{\mathrm{cal}}/b_q})\) (e.g., a quantized order-statistics summary or a GC-FCP / Fed-CCP coreset [15], [16]) and a statistical piece \(|q_t^\star - q_t^{\mathrm{pop}}|\) from the i.i.d.buffer. To control the latter pointwise rather than only on average, we fix a summable per-step calibration budget \(\{\delta_t^{\mathrm{cal}}\}_{t \ge 1}\) with \(\sum_{t \ge 1} \delta_t^{\mathrm{cal}} \le \delta_{\mathrm{cal}} \in (0,1)\); the canonical choice is \(\delta_t^{\mathrm{cal}} = 6\delta_{\mathrm{cal}}/(\pi^2 t^2)\). Sub-Gaussian deviation of the empirical \((1-\alpha)\)-quantile gives, with probability at least \(1 - \delta_t^{\mathrm{cal}}\), \[\label{eq:cal-deviation} \big|\widehat q_t - q_t^{\mathrm{pop}}\big| \;\le\; c_q\,\sqrt{\log(2/\delta_t^{\mathrm{cal}})\big/n_{\mathrm{cal},t}} \;+\; \phi(B_t^{\mathrm{cal}}),\tag{1}\] for an absolute constant \(c_q\). Pushing 1 through the score CDF yields \[\label{eq:delta-fl} \Delta_{\mathrm{FL},t} \;=\; f_{\max,t}\!\left(c_q\,\sqrt{\log(2/\delta_t^{\mathrm{cal}})\big/n_{\mathrm{cal},t}} \;+\; \phi(B_t^{\mathrm{cal}})\right).\tag{2}\] The event \(G_t\) that 1 holds at time \(t\) is \(\mathcal{F}_{t-1}\)-measurable and has \(\mathbb{P}(G_t) \ge 1 - \delta_t^{\mathrm{cal}}\); the cumulative cal-good event \(\Omega_{\mathrm{cal}} := \bigcap_{t \ge 1} G_t\) satisfies \(\mathbb{P}(\Omega_{\mathrm{cal}}) \ge 1 - \delta_{\mathrm{cal}}\). The conditional theorems below take the form “on \(G_t\)” or “on \(\Omega_{\mathrm{cal}}\)” and absorb \(\delta_{\mathrm{cal}}\) into the final probability budget. The conditional bound 2 is strictly weaker than the marginal-over-calibration form of [1] Theorem 2 by a factor \(\sqrt{\log(2/\delta_t^{\mathrm{cal}})/\log(2/\delta_{\mathrm{cal}})}\), which grows from \(\approx 1.07\times\) at \(t = 1\) to \(\approx 2.48\times\) at \(t = 10^4\) for \(\delta_{\mathrm{cal}} = 0.05\), a slow \(O(\sqrt{\log t})\) price for converting marginal validity into pointwise-conditional validity.
Let \(\varepsilon_{\mathrm{train},t}\) denote the current training-side approximation level, an upper bound on \(\mathbb{E}_X[\mathrm{KL}(P^\star(\cdot\mid X)\,\|\,\widehat P_t(\cdot\mid X))]\), and let \(\Delta_t(x,y) := s_t^\star(x,y) - s^\star_{\mathrm{ideal}}(x,y) = \log(P^\star(y\mid x)/\widehat P_t(y\mid x))\) be the score-level training residual. The indicator-difference + FTC + Pinsker chain of [1], under the (B5’) conditional-density clause inherited in Assumption 3, yields the two-term bound \[\Delta_{\mathrm{train},t} \;=\; f_{\max,t}\,\big(\varepsilon_{\mathrm{train},t} \;+\; \sqrt{2\,\varepsilon_{\mathrm{train},t}}\big).\] In the small-\(\varepsilon_{\mathrm{train},t}\) regime (typical post-FPLD-training), the second summand dominates and recovers the Pinsker \(\sqrt{\cdot}\) shape. Theorem 6 bounds \(\varepsilon_{\mathrm{train},t}\) via the FPLD rate at the most recent training event. Under adversarial models violating the conditional-density clause, only the weaker rate \(\Delta_{\mathrm{train},t} = O(f_{\max,t}\,\varepsilon_{\mathrm{train},t}^{1/4})\) is recoverable via Markov truncation; we adopt the (B5’) regime as the operating assumption.
The four slack objects entering the per-step coverage bound are summarized below.
The base-paper FC-RAG inequality, lifted to a one-step conditional bound on the per-query miscoverage, is \[\label{eq:deployment-null} \mathbb{E}[M_t \mid \mathcal{F}_{t-1}] \;\mathbf{1}_{G_t} \;\le\; b_t \,\mathbf{1}_{G_t}, \qquad b_t \;:=\; \alpha \;+\; \frac{1}{n_{\mathrm{cal},t}+1} \;+\; \Delta_{\mathrm{FL},t} \;+\; \Delta_{\mathrm{RAG},t} \;+\; \Delta_{\mathrm{train},t}.\tag{3}\] We call 3 the deployment null \(H_0\). Each \(\Delta_{\bullet,t}\) is \(\mathcal{F}_{t-1}\)-measurable by predictability, hence so is \(b_t\). The bound holds pointwise on the calibration-good event \(G_t\); off \(G_t\) the bound is vacuous and the residual probability is absorbed into the final \(\delta_{\mathrm{cal}}\) budget. The theorems in Section 4 test whether the realized stream is consistent with \(H_0\) on \(\Omega_{\mathrm{cal}}\). The classical \(1/(n_{\mathrm{cal},t}+1)\) split-conformal overshoot is included for compatibility with the marginal-coverage bound of [1] and is dominated by \(\Delta_{\mathrm{FL},t}\) for typical \(n_{\mathrm{cal},t}\).
Adversarial nodes, differential privacy, open-ended generation, arbitrarily delayed labels, and architecture-heterogeneous clients are excluded from the first version of the theory. Each is a natural extension, but none is necessary to formulate the sequential object above.
The protocol is synchronous, single-aggregator, and charges uplink communication only. Figure 1 shows the two coupled loops: a fast per-query inference loop in which \(K\) nodes serve a query under their bandwidth budgets, and a slow sequential-testing loop in which observed miscoverage drives a betting e-process and a predictable controller feeds back to budgets, threshold, and student.
Once the answer \(Y_t\) is observed, the hub computes \(M_t = \mathbf{1}\{Y_t \notin C_t(X_t)\}\) and updates the betting e-process \[\label{eq:eprocess} E_0 \;=\; 1, \qquad E_t \;=\; E_{t-1}\,\big(1 + \lambda_t \, Z_t\big), \qquad Z_t \;:=\; M_t - b_t,\tag{4}\] where \(\lambda_t \in [0, 1/b_t]\) is predictable. The bound \(\lambda_t \le 1/b_t\) ensures \(1+\lambda_t Z_t \ge 0\) on every realization, since \(Z_t \ge -b_t\). The alarm time is \[\tau_{\mathrm{alarm}} \;=\; \inf\{t \ge 1 : E_t \ge 1/\delta\}.\] Note the algorithm does not reset \(E_t\) after an alarm or refresh: \(E_t\) continues to accumulate evidence across interventions, which is what preserves the time-uniform Ville bound across the entire post-deployment trajectory. (A reset variant with budgeted \(\delta\) across epochs is also valid; we discuss this in Section 4.5.0.1.) The ordering inside Algorithm 2 is essential: \(\widehat q_t\) is set from \(\mathcal{F}_{t-1}\) before \(X_t\) is served, \(M_t\) is recorded, \(E_t\) is updated using \(\mathcal{F}_{t-1}\)-measurable \(b_t\) and \(\lambda_t\), and only then does the controller act, executing Algorithm 3 to recompute the threshold and choose the next round’s budgets. Each downstream object is therefore predictable with respect to the next round.
The controller may react to the monitoring state in three ways: (i) recalibration (request fresh compressed calibration summaries and update \(\widehat q_t\)); (ii) bandwidth escalation (increase selected \(B_{i,t}\) or the calibration-refresh budget \(B_t^{\mathrm{cal}}\)); (iii) retraining refresh (replace the current student with a refreshed model). All such actions must be \(\mathcal{F}_{t-1}\)-measurable. That predictability condition is what lets the same e-process both guide interventions and remain valid after them (Theorem 5).
Fixed-horizon FC-RAG [1] sits inside our framework as the degenerate case in which \(B_{i,t} \equiv B_i\), the threshold \(\widehat q_t\) is computed once from a one-shot calibration set, the student is frozen throughout deployment, no intervention is triggered, and only the terminal-time miscoverage matters. The bounds differ in conditioning structure: FC-RAG’s marginal-coverage Theorem 2 holds with high probability over the calibration draw, whereas the present paper’s per-step bound holds conditional on \(\mathcal{F}_{t-1}\) on the cal-good event \(G_t\) (with the residual probability collected into \(\delta_{\mathrm{cal}}\) via union bound). The new ingredients of Anytime-FC-RAG are exactly three: a rolling calibration state, a (truncated) e-process for time-uniform monitoring, and predictable control actions that change bandwidth or refresh the model only when justified by accumulated evidence.
Off-the-shelf sequential testing supplies Ville’s inequality and the Hoeffding-stitched envelope, and base FC-RAG [1] supplies the slack decomposition \(\Delta_{\mathrm{FL},t}, \Delta_{\mathrm{RAG},t}, \Delta_{\mathrm{train},t}\) in marginal-over-calibration form. Neither composes into a sound conditional supermartingale on its own, because base FC-RAG’s coverage bound is a marginal claim (it fails conditionally on adverse calibration realizations) and the betting e-process needs a strict conditional centering. Our load-bearing construction is the cal-deviation budget \(\{\delta_t^{\mathrm{cal}}\}\) together with the resulting calibration-good event \(G_t\) and the truncated supermartingale \(\widetilde{E}_t = E_t\,\mathbf{1}_{\bigcap_{s\le t} G_s}\): this is what converts the marginal slack form into a strict conditional bound (Lemma 1) and turns the obstructed e-process into a bona fide supermartingale on the entire probability space (Lemma 2). Once those two pieces are in place, the rest is reuse: Ville’s inequality (Theorem 1), Hoeffding-stitching (Theorem 3), predictability of the controller (Theorem 5), and the FTC + (B5’) chain of [1] applied to the FPLD KL rate (Theorem 6).
The analysis is organized in six results. Lemma 1 lifts the base-paper Theorem 2 to a one-step conditional bound on \(G_t\). Lemma 2 establishes that \(\widetilde{E}_t\) is a nonnegative supermartingale. Theorem 1 applies Ville’s inequality and a union bound on \(G_t\) to recover a time-uniform alarm guarantee. Theorem 3 converts the same construction into a time-uniform envelope on cumulative miscoverage. Theorem 5 shows that predictable interventions preserve both guarantees. Theorem 6 bounds \(\Delta_{\mathrm{train},t}\) using the FPLD training rate at the most recent refresh.
Assumption 1 (Predictable schedules). For every \(t\), the calibration index set \(I_t^{\mathrm{cal}}\), the budgets \(\{B_{i,t}\}_{i=1}^K\) and \(B_t^{\mathrm{cal}}\), the active student \(\widehat P_t\), the threshold \(\widehat q_t\), the slack quantities \(b_t\), and the betting fraction \(\lambda_t\) are all \(\mathcal{F}_{t-1}\)-measurable.
Assumption 2 (I.i.d.data and predictable buffer). The deployment sequence \((X_t, Y_t)_{t \ge 1}\) is i.i.d.from a fixed joint law \(\mathcal{P}\) on \(\mathcal{X}\times \mathcal{Y}\), and for every \(t\) the calibration index set satisfies \(I_t^{\mathrm{cal}} \subseteq \{1,\dots,t-1\}\) and is \(\mathcal{F}_{t-1}\)-measurable.
This is the analogue of Assumption (B1) in [1], adapted to the sequential setting. It is the cleanest sufficient condition for the per-step rank exchangeability between \((X_t, Y_t)\) and the buffer points to drive the conditional coverage analysis below; the buffer, being \(\mathcal{F}_{t-1}\)-measurable, is held fixed when conditioning, and randomness reduces to the test point and the buffer realization (the latter shared between \(\mathcal{F}_{t-1}\) and the implicit \(G_t\) event over the buffer draw). The assumption fails under genuine drift, and rejection of the deployment null is exactly the alarm event the e-process detects.
Assumption 3 (Bounded score and strengthened density regularity). The score \(s\) is bounded in \([0,S_{\max}]\) (clipping). The cumulative distribution function \(F_t\) of the oracle uncompressed score \(s_t^\star(X_t,Y_t)\) admits a density \(f_t \le f_{\max,t}\) on a fixed deterministic* neighborhood \([q_t^{\mathrm{pop}} - r_t,\, q_t^{\mathrm{pop}} + r_t]\) for some \(r_t > 0\); the analysis verifies a posteriori via 1 that \(\widehat q_t\) lies inside this neighborhood with probability at least \(1 - \delta_t^{\mathrm{cal}}\). Inheriting (B5’) of [1], the conditional density (on the same neighborhood) of the relevant “before-perturbation” score given the perturbation is also bounded by \(f_{\max,t}\) in two specific cases:*
**Retrieval-bandwidth dither* (Step 3 of Lemma 1): the conditional density of \(s_t^\star(X_t,Y_t)\) given the dither average \(\bar\xi_t(X_t,Y_t) := (1/K)\sum_i \xi_{i,t}(X_t,Y_t)\), where \(\bar\xi_t\) has \(\mathbb{E}[\bar\xi_t\mid\mathcal{F}_{t-1},X_t]=0\) and bounded conditional second moment by Assumption (B3) of [1].*
**Training residual* (FTC chain of Section 2 and Theorem 6): the conditional density of \(s^\star_{\mathrm{ideal}}(X_t,Y_t) := -\log P^\star(Y_t\mid X_t)\) given \(\Delta_t(X_t,Y_t) := \log(P^\star(Y_t\mid X_t)/\widehat P_t(Y_t\mid X_t)) = s_t^\star - s^\star_{\mathrm{ideal}}\).*
Assumption 4 (Slack admissibility). For every \(t\), \(b_t \in (0, 1-\eta]\) for some fixed \(\eta > 0\), and the betting cap \(\bar\lambda\) satisfies \(\bar\lambda \le 1/b_t\) uniformly in \(t\).
Assumption 4 ensures the e-process stays nonnegative; we typically take \(\bar\lambda \le 1/(\alpha + \Delta_{\max})\) for a slack upper bound \(\Delta_{\max}\) chosen by the operator.
Lemma 1 (One-step deployment null on the cal-good event). Under Assumptions 1–3, on the calibration-good event \(G_t\) defined in Section 2 (which satisfies \(\mathbb{P}(G_t) \ge 1 - \delta_t^{\mathrm{cal}}\)), \[\mathbb{E}[M_t \mid \mathcal{F}_{t-1}] \;\mathbf{1}_{G_t} \;\le\; b_t \,\mathbf{1}_{G_t}.\] Equivalently, on the event \(G_t\), \(\mathbb{E}[M_t \mid \mathcal{F}_{t-1}] \le b_t\).
Proof. The proof is provided in Appendix 7.1. ◻
This is the conditional bound the betting e-process needs as a supermartingale on top of FC-RAG. The pointwise form on \(G_t\) (not FC-RAG’s marginal Theorem 2 form) is the centering Ville’s inequality requires; without it, no anytime-valid alarm is possible. The price is a \(\sqrt{\log(2/\delta_t^{\mathrm{cal}})}\) inflation of \(\Delta_{\mathrm{FL},t}\), an \(O(\sqrt{\log t})\) factor amounting to \(1.07{\times}\)–\(2.72{\times}\) across \(t \in [1, 10^5]\). The cal-deviation budget is not optional: without it, the marginal centering fails the supermartingale property on adverse calibration draws and Ville’s inequality cannot be applied (Appendix 7).
Define the per-step centered residual \(Z_t := M_t - b_t \in [-b_t,\,1-b_t]\), an \(\mathcal{F}_t\)-measurable bounded random variable, and the cumulative cal-good event \(\Omega_{\mathrm{cal},t} := \bigcap_{s \le t} G_s\) (\(\mathcal{F}_{t-1}\)-measurable since each \(G_s\) is \(\mathcal{F}_{s-1}\)-measurable).
Lemma 2 (Truncated e-process is a supermartingale). Under Lemma 1 and Assumptions 1, 4, define the truncated e-process \[\widetilde{E}_t \;:=\; E_t \cdot \mathbf{1}_{\Omega_{\mathrm{cal},t}}, \qquad \widetilde{E}_0 := 1,\] where \(E_t\) is the betting e-process 4 with predictable \(\lambda_t \in [0,\,1/b_t]\). Then \((\widetilde{E}_t)_{t \ge 0}\) is a nonnegative supermartingale with \(\mathbb{E}\widetilde{E}_0 = 1\): \[\mathbb{E}[\widetilde{E}_t \mid \mathcal{F}_{t-1}] \;\le\; \widetilde{E}_{t-1} \qquad \text{for every } t \ge 1.\]
Proof. The proof is provided in Appendix 7.2. ◻
Theorem 1 (Time-uniform alarm validity). Under the assumptions of Lemma 2, for every \(\delta_e \in (0,1)\), \[\mathbb{P}\!\left(\sup_{t \ge 1} E_t \ge 1/\delta_e\right) \;\le\; \delta_e + \delta_{\mathrm{cal}}.\] In particular, with the canonical split \(\delta_e = \delta_{\mathrm{cal}} = \delta/2\), the alarm time \(\tau_{\mathrm{alarm}} = \inf\{t : E_t \ge 2/\delta\}\) satisfies \(\mathbb{P}(\tau_{\mathrm{alarm}} < \infty \mid H_0) \le \delta\).
Proof. The proof is provided in Appendix 7.3. ◻
Remark 2 (Total probability budget). Theorem 1’s budget \(\delta_e + \delta_{\mathrm{cal}}\) holds on the training-good event of Theorem 6, which has probability \(\ge 1 - \delta_{\mathrm{train}}\) over the training draws. Combining by union bound, the unconditional alarm guarantee is \(\mathbb{P}(\sup_t E_t \ge 1/\delta_e) \le \delta_e + \delta_{\mathrm{cal}} + \delta_{\mathrm{train}}\), matching the user-facing total stated in Section 1. The canonical equal split \(\delta_e = \delta_{\mathrm{cal}} = \delta_{\mathrm{train}} = \delta/3\) recovers a single \(\delta\)-level guarantee.
The alarm at \(\sup_t E_t \ge 1/\delta\) controls the probability of ever flagging a violation. To get a numerically interpretable coverage statement at any predictable stopping time \(\tau\) we use the cumulative residual.
Theorem 3 (Time-uniform Hoeffding envelope). Under Lemma 1 and Assumption 1, define \(S_t := \sum_{s=1}^t (M_s - b_s) = \sum_{s=1}^t Z_s\). Then \(|Z_s| \le 1\) deterministically, and for every \(\delta_e \in (0,1)\) there is an explicit boundary \(u_t(\delta_e)\) with \[\mathbb{P}\!\left(\exists\, t \ge 1: S_t > u_t(\delta_e)\right) \;\le\; \delta_e + \delta_{\mathrm{cal}}.\] A closed-form admissible choice (polynomial-stitching variant of [7]) is \[u_t(\delta_e) \;=\; c_H\,\sqrt{\tfrac{1}{2}\,t \,\Big(\log(1/\delta_e) + \log\big(1 + \log_2 t\big)\Big)}\] with absolute constant \(c_H \le 1.7\). In particular, with probability at least \(1-\delta_e-\delta_{\mathrm{cal}}\), for every stopping time \(\tau\) adapted to \((\mathcal{F}_t)\), \[\frac{1}{\tau} \sum_{s=1}^{\tau} M_s \;\le\; \alpha \;+\; \frac{1}{\tau}\sum_{s=1}^{\tau}\!\left( \tfrac{1}{n_{\mathrm{cal},s}+1} + \Delta_{\mathrm{FL},s} + \Delta_{\mathrm{RAG},s} + \Delta_{\mathrm{train},s} \right) \;+\; \frac{u_\tau(\delta_e)}{\tau}.\] The envelope width \(u_\tau(\delta_e)/\tau = O(\sqrt{\log\log \tau / \tau})\) vanishes with \(\tau\).
Proof. The proof is provided in Appendix 7.4. ◻
Remark 4. The envelope and the betting e-process are complementary: \(E_t\) accumulates evidence multiplicatively and is best when violations are persistent (high power for sustained drift), while \(u_t(\delta)\) controls the cumulative deviation at any finite horizon and is best for reporting an interpretable upper bound on the realized miscoverage rate. We track both and report whichever is tighter at the requested \(\delta\).
Theorem 5 (Safe adaptive control). Let \(\Pi\) be any controller that maps \(\mathcal{F}_{t-1}\) to a (possibly randomized) action in \(\mathcal{A}_t\), where \(\mathcal{A}_t\) ranges over: recalibration refreshes (changing \(\widehat q_t\)), per-node bandwidth changes (changing \(B_{i,t}\) or \(B_t^{\mathrm{cal}}\)), and student refreshes (changing \(\widehat P_t\)). If every action is \(\mathcal{F}_{t-1}\)-measurable, then under Assumptions 1–4 the supermartingale property of the truncated e-process \((\widetilde{E}_t)\) and the time-uniform Hoeffding envelope of Theorem 3 both continue to hold. Hence the alarm bound \(\mathbb{P}(\sup_t E_t \ge 1/\delta_e) \le \delta_e + \delta_{\mathrm{cal}}\) of Theorem 1 and the envelope bound of Theorem 3 apply to the controlled trajectory.
Proof. The proof is provided in Appendix 7.5. ◻
Algorithm 2 does not reset \(E_t\) after an alarm (sticky mode). A reset variant with per-epoch budgets is also valid; both variants preserve Theorem 1 (Appendix 7).
Theorem 6 (Training propagation under (B5’)). Fix a summable per-training-event budget \(\{\delta_r\}_{r \ge 1}\) with \(\sum_{r \ge 1} \delta_r \le \delta_{\mathrm{train}}\) (canonical choice \(\delta_r = 6\delta_{\mathrm{train}}/(\pi^2 r^2)\)). Suppose the student \(\widehat P_t\) used at deployment time \(t\) comes from FPLD training event \(r(t)\) [1] with parameters \((K, n_{r(t)}, m_{r(t)}, B_{r(t)}, V)\), and let \[\mathcal{R}_{r(t)} \;:=\; \frac{c_1 d}{K\,n_{r(t)}} + c_2\,\rho\,\frac{V \log(V/\delta_{r(t)})}{\sqrt{m_{r(t)}}} + c_3\,2^{-2 B_{r(t)}/V} + \varepsilon_{\mathrm{opt}} + \varepsilon_{\mathrm{fit}}\] denote the corresponding training-rate bound. Under the strengthened conditional-density clause of Assumption 3, with probability at least \(1 - \delta_{\mathrm{train}}\) over the training draws, simultaneously over all \(t \ge 1\), \[\Delta_{\mathrm{train},t} \;\le\; f_{\max,t}\,\big(\mathcal{R}_{r(t)} \;+\; \sqrt{2\,\mathcal{R}_{r(t)}}\big).\]
Proof. The proof is provided in Appendix 7.6. ◻
In the small-\(\mathcal{R}\) regime the \(\sqrt{2\,\mathcal{R}_{r(t)}}\) summand dominates, recovering the Pinsker shape. Absent (B5’), only the weaker \(O(\mathcal{R}^{1/4})\) rate is recoverable; the sequential-testing layer is unaffected either way (Appendix 7).
A communication-optimality oracle inequality of the form \(\mathbb{E}\sum_{t=1}^T \Gamma_t^{\Pi} \le \inf_{\Pi'\in\mathfrak{C}_{\mathrm{valid}}}\mathbb{E}\sum \Gamma_t^{\Pi'} + \mathrm{overhead}(T)\), where \(\mathfrak{C}_{\mathrm{valid}}\) is the class of validity-preserving controllers, would be substantially stronger than Theorem 5. We do not attempt it here and flag it as future work.
The experiments split into two qualitatively distinct roles. Synthetic experiments probe the sequential-testing layer in isolation on Bernoulli streams whose conditional miscoverage is set by hand, decoupled from the FC-RAG / FPLD pipeline. Real-world experiments deploy the GPT-2-small + MiniLM swarm of [1] on three benchmarks — MMLU [17] (in-cal-set redistribution), DBpedia ontology [18] (drift to a class the swarm has no node for), and AG News [19] (drift to a target recoverable from GPT-2 priors) — and compare against conformal test martingales [4], CUSUM, Shiryaev–Roberts, and online conformal prediction [12]. Real-world runs use \(3\) calibration splits \(\times\) \(5\) deployment seeds (\(15\) trajectories per benchmark); synthetic runs use \(200\)–\(2000\) seeds. All runs use \(\alpha = 0.10\), \(\delta_e = 0.05\), \(\delta_{\mathrm{cal}} = 0.05\); per-experiment details and expanded ablations are in Appendix 8.
On synthetic Bernoulli streams, Type-I rate is \(0.0105\) at the boundary of \(H_0\) and \(0.0025\) in the interior, both an order of magnitude below the \(\delta_e + \delta_{\mathrm{cal}} = 0.10\) budget (Table 2). Detection power saturates at \(\ge 99.6\%\) from drift \(+0.04\) onward, with median delay \(3047 \to 687\) steps as drift grows from \(+0.04\) to \(+0.15\) (Figure 4, Table 3), matching the predicted \(\Omega(\log\log T / \mathrm{drift}^2)\) rate. The Hoeffding-stitched envelope holds across all \(4000\) null trajectories with \(0.0000\) breach rate, and a Monte-Carlo sanity check confirms Lemma 1 pointwise (Appendix 8.1).
| Regime | alarm rate | median \(\sup_t E_t\) | \(p_{95}\sup_t E_t\) | \(p_{99}\sup_t E_t\) |
|---|---|---|---|---|
| Boundary (\(\mathbb{E}[Z_t\mid\mathcal{F}_{t-1}] = 0\)) | 0.0105 | 1.131 | 6.381 | 21.230 |
| Interior (\(\mathbb{E}[Z_t\mid\mathcal{F}_{t-1}] < 0\)) | 0.0025 | 1.000 | 3.063 | 6.622 |
| Drift size | fraction detected | median delay | \(p_{95}\) delay |
|---|---|---|---|
| \(+0.02\) | 0.310 | 5076 | 5938 |
| \(+0.04\) | 0.996 | 3047 | 4464 |
| \(+0.06\) | 0.998 | 1904 | 2685 |
| \(+0.08\) | 0.998 | 1361 | 1864 |
| \(+0.10\) | 0.998 | 1057 | 1480 |
| \(+0.15\) | 0.998 | 687 | 931 |
Empirical \(\Delta_{\mathrm{RAG}}\) tracks the variance form on a log-log axis with slope exactly \(-0.5000\) for \(K \in \{1, \dots, 128\}\) (Figure 5, left). The training slack’s two-term form holds tightly: empirical-to-theory ratio \(\le 0.91\) (mean \(0.65\)) across \(20\) Bernoulli pairs with \(\mathrm{Beta}(2,2)\) residuals. The cost-overhead factor \(R(t)\) grows from \(1.07\) at \(t=1\) to only \(2.72\) at \(t = 10^5\) (Figure 5, right), so the conditional construction is essentially free at any practical horizon.
Additional necessity checks (cal-deviation budget, predictability of the controller layer) and robustness studies (aGRAPA vs.constant-\(\lambda\) bettors, hyperparameter sensitivity) are reported in Appendix 8.4.
On a synthetic stream with \(+0.20\) drift at \(t=2500\), all three bandwidth regimes (low-only, high-only, adaptive) reach \(100\%\) alarm rate, but the adaptive regime pays only \(1.708\) on average — a \(57\%\) cost saving (Table 4). On the GPT-2-small swarm (\(B_i = 8\) low, \(B_i = 12\) high), the same pattern holds: on DBpedia all three regimes alarm at \(100\%\) and adaptive saves \(14\%\) (\(41.5\) vs.\(48.0\)); on MMLU and AG News (no genuine drift) the controller correctly does not escalate (Figure 6). Predictable escalation is a real operational lever: alarm validity is preserved, and high bandwidth is paid only where needed.
| Regime | mean cost | alarm rate |
|---|---|---|
| Low only (\(\Delta_{\mathrm{RAG}} = 0.04\)) | 1.000 | 1.000 |
| High only (\(\Delta_{\mathrm{RAG}} = 0.005\)) | 4.000 | 1.000 |
| Adaptive (\(0.5/\delta_e\) warning trigger) | 1.708 | 1.000 |
The end-to-end result (Figure 7, Table 5) on the GPT-2-small + MiniLM swarm with sudden drift at \(t = 500\) over \(T = 2000\) (\(15\) trajectories): the e-process is silent on MMLU in-cal redistribution (alarm \(0.00\), post-miscov \(0.149\)) and AG News priors-recoverable drift (\(0.00\), \(0.127\)), and fires on DBpedia genuine drift (\(0.33\), rising to \(1.00\) in the head-to-head; post-miscov \(0.370 > b \approx 0.21\)). Three robustness ablations (drift schedule, heterogeneous bandwidth, FPLD multi-refresh) preserve this discriminative behavior (Appendix 8.3).
Against prior monitoring methods, our e-process matches conformal test martingales [4] in null validity (\(0.009\) vs.\(0.000\)) and dominates on drift (\(1.000\) vs.\(0.000\) on large drift), because CTM does not exploit the slack-decomposed null bound. Parametric CUSUM and Shiryaev–Roberts match or exceed at every drift by assuming the alternative is known; our nonparametric e-process is slower only at borderline drift, the expected price of distribution-free testing [7]. Online conformal [12] adapts \(\alpha_t\) to maintain coverage rather than alarming, so the two are complementary.
The real-world head-to-head makes this concrete (Figure 8): on DBpedia our alarm fires in every trajectory while OC compresses \(\alpha_t\) from \(0.10\) to \(0.031\); on MMLU and AG News both heads correctly stay quiet. Aggregated (Table 5), the alarm fires when and only when coverage genuinely breaks.
| Experiment | MMLU | DBpedia | AG News |
|---|---|---|---|
| End-to-end | 0.00 / 0.149 | 0.33 / 0.370 | 0.00 / 0.127 |
| Sudden-drift ablation | 0.33 | 1.00 | 0.00 |
| Adaptive controller | 0.33 / 34.4 | 1.00 / 41.5 | 0.00 / 24.0 |
| Online-conformal head-to-head | 0.27 / 0.184 | 1.00 / 0.445 | 0.00 / 0.168 |
Anytime-valid testing via e-processes goes back to [20] and was modernized for nonparametric settings in [6], [7]; the betting interpretation we use [5], [8] and recent conformal extensions [9]–[12], [21] cover the sequential-testing layer. The alarm half of our construction is a federated, bandwidth-aware analogue of conformal test martingales [4]; classical CUSUM and Shiryaev–Roberts changepoint detection are recovered as the special case \(b_t = \alpha\) of our construction. Single-site conformal-RAG [2], [3] and federated conformal prediction [13]–[16] are proper subsets of our setting in distinct dimensions: the former assumes one model, one corpus, one calibration set; the latter takes the score as given and is silent on per-node retrieval bandwidth. Closest in spirit to our deployment null is the non-exchangeable coverage analysis of [22]. The construction is not RAG-specific: any sequential conformal protocol with i.i.d.deployment data, an \(\mathcal{F}_{t-1}\)-measurable threshold satisfying a high-probability quantile bound, and a predictable slack decomposition admits the same alarm and envelope guarantees at budget \(\delta_e + \delta_{\mathrm{cal}}\).
We extended FC-RAG from a fixed-horizon coverage guarantee to an anytime-valid deployment-time reliability framework via the cal-deviation budget and the truncated supermartingale; the four theorems cover alarm validity, cumulative-miscoverage envelope, safe adaptive control, and training-to-deployment propagation, and the empirical Type-I rate, detection power, envelope coverage, controller-cost saving, and discriminative real-LM behavior all match the predicted regimes (Section 5).
The main formulation assumes immediate label feedback, and switching to an empirical-Bernstein boundary [7] would tighten the envelope by typically \(1.5{\times}\). Under adversarial models that violate the (B5’) clause, only the weaker \(\Delta_{\mathrm{train},t} = O(f_{\max,t}\,\mathcal{R}_{r(t)}^{1/4})\) rate is recoverable; the sequential-testing layer is unaffected. The construction is not specific to RAG and can be attached to any sequential conformal protocol with a predictable slack decomposition. Two operational caveats: the conformal coverage guarantee holds in expectation across queries, not conditionally per input; and the protocol does not provide differential privacy on its own. Adversarial nodes, open-ended generation, architecture-heterogeneous clients, and a communication-optimality oracle inequality are out of scope.
Supplementary Material
This supplement contains complete proofs of all six results stated in the main paper (Lemmas 1–2 and Theorems 1–6), with extended remarks on the cal-deviation budget, sticky vs.resetting alarms, and the tightness of the conditional-density clause placed after the corresponding proofs; and extended experimental results including the Lemma 1 sanity check, the \(\Delta_{\mathrm{train}}\) two-term form verification, real-world ablations (A2–A4), and necessity and robustness studies. We use the notation of the main paper throughout.
The proof proceeds in three steps. We first bound the conditional miscoverage relative to the population quantile (Step 1), then absorb the empirical-vs-population deviation on the cal-good event \(G_t\) (Step 2), then absorb retrieval-bandwidth quantization (Step 3). Throughout, the test pair \((X_t,Y_t)\) is independent of \(\mathcal{F}_{t-1}\) by Assumption 2, and any quantity defined from past observations is \(\mathcal{F}_{t-1}\)-measurable.
Step 1 (population coverage). Recall the deployed-student score \(s_t^\star\) from Section 2, and let \(q_t^{\mathrm{pop}}\) denote the \((1-\alpha)\)-quantile of \(s_t^\star(X,Y)\) under \((X,Y) \sim P^\star\) (the score function is \(\mathcal{F}_{t-1}\)-measurable, so \(q_t^{\mathrm{pop}}\) is too). Since \((X_t,Y_t) \mid \mathcal{F}_{t-1}\) has the same conditional law as the deployment-law marginal, \[\mathbb{P}\!\left(s_t^\star(X_t,Y_t) \le q_t^{\mathrm{pop}} \,\Big|\, \mathcal{F}_{t-1}\right) \;\ge\; 1 - \alpha,\] with equality under continuity. (The classical \(1/(n_{\mathrm{cal},t}+1)\) split-conformal overshoot is a strictly weaker bound on the same quantity; we keep it inside \(b_t\) for compatibility with [1] but the high-probability conditional analysis below does not need it.)
Step 2 (quantile reconstruction on \(G_t\)). On the event \(G_t\), by the construction of \(G_t\) in 1 , \[\big|\widehat q_t - q_t^{\mathrm{pop}}\big| \;\le\; c_q\sqrt{\log(2/\delta_t^{\mathrm{cal}})/n_{\mathrm{cal},t}} \;+\; \phi(B_t^{\mathrm{cal}}).\] By Assumption 3 the score CDF is \(f_{\max,t}\)-Lipschitz in a neighborhood of \(q_t^{\mathrm{pop}}\) [1], so on \(G_t\), \[\big|\,\mathbb{P}(s_t^\star \le \widehat q_t \mid \mathcal{F}_{t-1}) - \mathbb{P}(s_t^\star \le q_t^{\mathrm{pop}} \mid \mathcal{F}_{t-1})\,\big| \;\le\; \Delta_{\mathrm{FL},t}.\]
Step 3 (score perturbation, variance-based). Under the dithered-quantization Assumption (B3) of [1] the swarm score decomposes as \(s_t^{\mathrm{swarm}}(X_t,y) = s_t^\star(X_t,y) + \bar\xi_t(X_t,y)\) with the dither average \(\bar\xi_t = (1/K)\sum_i \xi_{i,t}\) satisfying \(\mathbb{E}[\bar\xi_t \mid \mathcal{F}_{t-1}, X_t] = 0\) and, by independence of the per-node noise, \[\mathbb{E}[\bar\xi_t^2 \mid \mathcal{F}_{t-1}, X_t] \;\le\; V_{K,t} \;:=\; \frac{1}{K^2}\sum_{i=1}^K v(B_{i,t}).\] For any \(\mathcal{F}_{t-1}\)-measurable threshold \(u\) in the score-density neighborhood of Assumption 3, write \(F^\star_{|\bar\xi}(u) := F_{s_t^\star\mid\bar\xi_t}(u\mid \mathcal{F}_{t-1}, X_t)\). Then \[\mathbb{P}(s_t^{\mathrm{swarm}} \le u \mid \mathcal{F}_{t-1}, X_t) - \mathbb{P}(s_t^\star \le u \mid \mathcal{F}_{t-1}, X_t) \;=\; \mathbb{E}\!\big[F^\star_{|\bar\xi}(u - \bar\xi_t) - F^\star_{|\bar\xi}(u) \,\big|\, \mathcal{F}_{t-1}, X_t\big].\] The strengthened conditional-density clause of Assumption 3 (the conditional density of \(s_t^\star\) given \(\bar\xi_t\) is bounded by \(f_{\max,t}\) on the same neighborhood) makes \(u\mapsto F^\star_{|\bar\xi}(u)\) uniformly \(f_{\max,t}\)-Lipschitz, so by Cauchy–Schwarz, \[\big|\,\mathbb{P}(s_t^{\mathrm{swarm}} \le u \mid \mathcal{F}_{t-1}, X_t) - \mathbb{P}(s_t^\star \le u \mid \mathcal{F}_{t-1}, X_t)\,\big| \;\le\; f_{\max,t}\,\mathbb{E}[|\bar\xi_t| \mid \mathcal{F}_{t-1}, X_t] \;\le\; f_{\max,t}\sqrt{V_{K,t}}.\] Taking expectation over \(X_t \mid \mathcal{F}_{t-1}\) and recalling the definition of \(\Delta_{\mathrm{RAG},t}\), \[\big|\,\mathbb{P}(s_t^{\mathrm{swarm}}(X_t,Y_t) \le u \mid \mathcal{F}_{t-1}) - \mathbb{P}(s_t^\star(X_t,Y_t) \le u \mid \mathcal{F}_{t-1})\,\big| \;\le\; \Delta_{\mathrm{RAG},t}.\]
Combine. On \(G_t\), chaining Steps 1–3, \[\begin{align} \mathbb{P}(s_t^{\mathrm{swarm}}(X_t,Y_t) \le \widehat q_t \mid \mathcal{F}_{t-1}) &\ge \mathbb{P}(s_t^\star(X_t,Y_t) \le \widehat q_t \mid \mathcal{F}_{t-1}) - \Delta_{\mathrm{RAG},t} && \text{(Step 3)}\\ &\ge \mathbb{P}(s_t^\star(X_t,Y_t) \le q_t^{\mathrm{pop}} \mid \mathcal{F}_{t-1}) - \Delta_{\mathrm{RAG},t} - \Delta_{\mathrm{FL},t} && \text{(Step 2)}\\ &\ge 1 - \alpha - \Delta_{\mathrm{RAG},t} - \Delta_{\mathrm{FL},t} && \text{(Step 1)}\\ &\ge 1 - \alpha - \Delta_{\mathrm{RAG},t} - \Delta_{\mathrm{FL},t} - \Delta_{\mathrm{train},t} && (\Delta_{\mathrm{train},t} \ge 0). \end{align}\] Hence \(\mathbb{E}[M_t \mid \mathcal{F}_{t-1}] \le \alpha + \Delta_{\mathrm{FL},t} + \Delta_{\mathrm{RAG},t} + \Delta_{\mathrm{train},t} \le b_t\) on \(G_t\). The \(\Delta_{\mathrm{train},t}\) slack in \(b_t\) absorbs training-side conservatism: it is bounded explicitly via the FTC chain of Section 2 (training-side distortion paragraph), inheriting the (B5’) conditional-density clause of [1], and propagated across FPLD refresh events in Theorem 6.0◻
Suppose one tried to skip the cal-deviation budget and instead use the marginal-over-calibration high-probability bound of [1] with a single fixed \(\delta\), centering the e-process at \[b_t^{\mathrm{marg}} := \alpha + \tfrac{1}{n_{\mathrm{cal},t}+1} + f_{\max,t}\!\left(c_q\sqrt{\log(2/\delta)/n_{\mathrm{cal},t}} + \phi(B_t^{\mathrm{cal}})\right) + \Delta_{\mathrm{RAG},t} + \Delta_{\mathrm{train},t}.\] The marginal bound holds with probability \(\ge 1 - \delta\) at any single \(t\), but for sequential validity simultaneously over all \(t\) a union bound is needed and a fixed \(\delta\) does not deliver one. On buffer realizations whose quantile deviation exceeds the centering at sufficiently many \(t\), \(\mathbb{E}[M_t \mid \mathcal{F}_{t-1}]\) exceeds \(b_t^{\mathrm{marg}}\) pointwise and the supermartingale property fails; Ville’s inequality cannot be applied. The summable budget \(\{\delta_t^{\mathrm{cal}}\}\) remedies this at the cost of an \(O(\sqrt{\log t})\) inflation of \(\Delta_{\mathrm{FL},t}\).
Nonnegativity. \(Z_t \ge -b_t\), so \(1 + \lambda_t Z_t \ge 1 - \lambda_t b_t \ge 0\) when \(\lambda_t \le 1/b_t\); inductively \(E_t \ge 0\) and hence \(\widetilde{E}_t = E_t\,\mathbf{1}_{\Omega_{\mathrm{cal},t}} \ge 0\).
Conditional mean. Since \(E_{t-1}\), \(\lambda_t\), and \(\mathbf{1}_{\Omega_{\mathrm{cal},t-1}}\) are \(\mathcal{F}_{t-1}\)-measurable, and \(\mathbf{1}_{G_t}\) is \(\mathcal{F}_{t-1}\)-measurable by construction (Section 2), the indicator \(\mathbf{1}_{\Omega_{\mathrm{cal},t}} = \mathbf{1}_{\Omega_{\mathrm{cal},t-1}}\,\mathbf{1}_{G_t}\) is \(\mathcal{F}_{t-1}\)-measurable. Hence \[\begin{align} \mathbb{E}[\widetilde{E}_t \mid \mathcal{F}_{t-1}] &= \mathbb{E}\!\left[E_{t-1}\,(1+\lambda_t Z_t)\,\mathbf{1}_{\Omega_{\mathrm{cal},t-1}}\,\mathbf{1}_{G_t}\,\Big|\,\mathcal{F}_{t-1}\right]\\ &= E_{t-1}\,\mathbf{1}_{\Omega_{\mathrm{cal},t-1}}\,\mathbf{1}_{G_t}\,\big(1 + \lambda_t\,\mathbb{E}[Z_t \mid \mathcal{F}_{t-1}]\big). \end{align}\] On \(G_t\), \(\mathbb{E}[Z_t\mid\mathcal{F}_{t-1}] \le 0\) by Lemma 1, so \(\mathbf{1}_{G_t}\big(1 + \lambda_t\,\mathbb{E}[Z_t \mid \mathcal{F}_{t-1}]\big) \le \mathbf{1}_{G_t} \le 1\). Therefore \[\mathbb{E}[\widetilde{E}_t \mid \mathcal{F}_{t-1}] \;\le\; E_{t-1}\,\mathbf{1}_{\Omega_{\mathrm{cal},t-1}} \;=\; \widetilde{E}_{t-1}. \qed\]
\((\widetilde{E}_t)\) is a nonnegative supermartingale with \(\mathbb{E}\widetilde{E}_0 = 1\) by Lemma 2, so Ville’s inequality [6], [20] gives \(\mathbb{P}(\sup_t \widetilde{E}_t \ge 1/\delta_e) \le \delta_e\). On \(\Omega_{\mathrm{cal}}\), \(\widetilde{E}_t = E_t\) for all \(t\). Splitting on \(\Omega_{\mathrm{cal}}\), \[\mathbb{P}\!\left(\sup_t E_t \ge 1/\delta_e\right) \;\le\; \mathbb{P}\!\left(\sup_t \widetilde{E}_t \ge 1/\delta_e\right) + \mathbb{P}(\Omega_{\mathrm{cal}}^c) \;\le\; \delta_e + \delta_{\mathrm{cal}}. \qed\]
\(Z_s \in [-b_s, 1-b_s]\) has range \(1\), so by Hoeffding’s lemma applied conditionally on \(\mathcal{F}_{s-1}\), \[\mathbb{E}[\exp(\lambda Z_s) \mid \mathcal{F}_{s-1}] \;\le\; \exp\!\left(\lambda \mathbb{E}[Z_s\mid\mathcal{F}_{s-1}] + \tfrac{\lambda^2}{8}\right).\] By Lemma 1, on \(G_s\) we have \(\mathbb{E}[Z_s\mid\mathcal{F}_{s-1}] \le 0\), so \(\mathbf{1}_{G_s}\,\mathbb{E}[\exp(\lambda Z_s)\mid \mathcal{F}_{s-1}] \le \mathbf{1}_{G_s}\,\exp(\lambda^2/8)\). For each fixed \(\lambda \ge 0\) the truncated exponential process \[\widetilde{W}_t^\lambda \;:=\; \exp(\lambda S_t - \lambda^2 t/8)\,\mathbf{1}_{\Omega_{\mathrm{cal},t}}\] is therefore a nonnegative supermartingale with \(\mathbb{E}\widetilde{W}_0^\lambda = 1\) (the same indicator-truncation argument as Lemma 2). Ville’s inequality gives \(\mathbb{P}(\exists\, t: \widetilde{W}_t^\lambda \ge 1/\delta_\lambda) \le \delta_\lambda\), equivalently \(\mathbb{P}(\exists\, t: S_t \ge \lambda^{-1}\log(1/\delta_\lambda) + \lambda t/8 \text{ on } \Omega_{\mathrm{cal},t}) \le \delta_\lambda\). Stitching over a discrete grid \(\lambda_k = 2^{-k}\) (\(k = 0,1,2,\dots\)) with budgets \(\delta_k = \delta_e\,c_\zeta/(k+1)^2\) (\(c_\zeta = 6/\pi^2\)) by the union-bound argument of [7] produces the displayed boundary on the event \(\Omega_{\mathrm{cal}}\); the constant \(c_H \le 1.7\) is from their Eq. (14). Splitting on \(\Omega_{\mathrm{cal}}\) then yields \(\mathbb{P}(\exists t: S_t > u_t(\delta_e)) \le \delta_e + \delta_{\mathrm{cal}}\). The stopping-time corollary follows by evaluating at \(t = \tau\) inside the joint high-probability event.0◻
The post-action quantities \((\widehat q_t, B_{\bullet,t}, \widehat P_t)\) at time \(t\) are \(\mathcal{F}_{t-1}\)-measurable by hypothesis. Hence \(b_t = \alpha + 1/(n_{\mathrm{cal},t}+1) + \Delta_{\mathrm{FL},t} + \Delta_{\mathrm{RAG},t} + \Delta_{\mathrm{train},t}\) and the calibration-good event \(G_t\) both remain \(\mathcal{F}_{t-1}\)-measurable, with the same per-step bound \(\mathbb{P}(G_t) \ge 1 - \delta_t^{\mathrm{cal}}\) applying since \(\delta_t^{\mathrm{cal}}\) is fixed by the predictable budget schedule. The conclusion of Lemma 1 applies as written. The supermartingale calculation in Lemma 2 proceeds without modification, as does the truncated Hoeffding bound underlying Theorem 3. Validity is therefore preserved under the entire intervention path.0◻
Algorithm 2 does not reset \(E_t\) after an alarm. The sticky version preserves \(\mathbb{P}(\sup_t E_t \ge 1/\delta_e) \le \delta_e + \delta_{\mathrm{cal}}\) over the entire trajectory but does not give post-intervention re-tests. A reset variant divides \(\delta_e\) into per-epoch budgets \(\delta_{e,k}\) with \(\sum_k \delta_{e,k} \le \delta_e\) and starts a fresh e-process at each reset; Theorem 1 applies to each epoch independently and a union bound recovers the global guarantee. Both variants are safe; we default to sticky.
Theorem 1 of [1] gives, for each training event \(r\), \(\bar K^{(r)} := \mathbb{E}_{X\sim P^\star_X}[\mathrm{KL}(P^\star(\cdot|X)\,\|\,\widehat P^{(r)}(\cdot|X))] \le \mathcal{R}_r\) with probability \(\ge 1 - \delta_r\). The chain of the training-side distortion paragraph in Section 2 ((B5’) of Assumption 3 + indicator-difference + FTC + the splitting \(\mathbb{E}_{P^\star}|\Delta_t| \le \bar K_t + \sqrt{2\,\bar K_t}\) via \(\log(1+t) \le t\) and Pinsker; [1]) gives, on each training-good event with \(\bar K_t := \bar K^{(r(t))}\), \[\Delta_{\mathrm{train},t} \;\le\; f_{\max,t}\,\mathbb{E}_{X,Y}|\Delta_t(X,Y)| \;\le\; f_{\max,t}\big(\bar K_t + \sqrt{2\,\bar K_t}\big) \;\le\; f_{\max,t}\big(\mathcal{R}_{r(t)} + \sqrt{2\,\mathcal{R}_{r(t)}}\big).\] Union bound over training events with \(\sum_r \delta_r \le \delta_{\mathrm{train}}\) yields the simultaneous-in-\(t\) statement on an event of probability \(\ge 1 - \delta_{\mathrm{train}}\).0◻
The two-term form \(\Delta_{\mathrm{train},t} = f_{\max,t}(\mathcal{R}_{r(t)} + \sqrt{2\,\mathcal{R}_{r(t)}})\) follows [1] under the clause (B5’) of Assumption 3. In the small-\(\mathcal{R}\) regime the second summand dominates, recovering the Pinsker-\(\sqrt{\cdot}\) shape; the linear summand is the price paid for matching base FC-RAG’s exact form. Absent the clause (B5’), the FTC step fails and only the weaker \(O(f_{\max,t}\,\mathcal{R}^{1/4})\) rate is recoverable via Markov truncation. Both the alarm validity and the cumulative envelope are insensitive to this choice; only the absolute scale of \(b_t\) shifts.
We sample \(1000\) buffer realizations of size \(n_{\mathrm{cal}} = 100\), scores \(\mathrm{Uniform}[0,1]\), \(\alpha = 0.10\), \(\delta_{\mathrm{cal}} = 0.05\), \(f_{\max} = 1.0\). For each buffer we compute the empirical \((1-\alpha)\)-quantile \(\widehat q\), the cal-good bound \(b_t\) at \(t \in \{1, 10, 10^2, 10^3, 10^4\}\) (with \(\delta_t^{\mathrm{cal}} = 6\delta_{\mathrm{cal}}/(\pi^2 t^2)\)), and check whether the conditional miscoverage rate \(\mathbb{P}(s > \widehat q \mid \widehat q)\) exceeds \(b_t\). The violation rate is \(0.0000\) across all \(5000\) buffer/horizon pairs, confirming Lemma 1 pointwise. The bound widens with \(t\) as expected (\(b_1 = 0.280\), \(b_{10^4} = 0.471\), via the \(\sqrt{\log(2/\delta_t^{\mathrm{cal}})}\) inflation in \(\Delta_{\mathrm{FL},t}\)), while the realized conditional miscoverage stays at \(\alpha = 0.10\) regardless of \(t\). The horizon-dependent inflation is the price of the conditional (vs.marginal) form; E8 quantifies it as a \(1.07{\times}\)–\(2.72{\times}\) factor over the cost-relevant range.
For Bernoulli pairs \((p, q)\) on a \(5\times 5\) grid with FPLD-distillation residuals drawn from \(\mathrm{Beta}(2, 2)\), the empirical ratio of \(|p - q|\) to \(f_{\max}(\mathcal{R} + \sqrt{2\mathcal{R}})\) stays below \(1\) in \(100\%\) of \(20\) sampled pairs (max ratio \(0.91\), mean ratio \(0.65\)). The two-term form from [1] holds tightly under the clause (B5’): this is a numerical witness for Theorem 6.
On DBpedia (the genuine-drift benchmark), alarm rates rank by drift severity as expected: no_drift \(0.00\), sudden \(1.00\), gradual \(0.67\), periodic \(0.33\). Sudden onset gives the strongest signal; gradual interpolation slows evidence accumulation; periodic schedules return to cal subjects every period and
dilute the signal. On MMLU (in-cal) and AG News (priors-recoverable), all schedules including sudden stay near zero alarm, mirroring A1’s discriminative behavior.
On MMLU across four bandwidth configurations, alarm rate is monotone in average bandwidth: uniform_high (\(B_i = 10\), alarm \(0.27\)); mixed_one_weak (\(B_i = (3, 10, 10, 10)\), alarm \(0.00\)); mixed_two_weak (\(B_i = (3, 3, 10, 10)\), alarm \(0.00\));
uniform_low (\(B_i = 4\), alarm \(0.00\)). Higher bandwidth shrinks \(\Delta_{\mathrm{RAG}}\), tightens \(b_t\), and lifts detection power, exactly the dependence the slack decomposition predicts (Theorem 1).
On MMLU with two FPLD-style refresh events (perturbation noise schedule \(0.5 \to 0.25 \to 0\) with transitions at \(t \in \{500, 1500\}\)), \(b_t\) shrinks at each refresh as the training residual \(\mathcal{R}_{r(t)}\) drops, and the e-process growth rate recalibrates accordingly, the predicted behavior under Theorem 6. Final miscoverage is \(0.120 \pm 0.027\) across \(15\) trajectories.
Two construction choices in §3 are not optional. The cal-deviation budget \(\{\delta_t^{\mathrm{cal}}\}\) is needed because adversarial buffer realizations make a marginal-bound centering \(b_t^{\mathrm{marg}}\) fail the supermartingale property; on uniform-Bernoulli null streams both our construction and the marginal variant alarm at \(0.0000\) (the inflation only kicks in on the exponentially-decaying tail event the budget is designed to control), so the budget functions as a theoretical safety net rather than an empirical-power booster. Predictability is needed at the controller layer: a non-predictable controller that switches bandwidth based on \(M_t\) rather than \(\mathcal{F}_{t-1}\)-measurable evidence inflates Type-I from \(0.0070\) (predictable, within \(\delta_e\)) to \(0.1010\) (non-predictable, above \(\delta_e + \delta_{\mathrm{cal}} = 0.10\)), a \(14.4\times\) violation of Theorem 1 confirming the predictability assumption of Theorem 5 is empirically load-bearing.
The bettor and hyperparameters are robust. Against three constant-\(\lambda\) baselines (\(\lambda \in \{0.32, 1.61, \bar\lambda = 3.23\}\), the slack-admissibility cap), the predictable-plug-in aGRAPA bettor matches all three under the null and on large drift but dominates by \(4.9\times\) on small drift (\(0.348\) alarm rate vs.best constant-\(\lambda\)’s \(0.071\), Figure 9); aGRAPA is the only choice that is uniformly competitive across drift sizes. One-at-a-time sweeps over \(\alpha \in \{0.05, 0.10, 0.15, 0.20\}\), \(\delta_e \in \{0.01, 0.025, 0.05, 0.10\}\), and the \(\lambda\)-cap factor \(\in \{0.25, 0.5, 0.75, 1.0\}\) (12 configurations in total) leave Type-I within the corresponding \(\delta_e\) in every configuration and detection power \(\ge 0.81\) throughout (Figure 10), so the canonical choices used elsewhere in the paper are not load-bearing in their specific values.