July 19, 2026
Long-horizon rollout generation has become the dominant systems bottleneck in agentic reinforcement learning (RL). As agents interact with environments over many turns, trajectories rapidly grow to tens of thousands of tokens, making synchronous RL training increasingly constrained by rollout. We propose WAR, a workload-aware rollout system that substantially accelerates synchronous agentic RL by jointly optimizing decoding and scheduling.
WAR is built on a key observation: the optimal rollout optimization strategy depends on runtime load:
(1) Under low load, WAR enables model-free speculative decoding with SuffixDecoding, which reuses suffix patterns from previously completed trajectories as speculative drafts for future rollouts. Unlike model-based drafters, SuffixDecoding introduces no additional draft model and avoids GPU contention with rollout generation.
(2) Under high load, where saturated batched decoding leaves limited room for speculative speedup, WAR shifts the optimization focus to cache-aware scheduling. A global scheduler places requests across rollout replicas based on cache locality, trajectory progress and server load, reducing redundant KV-cache recomputation and mitigating load imbalance.
By combining decoding-level suffix reuse with system-level rollout scheduling, WAR delivers robust throughput improvements across workload regimes without changing the underlying RL algorithm. WAR improves long-context agentic rollout throughput by \(1.4\times\) under low load and up to \(1.6\times\) under high load. These results show that WAR removes a major rollout bottleneck in synchronous agentic RL and provides a practical path toward scalable long-context agent training.
Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of domains, including coding [1], mathematics [2] and agent [3]. Reinforcement learning (RL) has become an effective approach for eliciting and enhancing these capabilities [4]. A typical RL training step consists of three stages: rollout generation, reward computation, and model update. Among them, rollout generation is increasingly becoming the dominant systems bottleneck for agentic RL.
This bottleneck is particularly severe in long-horizon agentic tasks. Unlike single-turn generation, agentic rollouts require the model to interact with external environments over multiple turns. As tasks become more complex, agents need more rounds of interaction to gather observations, execute tools, inspect feedback and refine their responses. As a result, the context length of a trajectory can quickly grow to tens of thousands of tokens. Prior work has observed that rollout generation can dominate the end-to-end RL training time, contributing a large fraction of the total runtime [5]. The problem is further amplified by the long-tail nature of agentic rollouts: while most requests finish early, a small number of trajectories require substantially more computation due to longer reasoning chains or more environment interactions. In synchronous RL training, these stragglers force other workers to wait, leading to poor resource utilization and reduced rollout throughput.
One way to alleviate this issue is asynchronous rollout, which overlaps rollout generation with training [6]. While asynchronous execution can improve hardware utilization, it introduces new challenges. The most important one is policy staleness: trajectories may be generated using outdated policy parameters, making the collected data less consistent with the current model. Asynchronous rollout also requires more complex coordination among rollout workers, reward computation, and model updates. Therefore, improving the efficiency of synchronous rollout remains an important and practical direction.
Existing synchronous RL framework, such as veRL [7], typically relies on simple request scheduling policies, such as least-request scheduling combined with sticky sessions. Such policies work reasonably well under light workloads, where GPU resources are not fully saturated and request placement has limited impact. However, under heavy workloads, many concurrent long-context requests compete for limited GPU memory and computation. In this regime, simple load-based scheduling becomes insufficient: request placement must also account for KV-cache locality, trajectory progress, and server load. Otherwise, trajectories may be repeatedly scheduled to suboptimal replicas, causing redundant KV-cache recomputation and exacerbating load imbalance.
Rollout generation is essentially batched LLM inference, and can potentially benefit from inference acceleration techniques. Speculative decoding is a widely used lossless acceleration method, where a drafter proposes multiple tokens and the target model verifies them in parallel [8]. The drafter can be model-based [9] or model-free [10]. However, directly applying model-based speculative decoding to RL rollouts is challenging. A model-based drafter may compete with the target model for GPU computation and memory, reducing overall rollout throughput. In addition, RL training continuously updates the target model. If the drafter is kept fixed, the mismatch between the drafter and the updated target model may reduce the acceptance length, wasting GPU computation during verification. If the drafter is updated together with the target model, the RL training framework must maintain and synchronize an additional model, introducing extra complexity and GPU overhead. Speculative decoding is most effective when GPU resources are underutilized, as it increases parallelism by verifying multiple draft tokens in a single forward pass. Under heavy workloads, however, standard batched decoding already saturates GPU resources, and the additional overhead of multi-token verification can outweigh the reduction in decoding steps, leading to degraded inference performance.
These observations suggest that rollout optimization should be workload-aware. Under low load, normal decoding often underutilizes the GPU, leaving room for speculative decoding to improve hardware utilization and reduce decoding steps. Under high load, however, batched decoding already saturates GPU resources, leaving limited room for speculative speedup. Instead, system-level scheduling becomes more important under heavy workloads. As many long-context trajectories concurrently compete for limited GPU memory and computation, poor request placement can cause redundant KV-cache recomputation, uneven server utilization, and long-tail stragglers.
Motivated by this insight, we propose WAR, a workload-aware rollout system for synchronous agentic RL training. WAR jointly optimizes decoding and scheduling, and adapts its optimization strategy according to runtime load. Under low load, WAR enables model-free speculative decoding with SuffixDecoding [11]. SuffixDecoding reuses suffix patterns from previously generated trajectories as speculative drafts for future rollouts. Since it is model-free, it introduces no additional draft model and avoids GPU contention with rollout generation. This makes it particularly suitable for agentic rollouts, where recurring code patterns may appear both across different prompts and among multiple rollouts of the same prompt.
Under high load, WAR shifts the optimization focus to cache-aware scheduling. A global scheduler assigns requests to rollout replicas based on cache locality, trajectory progress, and server load. By favoring placements with higher KV-cache reuse and by prioritizing shorter trajectories when appropriate, the scheduler reduces redundant KV-cache recomputation, mitigates load imbalance, and preserves cache capacity for long-running trajectories.
To keep the design simple and practical, WAR uses runtime effective batch size as a lightweight workload indicator and determines the switching thresholds empirically for different workloads and hardware settings.
WAR is designed to be lightweight and compatible with existing synchronous RL training pipelines. Built on top of veRL, it introduces a simple workload-aware control path that combines decoding-level suffix reuse with system-level rollout scheduling, without changing the underlying RL algorithm.
We make the following contributions:
We identify workload as a first-class factor in rollout optimization. We show that the effectiveness of rollout optimization techniques is workload-dependent: an optimization that improves throughput under one workload regime may provide limited benefit, or even introduce overhead, under another.
We design WAR, a workload-aware rollout system for synchronous agentic RL. WAR combines model-free SuffixDecoding with cache-aware scheduling, enabling decoding-level reuse under low load and reducing KV-cache recomputation and load imbalance under high load.
We demonstrate substantial rollout throughput improvements. We implement WAR on top of veRL with approximately 7K lines of Python code and evaluate it using Qwen3-32B on production agent task datasets with 16 H100 GPUs. Compared with the veRL baseline, WAR improves long-context agentic rollout throughput by \(1.4\times\) under low load and up to \(1.6\times\) under high load.
Reinforcement learning (RL) has become an important post-training technique for improving the reasoning and problem-solving capabilities of large language models. A typical RL iteration consists of three major stages [12]: rollout, reward computation, and training. During rollout, the actor model generates trajectories for a batch of prompts; during reward computation, the generated trajectories are evaluated by rule-based checkers, sandbox execution, reward models, or LLM-as-a-Judge [13]; finally, the actor model is updated using the collected training signals.
Many RL training systems perform rollout synchronously to preserve on-policy semantics. In this setting, the rollout stage must finish before the training stage starts, and all trajectories are generated by the most recent actor model. This design improves training stability and makes debugging and reproducibility easier. However, the synchronization barrier also exposes rollout inefficiency: when rollout lengths are imbalanced, workers that finish short trajectories early must wait for the slowest trajectories to complete. Prior systems have observed that rollout often dominates RL training time and becomes the primary source of GPU underutilization in synchronous RL post-training [5], [14].
Asynchronous rollout systems attempt to reduce such idle time by overlapping rollout and training [6], [15]–[17]. However, asynchronous execution may introduce policy staleness, because trajectories generated by older model parameters can be used to update a newer policy. It can also introduce distributional skew, where faster-to-generate short samples disproportionately appear in earlier training batches. These effects may harm algorithmic fidelity and complicate debugging and reproducibility. Therefore, optimizing synchronous rollout remains an important problem for practical agentic RL training.
Agentic RL further amplifies the rollout bottleneck. Unlike single-turn generation, an agent interacts with an external environment over multiple turns. For example, in software engineering (SWE) tasks, an agent may inspect files, execute commands, observe tool outputs, and iteratively refine its solution. As the number of interaction turns increases, the context length of each trajectory grows rapidly, often reaching tens of thousands of tokens.
Long-context rollout introduces two major system challenges. First, the KV-cache footprint grows with the sequence length. A long-reasoning request may start with a small memory footprint but expand to tens of gigabytes as decoding progresses [5]. This dynamic memory growth can force batch-size shrinkage, trigger request preemption, and cause expensive re-prefill operations, reducing overall rollout throughput.
Second, rollout lengths are highly imbalanced. Long-tail response length is a key source of inefficiency in synchronous rollout. Since the next training stage cannot proceed until all trajectories in the current rollout step finish, a small number of long trajectories can leave many rollout workers idle. This creates GPU bubbles and reduces end-to-end training efficiency. RollPacker reports that the longest response in a rollout batch can be tens of times longer than the median response, causing prolonged idle periods for GPUs assigned to shorter responses [14].
These challenges suggest that rollout optimization must address both computation and memory efficiency. Simply increasing parallelism or adding more workers is insufficient: poor request placement can increase KV-cache recomputation, while long-tail trajectories can still leave many workers idle near the end of a rollout step.
Existing rollout optimizations are not uniformly effective across all runtime workloads. In particular, the bottleneck of rollout generation changes with the serving load.
Under low load, GPU resources are often underutilized. In this regime, decoding-level acceleration such as speculative decoding [18], [19] can improve throughput by increasing parallelism within each target-model forward pass. Speculative decoding proposes multiple draft tokens and verifies them with the target model in parallel while preserving the target output distribution. This property makes it attractive for RL rollout, where changing the trajectory distribution may affect training correctness.
However, under high load, standard batched decoding already saturates GPU resources, leaving limited room for speculative speedup. In this regime, speculative decoding still introduces additional overhead, such as draft construction, suffix-tree lookup, verification, and auxiliary attention-mask handling. When this overhead outweighs the benefit of accepted draft tokens, speculative decoding can degrade throughput rather than improve it.
Instead, system-level scheduling becomes more important under high load. Many long-context trajectories concurrently compete for limited GPU memory and computation. Inefficient request placement can reduce KV-cache locality, trigger redundant prefill computation, and worsen load imbalance across rollout replicas. Therefore, high-load rollout requires scheduling policies that account for KV-cache reuse, trajectory progress, and server load.
The above observations motivate a workload-aware view of rollout optimization. Rather than applying a single optimization uniformly, a rollout system should adapt its strategy to the current workload regime. Under low load, it should exploit underutilized GPU resources with decoding-level acceleration. Under high load, it should avoid unnecessary speculative overhead and instead focus on cache locality and load balancing.
WAR follows this principle, as shown in Figure 1. Under low load, it enables model-free speculative decoding with SuffixDecoding [11], which reuses suffix patterns from previously completed trajectories without introducing an additional draft model. Under high load, it shifts the optimization focus to cache-aware scheduling, reducing redundant KV-cache recomputation and mitigating load imbalance. By combining decoding-level suffix reuse with system-level scheduling, WAR targets high rollout throughput across different workload regimes while preserving synchronous RL training semantics.
This section presents WAR, a workload-aware rollout system for long-context agentic RL training. We start with an overview of the architecture and then describe the design of its main modules in detail.
WAR is centered around two key techniques: cache-aware scheduling and model-free speculative decoding, as shown in Figure 2.
In agentic RL, agents interact with environments over multiple turns, causing the context length of a trajectory to grow to the order of 10K–100K tokens. A rollout step is typically served by multiple rollout replicas, i.e., inference instances. If different turns of the same trajectory are frequently scheduled across different replicas, the system may repeatedly recompute KV caches, introducing non-negligible overhead. This overhead becomes especially significant for long-horizon multi-turn tasks such as SWE agent. Therefore, WAR incorporates cache locality into rollout scheduling by considering the cache hit rate of each trajectory when assigning requests to rollout replicas.
WAR further accelerates rollout generation with model-free speculative decoding. Speculative decoding is a widely used inference acceleration technique in which a verifier, usually the target model, verifies draft tokens generated by a drafter. The drafter can be either model-based or model-free. In RL rollouts, however, a model-based drafter may compete with the target model for GPU resources, reducing the overall rollout throughput. In contrast, SuffixDecoding is model-free and avoids such resource contention. SuffixDecoding stores the historical trajectories and matches the current token pattern against them to generate draft tokens. The workflow of GRPO-like algorithms is naturally well suited for SuffixDecoding. In these algorithms, each request is rolled out \(n\) times to generate multiple trajectories. Since these trajectories originate from the same prompt and interact with similar task environments, they often exhibit repeated local string patterns. Based on this observation, WAR integrates SuffixDecoding into the rollout pipeline to improve generation efficiency without changing the underlying RL algorithm.
At the begining of rollout, all rollout prompts are queued into a global scheduler. When the number of total request in the global request scheduler \(N_{\text{req,tot}}\) is less than a threshold \(Thr_{\text{sched}}\). The scheduling policy falls back to least-request combined with sticky-session, which is original scheduling policy of veRL. When \(N_{\text{req,tot}} \ge Thr_{\text{sched}}\), we score all the queued prompts to quantify which queued request should be scheduled. An illustrative example can be found in Figure 3. The priority is scored around several dimensions:
Cache hit rate \(H\). A higher \(H\) of a request means that current routed rollout instance has a higher kv cache in its GPU global memory. Thus \(H\) is a locality affinity metric and can be directly obtained from rollout server.
Online estimated assistant turns \(\hat{T}_g\). Beside cache hit rate, the actual length of context is also important. For two requests with same cache hit rate, their actual length of kv cache may be different. It is observed that inside a rollout group, in which all trajectories have same initial prompt, their assistant turns are similar. Assistant turn is usually proportional to context length. Compared to exact context length, the assistant turns are easier to estimate. Once a multi-turn trajectory finishes rollout, the corresponding assist turn can be updated as a function of group id: \[\hat{T}^{\,\text{new}}_g = \alpha T^{\text{obs}}_g + (1-\alpha)\hat{T}^{\,\text{old}}_g,\] where \(T_g^{\text{obs}}\) is the observed number of assistant turns for group \(g\) and \(\alpha\) is the smoothing factor.
Number of inflight requests \(N_{\text{rep}}^{\text{inflight}}\) on each rollout replica \(\text{rep}\). It is a metric of workload on \(\text{rep}\).
Waiting time of each request \(t_{\text{req}}^{\text{waiting}}\) in the queue of global scheduler. \(t_{\text{req}}^{\text{waiting}}\) is a practical concern because the sandbox related to the request have a prescribed time to live \(\text{ttl}_{\text{sandox}}\). If the sandbox is not visited for \(\text{ttl}_{\text{sandox}}\), it will be released. Therefore, a request can not stay to long in the queue.
Above 4 factors are combined together linearly to give a final score of a request: \[\begin{align} S &= S_\text{waiting} + S_\text{turn} + S_\text{cache} + S_\text{inflight} \\ &= \mathbf{1}\!\left(t_{\text{req}}^{\text{wait}} > \text{scale}_\text{ttl} \cdot \text{ttl}_{\text{sandox}}\right) \cdot \text{scale}_\text{waiting} \cdot t_{\text{req}}^{\text{wait}} \\ &\quad+ \text{scale}_\text{turn} \cdot \hat{T}_g + \text{scale}_\text{cache} \cdot H + \text{scale}_\text{inflight} \cdot N_{\text{rep}}^{\text{inflight}}. \end{align}\] where \[S_\text{waiting} > 0, \\ S_\text{turn} < 0, \\ S_\text{cache} > 0, \\ S_\text{inflight} < 0, \\ 0 < \text{scale}_\text{ttl} < 1.\] A negative \(\text{scale}_\text{turn}\) biases the scheduler toward shorter trajectories, enabling them to complete earlier and freeing KV cache for longer trajectories. Similarly, a negative \(\text{scale}_\text{inflight}\) biases scheduling toward less loaded servers, reducing load imbalance across rollout replicas. \(\text{scale}_\text{waiting}\), \(\text{scale}_\text{turn}\), \(\text{scale}_\text{cache}\), and \(\text{scale}_\text{inflight}\) are determined experimentally such that the score components have the following order of magnitude: \[|S_\text{waiting}| \gg |S_\text{turn}| \gg |S_\text{cache}| \gg |S_\text{inflight}|.\] Once the sandbox reaches \(\text{ttl}_{\text{sandbox}}\), the trajectory is terminated, resulting in an incomplete rollout. Therefore, \(|S_\text{waiting}|\) is assigned the largest magnitude to prioritize requests approaching the sandbox timeout. The coefficient \(\text{scale}_\text{ttl}\) is set between 0 and 1 to account for additional overheads, such as scheduling and communication, ensuring that a request is dispatched before the sandbox shuts down. \(|S_\text{turn}|\) is the second largest because we want short trajectories finished as early as possible.
\(|S_\text{inflight}|\) is assigned the smallest magnitude because request load balancing is primarily handled by a stronger strategy inspired by CONCUR [20]. CONCUR is a cache-aware request control algorithm based on the Additive Increase Multiplicative Decrease (AIMD) algorithm in TCP Congestion Control. Whether or not a rollout replica has free slot to a request is based on the KV cache Usage \(U\): \[W_{t} = \begin{cases} W_{t-1} + \alpha, & \text{if } U_{t-1} < U_\text{low},\\ W_{t-1} \times \beta, & \text{if } U_{t-1} > U_\text{high} \text{ and } H_{t-1} < H_\text{thresh},\\ W_{t-1}, & \text{otherwise}. \end{cases}\] where \(W_t\) and \(W_{t-1}\) are the maximum number of request on each inference instance in successive update time. We modify the algorithm as follows: \[\begin{align} \hat{W}_{t} &= \begin{cases} W_{t - \delta} + \alpha, & \text{if } \bar{H} < H_\text{low},\\ W_{t - \delta} \times \beta, & \text{if } \bar{H} > H_\text{high},\\ W_{t - \delta}, & \text{otherwise}. \end{cases} \\[0.8em] W_{t} &= \min\left( \max\left(1, \hat{W}_{t}\right), N^{\text{inflight}}_\text{rep} + \frac{N^{\text{queue}}_\text{req}}{N_\text{rep}} \right). \end{align}\] where \(H_\text{low}\) and \(H_\text{high}\) are the low and high threshold of cache hit rate. \(\bar{H}\) is the mean cache hit rate for the successive \(\delta\) requests. \(W_{t - \delta}\) is the maximum number of request before \(\delta\) requests, meaning that \(W_t\) is not updated until \(\delta\) requests are finished in order to stablize the maximum limit on each rollout replica. We bound \(W_t\) between 1 and \(N^{\text{inflight}}_\text{rep} + (N^{\text{queue}}_\text{req} / N_\text{rep})\), meaning that at least 1 request can be served on each rollout replica and the increased requests should not be more than number of queued requests averaged by number of rollout replica.
SuffixDecoding [11] is a model-free speculative decoding algorithm based on suffix trees. It caches token-sequence patterns from previously generated trajectories and uses them during decoding to predict likely subsequent tokens. These predicted tokens are then verified by the target model, allowing the rollout process to advance by multiple tokens in a single forward pass when the drafts are accepted. SuffixDecoding is well suited for agentic RL rollouts for the following reasons:
It exploits token patterns from real generated trajectories, which can lead to high acceptance rates in workloads with repeated local structures.
It introduces no additional draft model and therefore avoids GPU contention with rollout generation.
It incurs modest memory overhead, as it stores token-sequence structures rather than extra model parameters.
In agentic RL training, the rollout workload varies across steps and across servers. When the effective batch size is small, SuffixDecoding can provide significant acceleration because draft verification introduces limited overhead while accepted drafts allow the model to advance multiple tokens per step. However, when the effective batch size becomes large, the benefit of speculative decoding can diminish: the additional verification for draft tokens may offset the speedup from accepted tokens, leading to reduced throughput. This motivates an adaptive rollout system that applies SuffixDecoding when it is beneficial while avoiding unnecessary overhead under high-load conditions.
We implement SuffixDecoding in inference engine based on its open sourced code base with additional features:
Adaptive enabling and disabling. We dynamically enable or disable speculative decoding according to the current batch size, allowing the system to exploit suffix decoding under low-load conditions while avoiding unnecessary overhead under high load.
Stateless mode switching. We ensure safe switching between speculative and non-speculative execution without introducing state inconsistency or memory leaks.
CUDA Graph compatibility. We support separate CUDA Graph capture and replay for speculative and non-speculative execution modes, enabling efficient execution under both settings.
External update interface. We expose APIs that allow external frameworks to proactively monitor and update the suffix tree with newly generated trajectories, enabling cross-step and cross-worker suffix reuse.
In RL training, each rollout worker maintains its own suffix tree independently. When a request is dispatched to a worker, the generated trajectory is only used to update the local suffix tree of that worker. This design leads to several limitations:
Generated sequence patterns cannot be shared across different rollout workers.
Responses from similar prompts cannot be effectively reused by other workers.
Each worker can only leverage its own local history, resulting in limited historical coverage.
To address these limitations, we design a cross-worker suffix tree synchronization mechanism in the RL framework:
After each RL step, broadcast the complete prompt-response sequences generated in that step to all rollout workers.
Use the HTTP interface of the Suffix Cache Server to asynchronously update the suffix trees.
Avoid blocking the main training pipeline whenever possible.
WAR is implemented on top of veRL v0.8.0. We use SGLang v0.5.5 as the rollout inference engine and integrate SuffixDecoding into its inference pipeline. The reward and training components of veRL remain unchanged, with Megatron-LM v0.13.1 used as the training backend. To support software engineering tasks, we additionally implement an SWE agent in veRL, enabling multi-turn interaction with a customized sandbox environment.
WAR is evaluated with Qwen3-32B on an SWE dataset containing 500 prompts. All experiments are conducted on an InfiniBand-connected H100 cluster with 2 nodes and 8 GPUs per node. Qwen3-32B is extended with YaRN [21] to support a 96K context length.
The training-side tensor parallelism (TP), pipeline parallelism (PP), and context parallelism (CP) are configured as 4, 2, and 2, respectively, while the rollout-side TP is set to 4. These parallelism settings are kept unchanged throughout the evaluation. To evaluate WAR across different workload regimes, we vary the training batch size from 4 to 64, specifically using batch sizes of 4, 8, 16, 32, and 64. Each prompt is rolled out 8 times, and each experiment is run for 5 training steps.
We use end-to-end rollout time, prefill throughput, and decode throughput as the primary metrics. We also report cache hit rate and acceptance length as auxiliary metrics to analyze the behavior of cache-aware scheduling and SuffixDecoding.
| sched_only | spec_only | sched_spec | |
|---|---|---|---|
| Prefill tokens | \(2 \pm 4\) | \(1 \pm 2\) | \(-2 \pm 1\) |
| Decode tokens | \(1 \pm 4\) | \(0 \pm 4\) | \(-2 \pm 2\) |
| Assistant turns | \(1 \pm 2\) | \(0 \pm 2\) | \(-1 \pm 1\) |
We set the temperature, \(\text{top}_p\), and \(\text{top}_k\) to 1.0, 1.0, and 200, respectively, to maintain the diversity of rollout trajectories. Data shuffling is disabled to facilitate fair comparison across different experiment settings. The maximum number of assistant turns is set to 120 to allow sufficient interaction between the agent and the sandbox.
To ensure that each experiment setting is comparable to the veRL baseline, we select only runs whose relative differences in prefill tokens, decode tokens, and assistant turns are all within 5% compared with the veRL baseline, as shown in Table 1. This filtering reduces the impact of workload variation and ensures that the measured throughput differences mainly reflect system-level optimizations rather than differences in generated workload.
To understand the contribution of each component, we evaluate four configurations that enable or disable cache-aware scheduling and SuffixDecoding independently:
veRL_base. The baseline configuration where both cache-aware scheduling and SuffixDecoding are disabled.
sched_only. Only cache-aware scheduling is enabled.
spec_only. Only SuffixDecoding is enabled.
sched_spec. Both cache-aware scheduling and SuffixDecoding are enabled.
Sched_only becomes increasingly effective as workload grows. When the batch size increases from 4 to 64, relative improvement of in prefill throughput (Figure 4) increases from 0% to 47%, while its
improvement in decode throughput (Figure 5) increases from -4% to 47%. This indicates that cache-aware scheduling is more effective under heavier load, especially when the batch size is no smaller than 16.
In contrast, spec_only shows the opposite trend. Its relative improvement in prefill throughput decreases from 53% to 5%, and its improvement in decode throughput decreases from 48% to 5% as the batch size increases from 4 to 64. This
suggests that SuffixDecoding is most effective under lower load, namely when the batch size is smaller than 16.
By combining the two techniques, sched_spec achieves robust throughput improvements across the full workload range. Its prefill throughput improvement increases from 35% to 63%, and its decode throughput improvement increases from 32% to
63% as the batch size grows from 4 to 64. A similar complementary effect is also observed in end-to-end rollout time, as shown in Figure 6.
We further investigate cache hit rate and acceptance length, which are closely related to cache-aware scheduling and SuffixDecoding, respectively.
As shown in Figure 7, when the load is low, namely when the batch size is smaller than 16, all experiment settings achieve a very high cache hit rate, reaching 0.94 and 0.93 for batch sizes 4 and 8, respectively.
As the load increases, the cache hit rates of veRL_base and spec_only drop significantly, decreasing from around 0.9 at low load to 0.1 when the batch size reaches 64. This is because cache-aware scheduling is disabled in these
two settings, and the default scheduling policy relies on least-request scheduling with sticky sessions. Least-request scheduling is unaware of KV-cache locality. Sticky sessions provide implicit cache awareness by routing requests from the same session to
the same rollout instance, but the corresponding KV cache can still be evicted during long environment interactions when it is not revisited for a long time.
In contrast, with cache-aware scheduling enabled, both sched_only and sched_spec maintain a relatively high cache hit rate even under high load, achieving around 0.5 when the batch size is 64.
| batch size | 32 | 64 |
|---|---|---|
| Prefill tokens | -11.7 | -44.5 |
| Decode tokens | -8.6 | -41.3 |
| Assistant turns | -9.4 | -38.9 |
| Rollout end-to-end time | -12.0 | -46.9 |
We perform an ablation study to evaluate the effect of the four scheduling factors introduced in Section 3.2. We first examine the role of the waiting factor, \(\text{scale}_{\text{waiting}}\), by setting it to either zero or a non-zero value while keeping all other factors unchanged, as shown in Table 2.
As the rollout load increases, some queued requests may wait for a long time before being scheduled. If a request is not dispatched before the corresponding sandbox reaches its TTL, the sandbox shuts down and the agent trajectory is terminated prematurely. This leads to incomplete rollouts, which is reflected in the overall reduction of prefill tokens, decode tokens, and end-to-end time. As the load becomes heavier, i.e., when the batch size increases from 32 to 64, this rollout degradation becomes more severe.
For the ablation study of \(S_\text{turn}\), \(S_\text{cache}\), and \(S_\text{inflight}\), we keep \(S_\text{waiting}\)
enabled to preserve the integrity of rollout trajectories, as shown in Figure 8. Among the three ablation settings, namely turn_only, cache_only, and inflight_only,
cache_only achieves the best performance, with the highest throughput and cache hit rate. This highlights the importance of cache locality in rollout scheduling.
Notably, none of the ablation settings outperforms sched_only, which corresponds to the full scheduling configuration. This indicates that the three scheduling components are complementary and jointly contribute to the overall
performance.
As shown in Figure 9, when SuffixDecoding is enabled, i.e., in spec_only and sched_spec, the acceptance length reaches around 2.8 when the batch size is smaller than 16, and around 6.0 when the
batch size is no smaller than 16. This demonstrates that SuffixDecoding can effectively generate useful draft tokens for SWE tasks. We further observe from the rollout logs that the model often generates recurring code snippets, which provide favorable
matching opportunities for SuffixDecoding.
However, the high acceptance length does not always translate into higher throughput. A key observation of speculative decoding is that its benefit depends on the batch size. When the batch size exceeds a certain threshold, speculative decoding may no longer yield performance gains, and disabling it can result in better throughput. This is because speculative decoding is most beneficial when normal decoding underutilizes the GPU. At small batch sizes, verifying multiple draft tokens in one forward pass increases parallelism and improves hardware utilization. As the batch size grows, normal batched decoding already saturates GPU resources, leaving limited room for additional speculative speedup. Meanwhile, speculative decoding still incurs extra overhead from draft generation, verification, and tree-mask management. When this overhead outweighs the reduction in decoding steps, disabling speculative decoding becomes more efficient.
We observe the same phenomenon for SuffixDecoding. As shown in Figure 10, once the batch size exceeds 16, SuffixDecoding begins to underperform the non-SuffixDecoding baseline in decode throughput. This result indicates that SuffixDecoding should be enabled under low load but disabled under high load.
We conduct an ablation study to analyze the sources of the high acceptance length, as shown in Table 3. The results indicate that SuffixDecoding benefits from both cross-prompt similarity and intra-prompt rollout similarity. Different prompts may generate similar code fragments, while multiple rollouts of the same prompt often share similar local structures. Both types of recurring patterns provide useful suffix matches for draft generation and contribute to the high acceptance length.
| batch size | rollout_n | ||
|---|---|---|---|
| 2-4 | 4 | 8 | |
| 8 | 2.7 | 2.7 | |
| 16 | 2.6 | 5.5 | |
Recent RL training systems improve the efficiency of LLM post-training from different perspectives. Stage-overlap systems, such as RLHFuse [22], reduce pipeline bubbles by fusing rollout, reward computation, and training-related stages. Other systems improve resource utilization through asynchronous execution or relaxed synchronization [6], [15]–[17]. These approaches can reduce end-to-end training latency, but they either overlap non-dominant stages with rollout or relax strict on-policy semantics. In contrast, WAR focuses on accelerating synchronous rollout itself, without changing the underlying RL algorithm or relaxing the synchronization boundary.
Several recent systems address the long-tail inefficiency of synchronous rollout through scheduling. RollPacker [14] proposes tail batching, which consolidates prompts likely to produce long responses into a small number of long rounds while keeping most rounds short and balanced. Seer [5] decomposes prompt groups into finer-grained chunks and uses context-aware scheduling to reduce load imbalance and KV-cache pressure. These systems show that scheduling is crucial for efficient synchronous RL rollout. WAR is complementary to them: instead of focusing only on long-tail scheduling, it treats rollout optimization as workload-dependent and combines cache-aware scheduling with decoding-level acceleration.
Speculative decoding accelerates autoregressive generation by drafting multiple tokens and verifying them with the target model in parallel. For RL rollout, recent systems adapt speculative decoding to address model drift, long-tail latency, and distributional correctness. TLT [23] maintains an adaptive model-based drafter using idle GPUs during the long-tail phase. DAS [24] constructs a distribution-aware suffix-tree drafter from recent rollouts and allocates more speculative budget to long-tail trajectories. These methods demonstrate the potential of speculative decoding in RL training, but speculative decoding is not uniformly beneficial across all workload regimes. WAR therefore applies model-free SuffixDecoding under low load, where GPU resources are underutilized, and disables it under high load, where speculation overhead can outweigh its benefits.
Efficient LLM serving systems improve throughput through continuous batching, KV-cache management, prefix caching, and request scheduling. These techniques are largely orthogonal to RL training, but they become increasingly important for long-context agentic rollouts, where KV-cache recomputation and memory pressure dominate high-load performance. WAR brings cache-aware scheduling into synchronous agentic RL rollout by considering cache locality, trajectory progress, and server load when assigning requests to rollout replicas. Unlike general-purpose serving systems, WAR exploits RL-specific workload structure and adapts its optimization strategy according to runtime load.
This paper argues that workload should be treated as a first-class factor in optimizing synchronous agentic RL rollouts. Long-horizon agent trajectories create distinct bottlenecks under different runtime loads: low-load rollouts suffer from insufficient GPU utilization, while high-load rollouts are dominated by KV-cache pressure, redundant recomputation, and load imbalance. We propose WAR, a workload-aware rollout system that addresses these bottlenecks by jointly optimizing decoding and scheduling.
WAR applies model-free SuffixDecoding under low load to reuse historical suffix patterns as speculative drafts without introducing draft-model GPU contention. Under high load, WAR shifts the optimization focus to cache-aware scheduling, placing requests across rollout replicas based on cache locality, trajectory progress, and server load. This simple design preserves synchronous RL semantics and remains compatible with existing RL training pipelines. Experiments on long-context agentic workloads show that WAR improves rollout throughput by \(1.4\times\) under low load and up to \(1.6\times\) under high load over the veRL baseline. These results show that workload-aware rollout optimization provides a practical path toward scalable synchronous agentic RL training.