July 05, 2026
Large language model (LLM) agents are usually improved by changing prompts, models, or hand-written workflows, while the execution harness around the model is treated as fixed infrastructure. We argue that this harness is itself a learnable control layer. We formalize harness operation as a finite-horizon Harness MDP, where a lightweight controller selects structural execution actions while the LLM executor remains frozen. The controller is trained from offline rollouts using advantage-weighted regression with only terminal task-rubric rewards. We also separate final task quality from a post-hoc Harness Maturity Score, which measures whether the harness follows reliable execution patterns rather than only whether the final answer is correct. This separation gives a finite-buffer view of harness learning: final-quality gains require high-return support in the offline buffer, while process behavior can shift whenever it correlates with advantage weights. Across six controlled domains and two public-benchmark adapters, the learned controller consistently improves verification behavior and selectively improves final quality, with the largest gains on adapted \(\tau\)-bench retail, adapted AgentBench DB-Bench, and coding with a calibrated structural verifier. Ablations against behavior cloning and Forced CHECK show that the gains are not explained by imitation or by simply adding checks. These results identify harness control as a learnable layer for frozen LLM agents, while showing that offline support limits when better process control becomes better final answers. Our code is available at https://github.com/Hik289/Agentic-RL-harness.git.
An LLM agent is more than a language model with a prompt. In deployed systems, a surrounding harness decides when the model observes the environment, retrieves evidence, calls tools, drafts an answer, checks intermediate work, revises, and stops [1]–[4]. Holding the executor fixed, this control layer can change the agent’s behavior substantially: a coding agent may or may not run tests before submission, a research agent may or may not gather evidence before making a claim, and a tool-use agent may or may not recover from a failed call [4]–[9]. The harness is therefore not a cosmetic wrapper around inference; it is part of the policy that determines how work is carried out.
The difficulty is that harnesses are usually built as static artifacts. Existing systems compose hand-written rules, prompt templates, multi-agent scaffolds, or fixed execution graphs [3], [10], [11]. Such designs can be highly effective in a narrow regime, but they encode the same control logic for easy and hard tasks, for early and late trajectory states, and for successful and failed intermediate attempts. A fixed rule such as “always check” wastes budget when a solution is already complete; a rule such as “submit after drafting” fails when the draft is plausible but unverified. The failure mode is often not missing model knowledge, but a poor sequence of external control decisions.
This paper studies whether that sequence can be learned directly. We define a Harness Markov decision process in which the state summarizes the current trajectory, draft status, collected evidence, tool outputs, verifier feedback, previous failures, and remaining budget. At each step, a controller chooses one structural action from observe, retrieve, call-tool, draft, check, revise, and submit. The chosen action may invoke the frozen LLM executor, but the LLM parameters and prompts are not updated. Learning is thus confined to the external control policy: the controller learns how the agent should move through observation, evidence collection, verification, revision, and termination.
Online exploration in agent environments is expensive and can produce invalid or low-quality trajectories. We therefore train from a finite offline rollout buffer using advantage-weighted regression (AW) [12]. Each trajectory receives a terminal task-rubric return, and the controller increases the likelihood of state-action pairs that appear in higher-advantage trajectories. This choice gives a conservative learning rule: the learned policy can reuse good control patterns already present in the buffer, but it cannot reliably invent outcome-improving trajectories that the buffer never contains. Figure 1 gives the central intuition.
The main conceptual distinction is between final task quality and process quality. We train only on terminal task-rubric rewards, denoted by \(\mathcal{Q}\), and evaluate harness process behavior separately with the Harness Maturity Score \(\mathrm{HMS}\). \(\mathrm{HMS}\) measures events such as checking before submission, testing before submission, grounding claims in evidence, revising after failure, valid tool use, sufficient stopping, and premature submission. This separation is deliberate. A controller can learn a recurring process behavior, such as checking before submission, whenever that behavior is positively associated with advantage-weighted trajectories. In contrast, final-quality improvement requires high-return trajectories with enough support in the offline buffer.
Our empirical story follows this separation. Across six controlled domains and two public-benchmark adapters, AW increases verification before submission in every evaluation setting. Aggregate \(\mathrm{HMS}\) improves in five of six controlled domains and in both adapter evaluations, but the improvement is localized: it is driven mainly by CheckBeforeSubmit rather than by uniform gains across all process events. Final task quality improves selectively. Coding improves under the calibrated structural verifier; the \(\tau\)-bench retail and AgentBench DB-Bench adapters show the largest external gains, although these are adapter-level evaluations rather than official upstream benchmark scores. The pattern suggests that harness process can be learned broadly from offline traces, while outcome gains appear where the buffer contains useful high-return support.
Our contributions are summarized as follows:
We formulate the external control layer of a frozen LLM agent as a Harness MDP, in which a learned policy selects the next structural operation according to the current trajectory state.
We develop an offline RL method based on advantage-weighted regression that learns harness-control decisions from finite rollout buffers and task-rubric terminal rewards, without updating the LLM or directly rewarding predefined action patterns.
We separate final task quality \(\mathcal{Q}\) from harness process quality \(\mathrm{HMS}\) and provide a finite-buffer analysis showing why process behavior may improve even when gains in \(\mathcal{Q}\) remain limited by offline trajectory support.
We evaluate the framework across six controlled domains and two public-benchmark adapters using a shared action space, task-specific structural verifiers, and a seven-event process diagnostic. The results show reliable improvements in verification-before-submission and selective final-quality gains when the offline buffer contains useful high-return trajectories.
AgentBench, \(\tau\)-bench, GAIA, WebArena, WebShop, SWE-bench, AgentGym, ToolBench, ToolSandbox, and long-memory benchmarks evaluate tool use, web navigation, coding, retrieval, state tracking, and API-mediated interaction [5]–[9], [13]–[17]. These benchmarks reveal long-horizon failures of complete agent systems. Our goal is narrower: we isolate the external control policy around a fixed executor and ask whether that policy can be learned.
Prompt and workflow optimizers search over instructions, demonstrations, programs, or code-represented scaffolds [18]–[25]. Automatic harness-evolution methods further revise the program around an agent using trajectory feedback [26]–[28]. These approaches optimize an artifact before execution. We instead learn a state-conditioned controller that chooses the next harness operation during execution, so two tasks can follow different control paths even under the same frozen LLM.
Reflection and planning systems such as Reflexion, Voyager, LATS, and ADaPT provide structured feedback, revision, search, and decomposition procedures [29]–[32]. WebShop and AgentGym show that interactive agents can be trained or fine-tuned with reinforcement learning and behavior cloning [14], [15]. Our setting differs in two ways: the executor is frozen, and the learned policy acts only over a compact set of harness operations rather than over natural-language solution tokens or environment actions.
Offline RL studies how to learn from a fixed dataset without collecting additional experience [33], [34]. We use advantage-weighted regression [12], a simple weighted-imitation objective related to advantage-weighted actor-critic methods [35] and to conservative offline policy improvement methods such as implicit Q-learning [36]. Behavior cloning is a natural imitation baseline [37], [38]; Forced CHECK is a harness-specific baseline that tests whether the observed gains can be explained by mechanically inserting verification rather than by learning when verification is useful.
Final-answer accuracy can hide failures in verification, evidence use, revision, and stopping behavior [39]–[41]. We therefore report \(\mathrm{HMS}\) as a process diagnostic, but we do not optimize it directly. The distinction follows the reward-shaping principle that auxiliary rewards should not change the task-optimal policy set unless they are potential-based [42]. Our analysis makes this separation explicit for finite offline harness buffers.
The method learns a small controller around a fixed LLM executor. The controller observes compact trajectory features, selects the next harness operation, and receives supervision only through terminal task-rubric rewards collected in an offline buffer. Figure 2 summarizes the pipeline.
A state vector summarizes the information needed for external control: trajectory progress, draft availability, evidence coverage, tool outputs, verifier feedback, recent failures, remaining budget, last action, and domain-specific structural flags. The state does not expose hidden model activations or update the LLM. It is designed to capture whether the agent has enough information to continue, verify, revise, or stop.
The controller selects among seven structural actions: observe, retrieve, call-tool, draft, check, revise, and submit. Domain adapters implement the semantics of each action and mask invalid actions, while the learned policy shares a common interface across domains. The action space is intentionally structural: it controls the execution procedure rather than the content of the LLM’s generated answer.
Each domain maps a terminal artifact to criterion-level rubric scores and a scalar task reward. The verifier interface is shared, but the criteria are domain-specific: coding uses tests and patch constraints, research uses answer and evidence criteria, multi-tool tasks use numerical and tool-structure criteria, long-memory tasks use source-session matching, and planning tasks use state-transition constraints. This design keeps the reward tied to task quality rather than to a preferred action pattern. The multi-tool verifier is structural and numeric: it validates the final artifact and required tool structure, but it does not claim that every intermediate tool call was uniquely necessary.
Finite rollout buffers are collected from base and exploratory harnesses. Terminal rubric rewards define trajectory advantages, and the controller is trained by weighted behavior cloning: \[\begin{align} \mathcal{L}(\theta) &=-\mathbb{E}_{(s,a)\sim \mathcal{D}}\big[ \exp(A(s,a)/\beta)\\ &\qquad\qquad\cdot \log \pi_\theta(a\mid s) \big]. \end{align}\] The exponential weight increases the likelihood of actions observed on higher-rubric trajectories. Because learning is offline, AW is expected to recover useful control patterns present in \(\mathcal{D}\), not to solve exploration gaps by itself.
\(\mathrm{HMS}\) is a normalized weighted diagnostic over seven process events: CheckBeforeSubmit, EvidenceBeforeClaim, TestBeforeSubmit, RevisionAfterFailure, ValidToolUse, StopWhenSufficient, and EarlySubmit as a penalty. \(\mathrm{HMS}\) is never used as a training reward. It is reported to identify which harness behaviors changed after optimizing the terminal task objective.
The EarlySubmit detector uses threshold 0.25. Coding uses the same calibrated structural verifier family as the other domains rather than the original strict deterministic coding rubric. This choice changes the coding baseline and therefore the measured coding gain; Section 5 and Appendix 12 report the sensitivity analysis.
The harness controller is a one-hidden-layer MLP with 64 hidden units and a softmax over seven candidate actions, with invalid actions masked by the domain adapter. The state vector contains trajectory progress, draft status, rubric coverage, error count, remaining budget, recent test outcomes, and the previous action. Training uses Adam with learning rate \(10^{-3}\), batch size 256, 20 epochs, AW temperature \(\beta=0.2\), clipped AW weights in \([0.1,10.0]\), and entropy regularization coefficient 0.01. Three independent seeds are trained per domain. Hyperparameters are fixed across domains to avoid attributing domain variation to per-domain tuning.
The process-error normalizer \(P_{\mathrm{error}}\) uses max_steps rather than observed steps, preventing shorter trajectories from receiving a mechanical process advantage. The same rubric scorer is used to
produce terminal training rewards and evaluation scores; calibration mode is used only for scorer-validation diagnostics. Multi-tool criteria are structural and numeric: they verify the final artifact and required tool structure but do not certify ideal
intermediate reasoning. All primary claims report mean changes and effect magnitudes, with bootstrap intervals for final-quality changes.
We first require policy consistency: any reward used for harness training should preserve the optimal policies of the original terminal task objective. The statement below is the finite-horizon analogue of potential-based reward shaping [42]. Our experiments use the special case \(\Phi_t\equiv 0\), i.e., terminal task-rubric reward without intermediate action bonuses.
Definition 1 (Task-rubric return). Consider a finite-horizon Harness MDP with horizon \(H\), initial-state distribution \(\rho_0\), and trajectory \(\tau=(s_0,a_0,\ldots,s_H)\). The task-rubric return of a policy \(\pi\) is \[J_{\mathcal{Q}}(\pi)=\mathbb{E}_{\tau\sim p_\pi}\left[\mathcal{Q}(s_H)\right], \label{eq:task-rubric-objective}\qquad{(1)}\] where \(p_\pi\) is the trajectory distribution induced by \(\rho_0\), the harness policy \(\pi\), and the frozen-LLM transition kernel.
Definition 2 (Potential-shaped task reward). Let \(\{\Phi_t:\mathcal{S}\rightarrow\mathbb{R}\}_{t=0}^{H}\) be a sequence of bounded potential functions satisfying \(\Phi_H(s)=0\) for every terminal state \(s\). Define the shaped per-step reward as \[\begin{align} r_t^{\Phi}(s_t,a_t,s_{t+1}) &=r_t(s_t,a_t,s_{t+1})\\ &\quad+\Phi_{t+1}(s_{t+1})-\Phi_t(s_t), \end{align} \label{eq:potential-shaped-reward}\qquad{(2)}\] where \(r_t=0\) for \(t<H-1\) and \(r_{H-1}(s_{H-1},a_{H-1},s_H)=\mathcal{Q}(s_H)\). The corresponding shaped objective is \[J_{\Phi}(\pi)=\mathbb{E}_{\tau\sim p_\pi}\left[\sum_{t=0}^{H-1}r_t^{\Phi}(s_t,a_t,s_{t+1})\right]. \label{eq:potential-shaped-objective}\qquad{(3)}\]
Theorem 1 (Optimal-policy invariance). Under Definitions 1 and 2, there exists a constant \(C_{\rho_0}\) independent of \(\pi\) such that \[J_{\Phi}(\pi)=J_{\mathcal{Q}}(\pi)-C_{\rho_0}, \label{eq:objective-equivalence}\qquad{(4)}\] where \(C_{\rho_0}=\mathbb{E}_{s_0\sim\rho_0}[\Phi_0(s_0)]\). Consequently, \[\operatorname*{argmax}_{\pi}J_{\Phi}(\pi)=\operatorname*{argmax}_{\pi}J_{\mathcal{Q}}(\pi). \label{eq:optimal-policy-set}\qquad{(5)}\]
Theorem 1 states that task-rubric reward can be augmented with potential-based intermediate signals without changing the task-optimal policy set. Our experiments use no action-pattern bonus: checking, revising, or calling tools is valuable only insofar as it contributes to terminal task quality. The proof is provided in Appendix 8.1.
Proposition 1 (Action-pattern rewards can change the optimum). There exists a finite-horizon Harness MDP and an action-pattern reward \(b(a)\) such that \[\begin{align} &\operatorname*{argmax}_{\pi}\mathbb{E}_{\pi}\left[\mathcal{Q}(s_H)\right]\\ &\qquad\neq \operatorname*{argmax}_{\pi}\mathbb{E}_{\pi}\left[\mathcal{Q}(s_H)+b(a_0)\right]. \end{align} \label{eq:action-pattern-noninvariance}\qquad{(6)}\]
Proposition 1 shows that direct bonuses for actions such as check or revise can favor process forms that do not improve the final artifact. Its proof is provided in Appendix 8.2.
We next characterize the ideal support-restricted AW update induced by a finite rollout buffer. The result gives an outcome ceiling and an exact expression for the shift in any bounded process statistic.
Definition 3 (Finite rollout buffer and AW distribution). Let \(B=\{\tau^{(1)},\ldots,\tau^{(N)}\}\) be a finite rollout buffer with empirical distribution \(\mu_B(\tau)=N^{-1}\sum_{i=1}^{N}\mathbb{I}\{\tau=\tau^{(i)}\}\). For terminal return \(G:B\rightarrow\mathbb{R}\), define the empirical mean return \(\overline{G}_B=\mathbb{E}_{\mu_B}[G]\), the support ceiling \(G_B^\star=\max_{\tau\in B}G(\tau)\), and the buffer slack \(\sigma_B=G_B^\star-\overline{G}_B\). Let \(A:B\rightarrow\mathbb{R}\) be an advantage estimate and let \(\beta>0\). Define \(w(\tau)=\exp(A(\tau)/\beta)\) and \(Z_B=\mathbb{E}_{\mu_B}[w]\). The support-restricted AW distribution is \[q_{\mathrm{AW}}(\tau)=\frac{\mu_B(\tau)w(\tau)}{Z_B}. \label{eq:aw-trajectory-distribution}\qquad{(7)}\]
Theorem 2 (Finite-buffer outcome and process characterization). Let \(q_{\mathrm{AW}}\) be defined by Equation ?? .
(i) Outcome bound. Define \(\Delta\mathcal{Q}_B=\mathbb{E}_{q_{\mathrm{AW}}}[G]-\mathbb{E}_{\mu_B}[G]\). Then \[\mathbb{E}_{q_{\mathrm{AW}}}[G]\leq G_B^\star, \qquad \Delta\mathcal{Q}_B\leq\sigma_B. \label{eq:finite-buffer-outcome-bound}\qquad{(8)}\]
(ii) Process identity. For any bounded process statistic \(\Psi:B\rightarrow\mathbb{R}\), define \(\Delta\Psi_B=\mathbb{E}_{q_{\mathrm{AW}}}[\Psi]-\mathbb{E}_{\mu_B}[\Psi]\). Then \[\Delta\Psi_B=\frac{\operatorname{Cov}_{\mu_B}(w,\Psi)}{\mathbb{E}_{\mu_B}[w]}. \label{eq:exact-process-shift}\qquad{(9)}\] Since \(\mathbb{E}_{\mu_B}[w]>0\), \(\Delta\Psi_B\geq 0\) if and only if \(\operatorname{Cov}_{\mu_B}(w,\Psi)\geq 0\).
Corollary 1 (Process improvement under monotone advantage weighting). Suppose there exists a measurable non-decreasing function \(h:\mathbb{R}\rightarrow\mathbb{R}\) such that \(\mathbb{E}_{\mu_B}[\Psi(\tau)\mid A(\tau)]=h(A(\tau))\). Then \(\operatorname{Cov}_{\mu_B}(w,\Psi)\geq 0\), and hence \(\Delta\Psi_B\geq 0\).
Proposition 2 (Terminal-return correlation does not determine the process shift). There exist finite buffers, returns \(G\), process statistics \(\Psi\), and advantage estimates \(A\) such that \[\operatorname{Cov}_{\mu_B}(G,\Psi)<0, \qquad \operatorname{Cov}_{\mu_B}(w,\Psi)>0. \label{eq:opposite-process-covariances}\qquad{(10)}\] For such a construction, Equation ?? implies \(\Delta\Psi_B>0\) despite a negative return–process covariance.
Theorem 2 separates the two channels of offline harness learning. Outcome improvement is bounded by the return support of the buffer. Process improvement is governed by the covariance between the process statistic and the AW weights. In the experiments, \(\Psi\) is either an individual harness event or aggregate \(\mathrm{HMS}\); for domain \(D\), \(\sigma_D:=\sigma_{B_D}\) denotes empirical buffer slack. Proofs are provided in Appendix 9, and Figure 5 evaluates the empirical relationship between slack and outcome improvement.
The experiments test whether a learned harness controller changes the way a frozen LLM agent works, and whether those process changes improve final task quality. We evaluate six controlled domains: knowledge-work, coding, research QA, multi-tool, long-memory, and planning. Each domain has 100 human-annotated tasks, with 80 used for buffer collection and 20 held out for evaluation. The split is stratified by difficulty and uses disjoint task identifiers, templates, entities, and starter artifacts; Appendix 11 details the construction and leakage controls. We also evaluate two public-benchmark adapters: \(\tau\)-bench retail for policy-constrained tool interaction [8], and AgentBench DB-Bench for interactive database reasoning [7]. These adapters map selected tasks into the Harness MDP so that we can isolate harness control rather than claim official benchmark scores. Each adapter uses 16 training tasks and 20 held-out evaluation tasks. All evaluations use three independent seeds and three rollouts per held-out task.
We compare AW with behavior cloning (BC) and Forced CHECK (FC). BC tests whether copying observed trajectories is enough [37], [38]; FC tests whether simply adding verification explains the gains. Table 1 summarizes the empirical message: AW reliably learns verification behavior, but final-quality gains appear only when the offline buffer contains useful high-return trajectories.
| Finding | Result | Evaluation |
|---|---|---|
| Consistent improvement in verification | CheckBeforeSubmit increases from near-zero base rates to 5.6–17.8% across the controlled domains, 17.2% on \(\tau\)-bench retail, and 16.7% on AgentBench DB-Bench. | Six controlled domains and two benchmark adapters |
| Largest adapter-level gain on \(\tau\)-bench retail | Final quality improves by 18.2% under the adapted plan-quality rubric. This is external-validation evidence under our rubric, not an official \(\tau\)-bench score. | 16 training / 20 held-out tasks |
| Positive transfer to AgentBench DB-Bench | Final quality improves by 13.2% under the adapted deliberative-reasoning rubric. This is external-validation evidence, not an official AgentBench score. | 16 training / 20 held-out tasks |
| Strongest controlled-domain gain in coding | Final quality improves by 10.0% under the calibrated structural verifier. Under the original strict coding rubric, the base harness is near ceiling, so this gain is tied to the calibrated protocol. | 80 training / 20 held-out tasks |
6pt
We first ask whether AW changes the execution procedure at all. The clearest answer is verification before submission. Under the base harness, CheckBeforeSubmit is nearly absent across the controlled domains and the two adapter evaluations. After AW training, Table 2 shows that the event appears in every setting. This is the central process result: the controller learns a reusable control habit from offline trajectories even when the downstream score change is small.
| Domain | Base CBS | AW CBS |
|---|---|---|
| Knowledge-work | 0% | 5.6% |
| Coding | 0% | 17.8% |
| Research | 0% | 13.9% |
| Multi-tool | 0% | 8.9% |
| Long-memory | 0% | 10.0% |
| Planning | 0.6% | 6.7% |
| \(\tau\)-bench retail | 0% | 17.2% |
| AgentBench DB-Bench | 0% | 16.7% |
Process gains do not automatically imply better final artifacts, so we next examine final task quality. Figure 3 and Table 3 show that the strongest outcome gains occur in the two public-benchmark adapters and in coding. The \(\tau\)-bench retail adapter improves most under our plan-quality rubric, AgentBench DB-Bench also improves, and coding is the strongest controlled-domain result under the calibrated structural verifier. These settings support the main positive claim: when the buffer contains trajectories that connect better control decisions to better terminal artifacts, AW converts process learning into final-quality improvement.
| Setting | Base \(G\) | AW \(G\) | \(\Delta G\) | CBS |
|---|---|---|---|---|
| \(\tau\)-bench retail | 0.337 | 0.519 | \(\mathbf{+18.2}\) | 17.2% |
| interval | [\(+15.1\), \(+21.3\)] | |||
| DB-Bench | 0.415 | 0.547 | \(\mathbf{+13.2}\) | 16.7% |
| interval | [\(+10.2\), \(+16.2\)] | |||
| Coding | 0.712 | 0.812 | \(\mathbf{+10.0}\) | 17.8% |
| interval | [\(+5.9\), \(+14.6\)] | |||
2pt
As shown in Table 4 and Figure 4, AW provides stronger gains than BC in all eight settings and outperforms both BC and FC in five. This rules out two simpler explanations. The effect is not pure imitation, because BC is consistently weaker; it is not mechanical verification, because FC does not reproduce the adapter and coding gains. AW’s advantage is state dependence: the controller learns when checking and revision are useful rather than applying them uniformly.
| Setting | AW lift | BC lift | FC lift | \(\Delta\Delta_{\mathrm{AW-BC}}\) | \(\Delta\Delta_{\mathrm{AW-FC}}\) |
|---|---|---|---|---|---|
| Knowledge-work | \(+1.4\) | \(-0.4\) | \(+0.1\) | \(+1.8\) | \(+1.3\) |
| Coding | \(+10.0\) | \(-8.3\) | \(+0.0\) | \(+18.3\) | \(+10.0\) |
| Research | \(-0.3\) | \(-4.2\) | \(-0.2\) | \(+3.8\) | \(-0.1\) |
| Multi-tool | \(-1.3\) | \(-6.8\) | \(+0.3\) | \(+5.5\) | \(-1.6\) |
| Long-memory | \(-0.3\) | \(-5.8\) | \(+0.0\) | \(+5.5\) | \(-0.3\) |
| Planning | \(+2.6\) | \(+0.0\) | \(-2.3\) | \(+2.6\) | \(+4.9\) |
| \(\tau\)-bench retail | \(+18.2\) | \(+8.2\) | \(+0.1\) | \(+10.0\) | \(+18.1\) |
| DB-Bench | \(+13.2\) | \(+5.8\) | \(-0.5\) | \(+7.4\) | \(+13.7\) |
4pt
On the held-out adapters, AW gains remain positive across the available task-complexity strata. The improvements persist on more complex tasks, suggesting that the learned harness policy is not merely exploiting easy instances.
Table 5 reports the full results, and Table 6 provides a compact interpretation. Coding shows the clearest calibrated-protocol gain. Under the original strict deterministic coding rubric, the base harness was already near saturation, leaving little room for improvement. The reported +10.0 points therefore depends on the calibrated structural verifier, which supports cross-domain comparison by scoring parseability, safety constraints, cost compliance, testing, and checking. Planning and knowledge-work show smaller gains, while the remaining domains are near zero or slightly negative. Aggregate \(\mathrm{HMS}\) improves in five of six domains, mainly through verification before submission; research is the exception because EarlySubmit offsets the checking gain.
| Domain | Base \(G\) | AW \(G\) | \(\Delta G\) | \(\Delta\HMS\) | 95% CI | \(p\) | CBS |
|---|---|---|---|---|---|---|---|
| Knowledge-work | 0.450 | 0.464 | \(+0.014\) | \(+0.059\) | [\(-0.002\), \(+0.031\)] | 0.097 | 5.6% |
| Coding | 0.712 | 0.812 | \(\mathbf{+0.100}\) | \(\mathbf{+0.054}\) | [\(+0.059\), \(+0.146\)] | \(\mathbf{<0.001}\) | 17.8% |
| Multi-tool | 0.528 | 0.515 | \(-0.013\) | \(+0.058\) | [\(-0.035\), \(+0.006\)] | 0.217 | 8.9% |
| Planning | 0.385 | 0.412 | \(+0.026\) | \(+0.010\) | [\(-0.012\), \(+0.064\)] | 0.175 | 6.7% |
| Research | 0.314 | 0.311 | \(-0.003\) | \(-0.026\) | [\(-0.009\), \(+0.002\)] | 0.238 | 13.9% |
| Long-memory | 0.453 | 0.451 | \(-0.003\) | \(+0.011\) | [\(-0.008\), \(+0.000\)] | 0.449 | 10.0% |
| Macro mean | – | – | \(+0.020\) | \(+0.028\) | – | – | 10.5% |
6pt
| Domain | Base \(G\) | AW \(G\) | \(\Delta G\) | \(\Delta\HMS\) | Summary |
|---|---|---|---|---|---|
| Knowledge-work | 0.450 | 0.464 | \(+0.014\) | \(+0.059\) | Directional quality gain; process improves |
| Coding | 0.712 | 0.812 | \(\mathbf{+0.100}\) | \(+0.054\) | Largest controlled-domain quality gain |
| Research | 0.314 | 0.311 | \(-0.003\) | \(-0.026\) | EarlySubmit offsets checking gain |
| Multi-tool | 0.528 | 0.515 | \(-0.013\) | \(+0.058\) | Process improves despite lower final score |
| Long-memory | 0.453 | 0.451 | \(-0.003\) | \(+0.011\) | Near-zero final-quality change |
| Planning | 0.385 | 0.412 | \(+0.026\) | \(+0.010\) | Directional quality gain |
| Macro mean | – | – | \(+0.020\) | \(+0.028\) | Process improves in five of six domains |
6pt
Table 7 decomposes \(\mathrm{HMS}\) into seven process events. The overall improvement is driven mainly by CheckBeforeSubmit, which increases from 1 of 1,080 base episodes to 113 of 1,080 AW episodes across the controlled suite. TestBeforeSubmit also improves in coding. In contrast, EarlySubmit worsens in several settings, most notably in research, while other events remain sparse or saturated. Thus, the observed process gain is concentrated in verification before submission rather than reflecting uniform improvement across all dimensions of process maturity.
| Event | KW | Coding | Research | Multi-tool | Long-mem | Planning |
|---|---|---|---|---|---|---|
| CheckBeforeSubmit | 0%\(\to\)5.6% | 0%\(\to\)17.8% | 0%\(\to\)13.9% | 0%\(\to\)8.9% | 0%\(\to\)10.0% | 0.6%\(\to\)6.7% |
| EvidenceBeforeClaim | 0%\(\to\)0% | – | 0%\(\to\)0% | 0%\(\to\)0% | 73.8%\(\to\)73.2% | – |
| TestBeforeSubmit | – | 76.7%\(\to\)86.7% | – | – | – | – |
| RevisionAfterFailure | 0%\(\to\)0% | 0%\(\to\)4.3% | 0%\(\to\)0% | 0%\(\to\)0% | 0%\(\to\)0% | 0%\(\to\)3.4% |
| ValidToolUse | – | – | 100%\(\to\)100% | 100%\(\to\)100% | – | – |
| StopWhenSufficient | 0%\(\to\)0% | 0%\(\to\)0% | 0%\(\to\)0% | 0%\(\to\)0% | 0%\(\to\)0% | 0.6%\(\to\)0% |
| EarlySubmit | 1.7%\(\to\)2.2% | 0%\(\to\)0.6% | 0%\(\to\)25.0% | 0%\(\to\)0% | 0%\(\to\)3.3% | 51.1%\(\to\)47.8% |
| \(\Delta\HMS\) | \(+0.059\) | \(+0.054\) | \(-0.026\) | \(+0.058\) | \(+0.011\) | \(+0.010\) |
3pt
We include public benchmarks only when the harness abstraction can be applied without changing the core task. Native benchmark interfaces usually evaluate the full agent stack, whereas our adapters isolate the harness-control layer. The \(\tau\)-bench retail adapter preserves policy-constrained tool decisions but uses an adapted plan-quality rubric rather than the native simulator score. The AgentBench DB-Bench adapter preserves database reasoning and deliberative control but not the full upstream evaluator. We therefore treat these results as external validation of the learned control pattern, not as official benchmark scores. We do not report results on AgentBench OS-Interaction, SWE-bench Lite, or GAIA L1/L2 because their native evaluation interfaces were unavailable in our setup or could not be represented faithfully by the current plan-generation harness.
Figure 5 relates final-quality change to buffer slack \(\sigma_D\) under two verifier configurations. The original strict domain-specific verifier makes the coding baseline nearly saturated, which produces a strong apparent relationship between slack and outcome gain. The calibrated structural verifier replaces that saturated coding rubric with the same scoring family used across domains; the relationship then weakens because coding moves from zero to moderate estimated slack. Buffer slack is therefore useful as an offline-support diagnostic, but it should not be treated as a calibration-invariant explanation of final-quality improvement.
Table 8 reports sensitivity to the EarlySubmit threshold used by the calibrated structural detector. Larger thresholds classify more submissions as early. At a threshold of 0.35, the detector marks all research episodes under the base harness as EarlySubmit, producing an artificially large positive change in \(\mathrm{HMS}\). We use 0.25 because it avoids this saturation while preserving the qualitative process conclusions.
| Domain / measure | \(t=0.25\) | \(t=0.30\) | \(t=0.35\) |
|---|---|---|---|
| Knowledge-work | \(+0.059\) | \(+0.066\) | \(+0.066\) |
| Coding | \(+0.054\) | \(+0.054\) | \(+0.054\) |
| Research | \(\mathbf{-0.026}\) | \(+0.002\) | \(+0.096\) |
| Multi-tool | \(+0.058\) | \(+0.058\) | \(+0.058\) |
| Long-memory | \(+0.011\) | \(+0.002\) | \(+0.002\) |
| Planning | \(+0.010\) | \(+0.007\) | \(+0.007\) |
| Macro mean | \(\mathbf{+0.028}\) | \(\mathbf{+0.032}\) | \(\mathbf{+0.047}\) |
| Positive domains | 5/6 | 6/6 | 6/6 |
| Research ES base | 0.0% | 20.0% | 100.0% |
4pt
Figure 6 shows that the within-domain relationship between final quality and \(\mathrm{HMS}\) changes sign across domains. A negative correlation does not prevent process improvement under AW because AW reweights trajectory regions rather than fitting a linear relationship between \(G\) and \(\mathrm{HMS}\). This is consistent with Proposition 2: process maturity can improve even when final quality is flat or when the local correlation structure is mixed. Figure 7 gives the complementary per-setting view, showing that the adapter settings combine outcome and process gains while the controlled domains exhibit a weaker outcome response.
The results support a process-control interpretation of harness-level offline RL. A frozen LLM agent exposes an optimizable control layer distinct from the model’s language and reasoning capacity: AW reliably changes when the harness checks, revises, and submits, even when final task quality is nearly unchanged.
Offline AW reweights actions observed in the rollout buffer. Recurring control patterns, especially checking before submission, can therefore be recovered across many tasks. Final-quality gains additionally require trajectories whose actions lead to better terminal artifacts. This accounts for the larger gains in coding, \(\tau\)-bench retail, and DB-Bench, and for the controlled domains where verification increases without comparable final-score movement.
The process improvement is not a broad increase in all \(\mathrm{HMS}\) events. It is concentrated in CheckBeforeSubmit, with smaller support from TestBeforeSubmit in coding. Some events are sparse, saturated, or insensitive under the current detector. The localized shift identifies the concrete behavior the controller learned and prevents overclaiming that the agent became generally more mature across all process dimensions.
The \(\tau\)-bench retail and AgentBench DB-Bench results show that the learned control pattern can help outside the controlled suite. However, both evaluations use adapted scoring protocols and preserve only the benchmark components needed for Harness MDP control. They support external validity of the harness-control mechanism, not a benchmark-level claim about the upstream tasks.
We formulate the control layer around a frozen LLM agent as a Harness MDP and train a lightweight offline AW controller over structural harness actions. The learned controller consistently increases verification before submission, while final-quality gains are concentrated in coding and two adapted public-benchmark settings; the latter are adapter-level results rather than official upstream benchmark scores. This process–outcome gap reflects the finite-buffer setting: recurring verification behavior can be learned broadly, but outcome gains require stronger trajectories in the offline data. Future work should move beyond fixed-buffer support through online improvement or targeted data collection, adopt native benchmark protocols, and develop broader, better-calibrated measures of harness process maturity.
For any trajectory \(\tau=(s_0,a_0,\ldots,s_H)\), the cumulative shaped return defined in Equation ?? satisfies \[\begin{align} \sum_{t=0}^{H-1}r_t^{\Phi}(s_t,a_t,s_{t+1}) &=\sum_{t=0}^{H-1}r_t(s_t,a_t,s_{t+1})\\ &\quad+\sum_{t=0}^{H-1}\Phi_{t+1}(s_{t+1}) \\ &\quad-\sum_{t=0}^{H-1}\Phi_t(s_t). \end{align} \label{eq:proof-shaped-sum}\tag{1}\] The potential terms telescope as \[\begin{align} &\sum_{t=0}^{H-1}\left(\Phi_{t+1}(s_{t+1})-\Phi_t(s_t)\right)\\ &\qquad=\Phi_H(s_H)-\Phi_0(s_0). \end{align} \label{eq:proof-telescoping}\tag{2}\] By Definition 2, \(\Phi_H(s_H)=0\), while the unshaped cumulative reward equals \(\mathcal{Q}(s_H)\). Hence, \[\sum_{t=0}^{H-1}r_t^{\Phi}(s_t,a_t,s_{t+1})=\mathcal{Q}(s_H)-\Phi_0(s_0). \label{eq:proof-trajectory-equivalence}\tag{3}\] Taking expectation under the trajectory distribution induced by \(\pi\) gives \[\begin{align} J_{\Phi}(\pi) &=\mathbb{E}_{\tau\sim p_\pi}[\mathcal{Q}(s_H)]\\ &\quad-\mathbb{E}_{s_0\sim\rho_0}[\Phi_0(s_0)]. \end{align} \label{eq:proof-policy-equivalence}\tag{4}\] Let \(C_{\rho_0}=\mathbb{E}_{s_0\sim\rho_0}[\Phi_0(s_0)]\). By Definition 1, \[J_{\Phi}(\pi)=J_{\mathcal{Q}}(\pi)-C_{\rho_0}. \label{eq:proof-constant-shift}\tag{5}\] Because \(\rho_0\) and \(\Phi_0\) are fixed, \(C_{\rho_0}\) is independent of \(\pi\). Thus, for any policies \(\pi_1\) and \(\pi_2\), \[J_{\Phi}(\pi_1)\geq J_{\Phi}(\pi_2)\Longleftrightarrow J_{\mathcal{Q}}(\pi_1)\geq J_{\mathcal{Q}}(\pi_2). \label{eq:proof-ranking-equivalence}\tag{6}\] Therefore, \[\operatorname*{argmax}_{\pi}J_{\Phi}(\pi)=\operatorname*{argmax}_{\pi}J_{\mathcal{Q}}(\pi), \label{eq:proof-argmax-equivalence}\tag{7}\] which proves Theorem 1.
Consider a one-step Harness MDP with initial state \(s_0\), action set \(\mathcal{A}=\{\mathrm{\small submit},\mathrm{\small check}\}\), and terminal states \(s_{\mathrm{good}}\) and \(s_{\mathrm{bad}}\). Action submit deterministically reaches \(s_{\mathrm{good}}\), whereas check deterministically reaches \(s_{\mathrm{bad}}\). Let \(\mathcal{Q}(s_{\mathrm{good}})=1\) and \(\mathcal{Q}(s_{\mathrm{bad}})=0\). Under the task-rubric objective in Equation ?? , the unique optimal policy is \(\pi_{\mathrm{submit}}\).
Now define the action-pattern bonus \(b(a)=\lambda\mathbb{I}\{a=\mathrm{\small check}\}\) for \(\lambda>1\). The modified returns are \[\begin{align} J_{\mathrm{pattern}}(\pi_{\mathrm{submit}})&=1,\\ J_{\mathrm{pattern}}(\pi_{\mathrm{check}})&=\lambda. \end{align} \label{eq:counterexample-modified-returns}\tag{8}\] Since \(\lambda>1\), the modified objective prefers \(\pi_{\mathrm{check}}\), although it produces lower terminal task quality. Therefore, \[\begin{align} \operatorname*{argmax}_{\pi}J_{\mathcal{Q}}(\pi)&=\{\pi_{\mathrm{submit}}\},\\ \operatorname*{argmax}_{\pi}J_{\mathrm{pattern}}(\pi)&=\{\pi_{\mathrm{check}}\}. \end{align} \label{eq:counterexample-different-optima}\tag{9}\] Thus, an action-pattern reward need not preserve the task-optimal policy set.
Since \(q_{\mathrm{AW}}\) is a probability distribution supported on \(B\), \[\begin{align} \mathbb{E}_{q_{\mathrm{AW}}}[G] &=\sum_{\tau\in B}q_{\mathrm{AW}}(\tau)G(\tau)\\ &\leq\sum_{\tau\in B}q_{\mathrm{AW}}(\tau)G_B^\star\\ &=G_B^\star. \end{align} \label{eq:proof-outcome-ceiling}\tag{10}\] Subtracting \(\mathbb{E}_{\mu_B}[G]=\overline{G}_B\) yields \[\Delta\mathcal{Q}_B\leq G_B^\star-\overline{G}_B=\sigma_B. \label{eq:proof-buffer-slack}\tag{11}\]
By Equation ?? , \[\begin{align} \mathbb{E}_{q_{\mathrm{AW}}}[\Psi] &=\sum_{\tau\in B}\frac{\mu_B(\tau)w(\tau)}{Z_B}\Psi(\tau)\\ &=\frac{\mathbb{E}_{\mu_B}[w\Psi]}{\mathbb{E}_{\mu_B}[w]}. \end{align} \label{eq:proof-aw-process-expectation}\tag{12}\] It follows that \[\begin{align} \Delta\Psi_B &=\frac{\mathbb{E}_{\mu_B}[w\Psi]}{\mathbb{E}_{\mu_B}[w]}-\mathbb{E}_{\mu_B}[\Psi]\\ &=\frac{\operatorname{Cov}_{\mu_B}(w,\Psi)}{\mathbb{E}_{\mu_B}[w]}. \end{align} \label{eq:proof-process-identity}\tag{13}\] Since \(w(\tau)=\exp(A(\tau)/\beta)>0\) for every \(\tau\in B\), we have \(\mathbb{E}_{\mu_B}[w]>0\). Hence, \(\Delta\Psi_B\geq 0\) if and only if \(\operatorname{Cov}_{\mu_B}(w,\Psi)\geq 0\).
By the law of total covariance, \[\begin{align} \operatorname{Cov}_{\mu_B}(w,\Psi) &=\operatorname{Cov}_{\mu_B}\left(w,\mathbb{E}_{\mu_B}[\Psi\mid A]\right)\\ &\quad+\mathbb{E}_{\mu_B}\left[\operatorname{Cov}_{\mu_B}(w,\Psi\mid A)\right]. \end{align} \label{eq:proof-total-covariance}\tag{14}\] Because \(w=\exp(A/\beta)\) is deterministic conditional on \(A\), the second term is zero. Under the condition \(\mathbb{E}_{\mu_B}[\Psi\mid A]=h(A)\), \[\operatorname{Cov}_{\mu_B}(w,\Psi)=\operatorname{Cov}_{\mu_B}\left(\exp(A/\beta),h(A)\right). \label{eq:proof-monotone-covariance}\tag{15}\] Both \(\exp(A/\beta)\) and \(h(A)\) are non-decreasing functions of \(A\). Chebyshev’s covariance inequality therefore gives \(\operatorname{Cov}_{\mu_B}(w,\Psi)\geq 0\). Theorem 2 then implies \(\Delta\Psi_B\geq 0\).
Consider \(B=\{\tau_1,\tau_2,\tau_3\}\) with \(\mu_B(\tau_i)=1/3\). Let \[\begin{align} (G(\tau_1),G(\tau_2),G(\tau_3))&=(1,0.5,0),\\ (\Psi(\tau_1),\Psi(\tau_2),\Psi(\tau_3))&=(0,1,1). \end{align} \label{eq:counterexample-return-process}\tag{16}\] Then \(\operatorname{Cov}_{\mu_B}(G,\Psi)<0\). Choose \((A(\tau_1),A(\tau_2),A(\tau_3))=(0,2,1)\). Let \(u=\exp(1/\beta)>1\), so the weights are \((1,u^2,u)\). A direct calculation gives \[\operatorname{Cov}_{\mu_B}(w,\Psi)=\frac{u^2+u-2}{9}>0. \label{eq:counterexample-weight-covariance}\tag{17}\] Equation ?? therefore implies \(\Delta\Psi_B>0\) despite \(\operatorname{Cov}_{\mu_B}(G,\Psi)<0\).
Each controlled domain contains 80 training tasks and 20 held-out evaluation tasks. We evaluate each policy using three random seeds and three rollouts per held-out task, resulting in 180 evaluation episodes per domain. Offline AW is trained on trajectories collected from the base and exploratory harnesses using exponential advantage weights, \(\exp(A/\beta)\) [12]. Terminal training rewards are produced by the same rubric scorer used for evaluation. Under the calibrated structural verifier, none of the controlled domains reaches the terminal-score ceiling at \(G=1\). For each public-benchmark adapter, we use 16 training tasks and 20 held-out evaluation tasks, following the same three-seed, three-rollout protocol. The adapter tasks are drawn from \(\tau\)-bench retail and AgentBench DB-Bench [7], [8]; the adapters preserve their tool-interaction or database-reasoning structure while replacing the native evaluator with Harness MDP scoring.
Several detector extensions may improve the coverage and reliability of \(\mathrm{HMS}\). First, reducing the StopWhenSufficient threshold from 0.85 to 0.70 may reduce event sparsity. Second, RevisionAfterFailure could be replaced by a graded measure that combines verifier-detected failure with the magnitude of the subsequent revision. Third, EvidenceBeforeClaim could be evaluated using claim-level annotations rather than surface-form heuristics.
All reported results are computed from per-task quality scores, event-level \(\mathrm{HMS}\) measurements, and public-benchmark adapter evaluations. The task-generation procedure, train–evaluation splits, detector thresholds, verifier configurations, and calibration analyses are reported in the following sections.
The controlled suite contains six domains: knowledge-work, coding, research, multi-tool, long-memory, and planning. The domains are inspired by benchmark families for web and tool interaction, software repair, multi-hop question answering, and long-context memory [5], [14], [17], [43]. Each domain contains 100 human-annotated tasks with a fixed difficulty distribution of 20 easy, 60 standard, and 20 hard tasks. The train–evaluation split is stratified by difficulty, yielding 80 training tasks and 20 held-out evaluation tasks per domain. Each training set contains 16 easy, 48 standard, and 16 hard tasks, while each evaluation set contains 4 easy, 12 standard, and 4 hard tasks. To reduce overlap between training and evaluation, tasks vary in difficulty, structural form, distractor content, starter artifacts, and required output format. Task identifiers are disjoint between buffer collection and held-out evaluation.
Task prompts, source constraints, reference criteria, and rubrics were produced through human annotation and curation. Held-out tasks are not paraphrases of training tasks. Train and evaluation splits use disjoint task identifiers, template instantiations, entities, distractors, starter artifacts, and required output formats. A manual review pass checked split assignment, rubric coverage, and duplicate or near-duplicate prompts before rollout collection.
Final task quality is computed as a normalized weighted sum of criterion-level scores, \[G_{\mathrm{norm}} = \frac{\sum_c \mathrm{score}_c}{\sum_c \mathrm{max\_score}_c}.\] The scoring criteria are domain-specific. Knowledge-work evaluates answer correctness, evidence support, criterion coverage, and output format, with an additional source-conflict criterion for hard tasks. Coding evaluates unit-test correctness, parseability, safety constraints, and cost compliance. Research evaluates answer correctness, evidence support, citation format, and output format. Multi-tool evaluates answer correctness, tool-call validity, and structural format requirements. Long-memory evaluates answer correctness, source-session matching, and output format. Planning evaluates constraint satisfaction, criterion coverage, and plan format.
Verifier coverage differs across domains. Coding is the only fully objective verifier domain because its main correctness criterion is determined by executable unit tests. The other five domains combine structural checks with rubric-based scoring against human-annotated criteria. The scorer receives the task prompt, rubric criterion, and candidate output, but not the reference answer. Calibration mode is used only for scorer-validation diagnostics. The remaining limitations include sensitivity of natural-language criteria to scorer calibration, limited coverage of regex and overlap-based checks for paraphrased content, and the use of fixed search observations rather than a live search system in the research domain.
This section provides the detector-level diagnostics underlying the compact sensitivity analysis in Section 5. We distinguish between the original domain-specific verifier, which uses a strict deterministic coding rubric, and the calibrated structural verifier, which applies a common structural scoring rule across domains and uses an EarlySubmit threshold of 0.25.
Table 9 reports final-quality changes, process changes, and buffer slack under the calibrated structural verifier. Coding exhibits the largest gain in final quality, while research is the only domain with a negative change in \(\mathrm{HMS}\). The variation across domains supports the use of \(\sigma_D\) as a diagnostic of offline trajectory support, but not as a calibration-invariant predictor of improvement.
| Domain | \(\sigma_D\) | Base \(G\) | \(\Delta G\) | \(\Delta\HMS\) |
|---|---|---|---|---|
| Knowledge-work | 0.279 | 0.450 | \(+0.014\) | \(+0.059\) |
| Coding | 0.217 | 0.712 | \(\mathbf{+0.100}\) | \(+0.054\) |
| Research | 0.060 | 0.314 | \(-0.003\) | \(\mathbf{-0.026}\) |
| Multi-tool | 0.175 | 0.528 | \(-0.013\) | \(+0.058\) |
| Long-memory | 0.000 | 0.453 | \(-0.003\) | \(+0.011\) |
| Planning | 0.552 | 0.385 | \(+0.026\) | \(+0.010\) |
| Macro mean | – | – | \(+0.020\) | \(+0.028\) |
5pt
Under the calibrated EarlySubmit threshold of 0.25, AW increases EarlySubmit on research tasks from 0.0% to 25.0%, corresponding to 45 of 180 evaluation episodes. Although CheckBeforeSubmit increases by 13.9%, the larger EarlySubmit penalty produces a net \(\Delta\mathrm{HMS}\) of \(-0.026\). AW therefore increases verification while also inducing earlier submission, showing that these two behaviors need not improve together.
Table 10 reports the full threshold sweep. Increasing the threshold causes more episodes to be classified as EarlySubmit and can substantially change the apparent process improvement. The effect is most pronounced in research: at a threshold of 0.35, all base episodes are labeled as EarlySubmit, which mechanically inflates the estimated AW improvement. We therefore use 0.25 as the primary threshold because it avoids this saturation while preserving the qualitative process conclusions.
| Threshold \(0.25\) | Threshold \(0.30\) | Threshold \(0.35\) | ||||
|---|---|---|---|---|---|---|
| 2-3(lr)4-5(lr)6-7 Domain | \(\Delta\HMS\) | ES Base | \(\Delta\HMS\) | ES Base | \(\Delta\HMS\) | ES Base |
| Knowledge-work | \(+0.059\) | 1.7% | \(+0.066\) | 16.7% | \(+0.066\) | 16.7% |
| Coding | \(+0.054\) | 0.0% | \(+0.054\) | 0.0% | \(+0.054\) | 0.0% |
| Research | \(\mathbf{-0.026}\) | 0.0% | \(+0.002\) | 20.0% | \(+0.096\) | 100.0% |
| Multi-tool | \(+0.058\) | 0.0% | \(+0.058\) | 0.0% | \(+0.058\) | 0.0% |
| Long-memory | \(+0.011\) | 0.0% | \(+0.002\) | 20.0% | \(+0.002\) | 20.0% |
| Planning | \(+0.010\) | 51.1% | \(+0.007\) | 51.1% | \(+0.007\) | 51.1% |
| Macro mean | \(+0.028\) | – | \(+0.032\) | – | \(+0.047\) | – |
| Positive domains | 5/6 | – | 6/6 | – | 6/6 | – |
6pt
The original coding evaluation uses a strict deterministic rubric, under which the base harness achieves 0.929 and is close to the verifier ceiling. Under this evaluation, AW changes final quality by approximately \(-0.006\). The calibrated analysis replaces this coding-specific rubric with the same structural verifier used across the other domains, reducing the base score to 0.712 and yielding an AW score of 0.812. The difference between the two base scores reflects verifier calibration rather than a change in the policy. The reported coding gain is therefore measured relative to the calibrated structural baseline.
On AgentBench DB-Bench, the EarlySubmit rate increases from 8.3% under the base harness to 27.8% under AW. CheckBeforeSubmit simultaneously increases by 16.7%, and the aggregate process score remains positive with \(\Delta\mathrm{HMS}=+0.030\). As in research, AW increases both verification and submission speed; however, the remaining process components offset the EarlySubmit penalty. These cases indicate that verification and stopping behavior should be modeled as distinct control objectives.
On the 20-task held-out \(\tau\)-bench retail split, final quality increases from 0.337 to 0.519, with \(\Delta\mathrm{HMS}=+0.075\). This is lower than the earlier four-task held-out estimate by 4.3 points, but remains well above the five-point decision threshold, preserving the adapter-level transfer conclusion.