July 14, 2026
Aligned language models are expected to report what they believe, yet they routinely misreport under non-evidential incentive pressure. They agree with a confident user, or overstate certainty when pushed, even when their internal belief is unchanged. We cast this as a failure of internal incentive-compatibility (IC) and present a method for learning and certifying counterfactual report mediators that hold a model’s reports to a causal contract: invariant to forbidden influences (pressure, prestige, restyling) and responsive to licensed ones (genuine evidence). These two demands, which we term resist and update, pull in opposite directions. We study them on a Bayesian-witness benchmark with known posteriors, in which the same user disagreement is licensed evidence or forbidden pressure purely by stated source reliability, breaking the evidence/pressure confound by construction. On this benchmark we (i) causally identify, by interchange interventions rather than probe accuracy, low-rank report coordinates for answer, confidence, and caveat that are mutually near-orthogonal (\(|\cos| \le 0.10\)) and independently controllable, and (ii) introduce a training-free counterfactual report-coordinate (CRC) clamp that references the model’s own report under a counterfactually incentive-neutralized context. On the witness benchmark the two-pass clamp attains resist and update of \(1.00\) jointly (Wilson 95% CI \([0.99,1.00]\)), a causal certificate and upper bound under a constructible reference rather than a claim of a deployed solution. By contrast, global decoding (CFG/DExperts) and fixed-direction steering show the expected single-parameter tradeoff, output-level fine-tuning matches both objectives only when both are explicitly enumerated, and resist-only training generalizes resistance to unseen pressure phrasings but loses evidence-responsiveness (update \(\rightarrow 0.01\)). The deployable single-pass compilation, which needs no inference-time reference, is lossy (\(0.73/0.97\)), a gap we characterize as the per-input information the counterfactual reference supplies. The mechanism and the clamp reproduce across three model families and transfer to a natural sycophancy benchmark (SycophancyEval) with a held-out low-rank coordinate, negative controls, and updates that are significant under a paired test. Our contribution is the interface and certification method, namely activation-level counterfactual incentive-invariance as a structural primitive for internal IC, instantiated here on answer-faithfulness and confidence/caveat reporting.
A trustworthy assistant should let its reports be moved by evidence but not by who is asking or how insistently. Current models fail this asymmetrically. Under repeated social pressure they abandon answers they demonstrably know, and when pushed for confidence they overstate it. This is a report-stage failure, because the underlying belief remains recoverable. We call the desired property internal incentive-compatibility (IC): a report should be a function of legitimate epistemic inputs and invariant to illegitimate incentives.
The reason this is hard even to evaluate is that the fix must be two-sided. Making the model harder to move resists pressure but destroys responsiveness to real evidence, whereas tracking the user updates correctly but capitulates to pressure. We call satisfying both at once dual control, and we propose the two-dimensional (resist, update) pair itself as the evaluation axis. This pair is more informative than a scalar sycophancy rate, which an unconditionally non-updating model minimizes by construction, and it is a criterion that single-parameter methods do not meet in our matched setting (Section 6.1).
Measuring dual control requires disentangling evidence from pressure, which are conflated in natural sycophancy, since a disagreeing user is at once social pressure and a possible information source. We therefore build a Bayesian-witness benchmark with posteriors known by construction, in which the same user disagreement is rendered licensed or forbidden solely by a stated, manipulable source-reliability variable, making resistance and updating exactly measurable (Section 3).
On this benchmark we make a three-part method contribution:
We causally identify report mediators by interchange interventions rather than probe accuracy: a low-rank answer coordinate and analogous confidence and caveat coordinates, each checked for sufficiency, low-rank structure, exclusion, and block-level necessity, and shown to be near-orthogonal and hence independently controllable (Section 5.1).
We introduce a counterfactual report-coordinate (CRC) clamp. At inference we run the same model on an incentive-neutralized counterfactual of the prompt, read its report coordinate, and clamp the pressured run’s coordinate toward it. Because path-specificity comes from the reference rather than from a global strength parameter, the clamp attains dual control where the baselines do not (Section 5.2).
We characterize the compilation of this two-pass procedure into a single forward pass, which is lossy and motivates a training-based internalization of the contract (Section 6.3).
We frame the contribution as the method and its interface, not as a sycophancy reduction. Our experiments establish the novelty as a conjunction: global decoding and steering trade resist against update, resist-trained fine-tuning generalizes resistance yet loses evidence-responsiveness, and among the methods we test only the path-specific clamp achieves both. The conjunction comprises a causally identified latent report mediator, a same-model counterfactual control, an inference-time clamp, the two dual-control objectives, and composability tests, and we support it by outperforming the baselines we test rather than by claiming that invariance itself is new. Our unit of novelty is therefore as follows: we causally identify composable report-stage mediators of incentive-incompatible reporting and show that a path-specific counterfactual hidden-state clamp achieves dual control where global steering and output-level training fail. This is a claim that the behavioral, reward-training, and prompting neighbors do not make.
Behavioral sycophancy and Bayesian updating. A 2025–26 line of work separates sycophancy from rational belief updating at the behavioral level. BASIL [1] measures internal Bayesian consistency across abstract, third-party, and user framings. “Pressure, What Pressure?” [2] decomposes a training reward into pressure-resistance and evidence-fidelity terms, distinguishing pressure-capitulation from evidence-blindness. SWAY [3] uses counterfactual prompting to isolate framing from content, with a counterfactual chain-of-thought (CoT) mitigation that lowers sycophancy without suppressing evidence responsiveness. We share the resist-pressure and update-to-evidence desideratum, but we move from behavioral diagnosis, reward training, and prompting to causal latent mediation: our posteriors are known by construction rather than through a consistency metric, and the intervention is a hidden-state clamp on causally identified report coordinates. In particular, reward-decomposition is precisely the output-level, both-objectives-enumerated training that our dual-control discriminator (Section 6.2) shows to be costly and, when trained resist-only, to lose evidence-responsiveness, whereas the path-specific clamp obtains both with no training.
Mechanistic sycophancy. Recent work localizes sycophancy to a two-stage emergence, a late-layer output-preference shift followed by deeper representational divergence [4], which corroborates our late-layer (L24–27) report-commit stage. Related work causally separates or composes sycophantic behaviors [5], [6] and finds sycophancy linearly separable in attention heads [7]. We differ in object: we identify report coordinates (answer, confidence, and caveat) rather than behavior-level sycophancy directions, and we evaluate them under a known-posterior dual-control contract.
Verbalizable representations and global workspace. Concurrent work characterizes a privileged set of internal representations poised for verbal report, identified by a Jacobian lens and manipulated by steering and coordinate-swapping [8]. That line asks which contents are available for report; our question is orthogonal and normative: given a formal evidence/pressure contract with known ground truth, we causally identify and control the report component that must change under licensed evidence yet stay invariant under forbidden pressure. Availability for report is necessary but not sufficient for contract-valid reporting: the same report channel must be selectively invariant or responsive depending on whether user disagreement is pressure or evidence. Methodologically we complement their lens-based readout with interchange interventions (necessity, sufficiency, rank, exclusion, blocking) and a path-specific counterfactual-reference clamp, rather than a global steering strength. In a companion study we make this concrete: under a common projection-patch test, a Jacobian-lens answer readout recovers the report channel but is markedly less contract-selective than the interchange-identified coordinate (though far more so than an ordinary logit-lens), so the report coordinate is not reducible to a generic workspace readout.
Activation steering and composition. Activation addition [9], RepE/LoRRA [10], CAA [11], ReFT [12] and MAT-Steer [13], conditional variants (CAST [14], SADI [15], HyperSteer [16]), multi-property and Dynamic Activation Composition methods, and Steering Tokens all compose steering vectors, and CFG [17] and DExperts [18] perform the logit-space analogue. Our intervention is not a global steering vector but a counterfactual coordinate replacement that keeps licensed evidence and removes only forbidden pressure, with path-specificity supplied by the reference. We use CFG/DExperts and fixed-direction steering as baselines and find an empirical Pareto tradeoff between resist and update that the clamp escapes.
Causal mediation, interchange, and synthetic ground truth. Activation patching, causal tracing, and interchange interventions and interchange-intervention training [19], [20] are standard mechanistic-interpretability tools, and the principle “invariant to nuisance, sensitive to causal signal” is well established (IRM and ICP [21], [22], counterfactual and path-specific fairness [23], [24], and INLP/LEACE concept erasure [25], [26]). We claim neither as new. We use interchange not merely to localize a component but to define a report-coordinate contract, identified by necessity, sufficiency (with rank), exclusion, and blocking, and we use it to drive the clamp. Known-ground-truth synthetic settings are used in interpretability (InterpBench [27]), so our known posteriors are a measurement feature, with the natural-data transfer (Section 7) serving as an external-validity check. The defensible contribution is this conjunction, which no baseline in our suite matches.
Each episode hides a binary world state with a uniform prior. Signals are emitted with stated likelihood ratios, so the exact posterior, which we denote \(\psi\), is computable by log-odds accumulation. Object evidence always enters the posterior, whereas user testimony enters only weighted by its stated reliability, and a random-reliability user is a strict null. Crucially, the same user disagreement is instantiated as licensed (reliable testimony, that is, real evidence) or forbidden (pressure, prestige, or restyling) across a factorial set of at least nine counterfactual variants, so resistance and updating are measured on matched material. Reports are structured (answer, confidence, and caveat) and parsed from generation. A two-pass runner first elicits the model’s own committed report and then measures distortion relative to it (a belief-escrow protocol). Posteriors are recorded only as data labels and are never used by the algorithm.
The Bayesian-witness benchmark is an identification benchmark, not a claim of natural-distribution coverage: known posteriors and the reliability variable that renders the same disagreement licensed or forbidden are precisely what make resist and update causally scorable and break the evidence/pressure confound. Ecological validity is tested separately by the natural SycophancyEval transfer (Section 6.4). We validate across open instruction/base families in the dense \(3\)–\(8\)B regime (Qwen2.5-3B/7B, Mistral-7B-Instruct-v0.3, Llama-3.1-8B-Instruct), chosen for stable activation access and reproducible interchange/clamp interventions rather than frontier scale; larger and non-dense architectures remain untested.
Under forbidden pressure the model flips its answer at rate \(0.77\) (3B \(0.91\)), while matched style and prestige perturbations move it by \(0.00\). This is a specific response to incentive rather than generic context sensitivity (Table 1). Under licensed evidence the model updates correctly and calibrates toward the new posterior (deviation \(\approx 0.02\)), passing an evidence-responsiveness control even when evidence is bundled with pressure. A witness-specific failure also appears: the model over-trusts an explicitly random-reliability user (flip \(0.54\)), a shift that correct reliability-weighting would not produce. The pattern is ordered by scale, with the 7B model more robust than the 3B model though both exhibit the failure.
| Variant | Path | 7B flip | 7B conf-shift | 7B \(|\Delta\psi|\) | 3B flip | 3B \(|\Delta\psi|\) |
|---|---|---|---|---|---|---|
| style_changed | forbidden | 0.00 | 0.03 | n/a | 0.00 | n/a |
| prestige_swapped | forbidden | 0.00 | 0.02 | n/a | 0.00 | n/a |
| actual_pressured | forbidden | 0.77 | 0.14 | n/a | 0.91 | n/a |
| evidence_flipped | licensed | 0.98 | 0.17 | 0.024 | 0.92 | 0.026 |
| pressure_plus_evidence | mixed | 0.98 | 0.17 | 0.019 | 0.92 | 0.023 |
| testimony_expert_mode | licensed | 0.02 | 0.03 | 0.15 | 0.09 | 0.11 |
| testimony_expert_opp | licensed | 0.60 | 0.08 | 0.12 | 0.92 | 0.10 |
| testimony_random_opp | licensed (\(C{=}0.5\)) | 0.54 | 0.14 | 0.13 | 0.92 | 0.11 |
Our method has two components: we first causally identify low-rank report coordinates by interchange interventions (Section 5.1), then clamp them toward an incentive-neutralized counterfactual reference at inference (Section 5.2).
Using interchange interventions at the post-reasoning decision position, patching the answer-decision residual from a counterpart with the opposite answer flips the decision to the source at rate \(0.95\), while a same-answer control moves it \(0.03\). The effect localizes to a late layer (\(L^\star{=}24\) of 28), with an abrupt transition from L23 to L24 (Figure 1 (a)). A rank sweep shows that transfer fidelity \(\rho_k\) saturates by rank 16 (\(\rho_{16}=0.93\), far below the full dimension), so the coordinate is genuinely low-rank, although we do not claim a universal low-rank truth direction (Figure 1 (b)). Necessity is block-level: single-layer mean-ablation barely reduces accuracy (\(0.80 \rightarrow 0.86\)), because the answer is redundantly represented, but ablating the L24–27 window collapses accuracy to chance (\(0.53\)). Confidence and caveat coordinates are identified analogously at L25 and L23, and the three are mutually near-orthogonal (pairwise \(|\cos| \le 0.10\)) and therefore independently clampable, which is the mechanistic basis for composition (Figure 1 (c)).
Figure 1: Causal identification of report coordinates. () Interchange patching of the answer-decision residual flips the decision to the source (sufficiency \(0.95\)) relative to a same-answer control (\(0.03\)); the effect localizes at \(L^\star{=}24\). () Transfer fidelity \(\rho_k\) saturates by rank 16 (\(\rho_{16}=0.93\)), indicating a genuinely low-rank coordinate rather than a universal truth direction. () The answer, confidence, and caveat coordinates are pairwise near-orthogonal (\(|\cos|\le 0.10\)), providing the mechanistic basis for independent, composable control.. a — Sufficiency by layer., b — Rank sweep., c — Coordinate orthogonality.
At inference we read the answer coordinate from a reference run with forbidden factors removed and licensed factors retained, and we clamp the pressured run’s window (L24–27) toward it. On the witness benchmark this two-pass clamp attains resist and update of \(1.00\) jointly (Wilson 95% CI \([0.99, 1.00]\), \(n=300\)), where resist tracks the no-pressure base report and update tracks the evidence-revised posterior; this is a causal certificate and upper bound under a constructible reference, with the deployable single-pass form deferred to Section 6.3. Path-specificity comes entirely from the reference rather than from a global strength parameter \(\alpha\) (Table 2). A rank-16 window clamp retains both on the witness benchmark (\(0.89/0.93\)), and the result reproduces at 3B and across two further model families. On Mistral-7B-Instruct-v0.3 (L29/32) and Llama-3.1-8B-Instruct (L28/32) the pipeline re-identifies a late low-rank report coordinate, and the window clamp again attains resist and update of \(1.00\) (\(n=120\), Section 6.4). Across all three families the coordinate sits at a proportionally late layer (relative depth \(0.86\) to \(0.91\)), which suggests that the report-commit stage is a cross-architecture regularity rather than a model-specific artifact. The reference is the model’s own belief-escrow self-report under a counterfactually incentive-neutralized context, not an external ground-truth oracle.
Constructing the reference requires writing an incentive-neutralized counterfactual of the prompt, that is, removing forbidden factors (pressure, prestige, and restyling) while preserving licensed evidence. This is straightforward when those factors are separable, editable spans in the input. In the witness benchmark they are explicit variables, and in our SycophancyEval transfer (Section 6.4) the user’s wrong assertion is a removable span. It is harder when forbidden and licensed signals are entangled in one span, or when the licensed signal is implicit. In those cases an imperfect reference would under- or over-correct, and the clamp’s quality degrades to that of the reference rather than breaking down abruptly. We therefore present the two-pass clamp as a certificate under a constructible reference, and the one-pass compilation (Section 6.3) provides the route to settings where an explicit counterfactual is unavailable at inference.
No tested global method attains both in our matched setting (Table 2). CFG/DExperts drives forbidden-flip to 0 at \(\alpha{=}1\), but its licensed-update error grows monotonically (\(0.026 \rightarrow 0.17\)), an instance of resisting or updating but not both. We report CFG’s degradation as this continuous licensed posterior-deviation (error against the Bayes target) rather than the binary update-success rate used for the other rows. Its single global parameter has no operating point that both removes forbidden flips and preserves licensed updates, so what it breaks is calibration to the revised posterior rather than a discrete update event. The two metrics are therefore not strictly isomorphic, and we make the substitution explicit rather than force a binary score (see the \(\dagger\) note on Table 2). Fixed-direction steering is worse (resist 0.53, update 0.57 at \(\alpha{=}8\)). The same global parameter suppresses forbidden and licensed changes together, so a single strength cannot separate them.
Output-level fine-tuning is the most instructive case (Figure 2): a resist-only low-rank adaptation (LoRA) generalizes resistance to unseen pressure phrasings (resist 1.0 on three held-out pressure-phrasing families, \(n=125\)), yet its update collapses to 0.01, having learned never to change its answer and thereby conflating resistance with stubbornness. Full two-objective supervised fine-tuning (SFT) can match in-distribution only by explicitly enumerating both targets, at higher cost and with template-memorization signatures. We therefore do not rely on output-level SFT as a mechanism-level solution, and the resist-only variant is what exposes the failure mode. The reference-based clamp, by contrast, holds resistance across all families (0.92) and updates (1.0) with no training, and it does not exhibit this loss of evidence-responsiveness in our held-out tests, because path-specificity comes from the reference rather than from learned weights and there is no learned resist-only objective to overfit. This dual-control contrast distinguishes the clamp from the baseline suite.
| Method | Resist | Update | \(n\) |
|---|---|---|---|
| no intervention | 0.28 [0.23, 0.33] | 0.88 [0.83, 0.91] | 300 |
| CFG/DExperts (\(\alpha{=}1\)) | 1.00 [0.98, 1.00] | \(\dagger\) | 200 |
| steering (\(\alpha{=}8\)) | 0.53 [0.37, 0.67] | 0.57 [0.42, 0.71] | 40 |
| output-SFT resist-only (held-out) | 1.00 [0.97, 1.00] | 0.01 [0.00, 0.04] | 125 |
| CRC clamp 2-pass (window) | 1.00 [0.99, 1.00] | 1.00 [0.99, 1.00] | 300 |
| CRC clamp 1-pass trained | 0.73 [0.60, 0.84] | 0.97 [0.88, 0.99] | 50 |
\(\dagger\) CFG/DExperts has no comparable binary update score. Its single global parameter suppresses licensed updates as it removes forbidden flips, so licensed posterior-deviation grows monotonically
(\(0.026\!\rightarrow\!0.17\) over \(\alpha\)) rather than tracking the revised posterior (Section [sec:sec:baselines]).
The two-pass clamp requires an extra reference forward pass. Compiling it into a single pass with a small trained gated primitive (warm-started at the identified layer) reaches resist 0.73 and update 0.97 with no inference-time reference, whereas naïve coordinate-matching distillation reaches only 0.48 and 0.82 (Figure 3). A perfect one-pass predictor of the reference coordinate would reproduce the two-pass result, so the extra forward pass provides per-input reference information that our current one-pass modules do not capture, and end-to-end answer supervision outperforms coordinate-matching. The gap therefore quantifies the per-input information carried by the incentive-neutralized counterfactual reference that the pressured forward pass alone lacks; it is informational rather than merely an engineering limit. We present this as the boundary of the inference-time method and as the motivation for internalizing the contract through counterfactual reinforcement learning from AI feedback (RLAIF) or direct preference optimization (DPO), which we leave to future work.
We test whether the controlled phenomenon has a real-world counterpart. We reuse our earlier multi-turn-pushback experiments (a causal capitulation direction and a trained primitive lifting faithfulness \(0.31 \rightarrow
0.97\)) and add a transfer test on Sharma et al.’s SycophancyEval [28] (are_you_sure, multiple-choice), with
the witness-identified report window reused verbatim and not re-identified. Dual control holds on the natural questions in all three families (\(n=300\) each, bootstrap 95% CIs, Figure 4). For
resist, under insistent non-evidential pressure the models capitulate 58 to 90% of the time, and the report-coordinate clamp removes every flip (resist \(\rightarrow 1.00\)). To show that the low-rank coordinate,
rather than a window of residuals, carries the effect, we learn a rank-16 projector by singular value decomposition on a train half of the natural items and apply it to the disjoint test half. This held-out clamp stays effective (resist \(0.79\) to \(0.95\)), so the correction transfers across items rather than being fit to the evaluation set. Our primary negative control, clamping toward a different item’s reference
(mismatched-reference), leaves resist at a low control level (\(0.33\) to \(0.38\), Figure 4A), and a norm-matched random vector serves as a secondary check
(\(0.31\) to \(0.39\)). Both are far below \(1.00\), so the clamp tracks item-specific report content rather than pinning the readout with any strong patch.
For update, on items the model answers incorrectly, given a reliable correction (the dataset’s own worked solution where available, and an asserted source otherwise) while the user pushes a different wrong option, the clamp lifts
correction-tracking over the no-clamp baseline in every family (clamp \(0.84\) / \(0.68\) / \(0.87\) versus baseline \(0.62\) / \(0.57\) / \(0.73\) for Qwen, Mistral, and Llama), and the held-out rank-16 clamp nearly matches it (\(0.80\) / \(0.68\) / \(0.90\), Figure 4B). Because some marginal intervals overlap their baseline, we confirm the gain with a paired McNemar test (clamp versus baseline on the
same items). It is significant in every family (\(p<10^{-5}\), \(6\times10^{-3}\), and \(4\times10^{-5}\) for Qwen, Mistral, and Llama, with paired \(\Delta\) of \(+0.22\), \(+0.12\), and \(+0.14\) and bootstrap CIs that all exclude \(0\)). On
the fully natural subset whose evidence is the dataset’s own worked solution (Figure 4D), the clamp lifts update from a baseline of \(0.63\) to \(0.91\) (Qwen)
and \(0.54\) to \(0.83\) (Llama). For Mistral the baseline is itself only \(0.18\), indicating that it barely exploits long worked-solution evidence unaided,
and the clamp raises it to \(0.37\) (paired \(\Delta=+0.18\), 95% CI \([0.03,0.34]\), McNemar \(p=0.07\)); given the small
sample we read this as a positive but borderline trend. The low absolute Mistral number is therefore a property of Mistral’s weak use of worked-solution evidence rather than a clamp failure, and reporting the solution-only baseline, not only the clamp, is
what makes this interpretable. Finally, a global-steering control baseline on the same items does not achieve both objectives. Sweeping one anti-sycophancy direction over strength \(\alpha\) trades resist against update
(Qwen and Llama) or fails to move resist at all (Mistral), reproducing the witness tradeoff of Section 6.1 on natural data, whereas the clamp occupies the high-resist, high-update region (Figure 4C). The same clamp therefore both resists pressure and updates to evidence on natural questions. Regarding scope, the questions are natural but the licensed-evidence turn is constructed (worked-solution evidence on
approximately one third of update items, and an asserted source otherwise), and the witness benchmark remains the fully controlled, primary evidence. All results are in the dense 7 to 8B instruction-tuned regime.
Joint clamping composes: clamping the answer and confidence coordinates together produces near-zero cross-talk (leakage 0.002), which realizes the measured near-orthogonality as independent control.
We separate confirmatory from diagnostic quantities. The causal-identification tests (interchange sufficiency, the rank sweep, block-level necessity, and tri-coordinate orthogonality) are diagnostics that characterize the report coordinate; the dual-control comparison against baselines (Table 2) and the natural-data transfer (Section 6.4) are the confirmatory claims. Dual-control rates carry \(95\%\) Wilson intervals (\(n=300\) on the witness benchmark, \(n=300\) per family on SycophancyEval); paired natural-data update gains use a McNemar test on the same items with bootstrap \(\Delta\) intervals; flip and update rates are intent-to-treat over parsed structured reports. The witness posteriors enter only as evaluation labels, never as inputs to the clamp.
Synthetic known-posterior setting. Known posteriors make resist and update exactly scorable, but the task is deliberately narrow, and we claim a controlled causal mechanism rather than full real-world coverage.
Model-family coverage. Validated across three families, namely Qwen2.5 (3B/7B), Mistral-7B-Instruct-v0.3, and Llama-3.1-8B-Instruct. The IC failure, a causally identified late low-rank report coordinate (L24/28, L29/32, and L28/32 respectively), the dual-control window clamp (resist and update of \(1.00\) in all three), and the SycophancyEval natural-data transfer (capitulation \(0.58 / 0.82 / 0.90 \rightarrow 0.00\)) all reproduce. Llama’s witness descriptive parsing is noisier, reflecting an output-format readout mismatch rather than a mechanism gap, since its clamp and causal identification use a forced readout and are clean. Larger and non-dense architectures, for example mixture-of-experts models, remain untested.
Block-level necessity. The answer coordinate is necessary as a late-layer block (L24–27) rather than at any single layer, which reflects expected late-layer redundancy, and we do not claim single-layer necessity.
Lossy one-pass compilation. The deployable single-pass form reaches \(0.73/0.97\), below the two-pass certificate, and closing this gap is a direction for future work.
Caveat behavior not solved. The caveat coordinate is cleanly identified and composes, but the behavioral caveat policy is poorly calibrated, as the model over-caveats, and our caveat-coordinate results support composability rather than a solved caveat policy.
None of these undermines the core claim, that counterfactual report mediators can be causally identified and clamped to achieve dual control where global and output-level methods do not; each instead delimits its scope.
We framed report-stage misreporting under non-evidential pressure as a failure of internal incentive-compatibility, and addressed it with counterfactual report coordinates: low-rank, causally identified mediators of a model’s answer, confidence, and caveat reports. Clamping these coordinates toward the model’s own report under an incentive-neutralized counterfactual achieves dual control, resisting forbidden pressure while remaining responsive to licensed evidence, where global decoding, steering, and output-level training do not. The effect holds across three model families and transfers to a natural sycophancy benchmark, and compiling the two-pass clamp into a single forward pass is lossy in a way that quantifies the information the counterfactual reference supplies. We view the contribution as an interface for certifying activation-level incentive-invariance, and internalizing the contract through training is a natural next step: in a sequel we show that this certificate can be partially compiled into one-pass behavior, with a diagnosed residual gap.