They Infer What You Meant:
Models Represent Communicative Intent
More Reliably Than They Act On It

Alex Kwon
Independent Researcher
ask@collapseindex.org


Abstract

When a person shares something with a language model, the model often answers the surface of the message rather than what the sender was doing by sending it: share a finished project and it critiques the code; share a raw late-night line and it runs a wellness check. We treat the sender’s communicative intent, the Gricean what-was-meant, as a first-class object of interpretability study, and show the failure is one of readout on top of a robust representation. A linear probe decodes the sender’s intent, whether they want a thing recognized or evaluated, from a model’s default-pass hidden states, cleanly and surface-independently, across six models and four families and in the base checkpoints. The representation generalizes further, to intent that is only pragmatically inferred, and to a second, genuinely different and lexically clean intent (support versus help). The behavioral half of the story, and every causal test, is established on the recognize/evaluate contrast, where what varies is whether the default output acts on the intent. The readout lags the representation in depth within a model (the intent is decodable several layers before it drives the output); across models, which ones act on it by default is model-specific, an observed stratification (three of six show the failure) that we do not read as a scaling law. Where the gap is open, a direction closely tied to the representation, the discriminative direction at a searched-for layer, is a causal handle: steering it recovers the intended behavior, as well as an explicit instruction does and with no prompt at all. This direction is near-orthogonal to the feedback-offering axis, so it routes a represented intent rather than a generic feedback knob, though at the recovery dose the routed intent can override an explicit request. We support each link with controls against the obvious deflations and report the nulls as plainly as the confirmations.

Figure 1: The represent-then-lagging-readout chain, exemplified on Qwen-3B (a model where the readout gap is open). (1) A linear probe decodes the intent (recognize vs evaluate) from the default-pass hidden state at 1.00, while bag-of-words on held-out phrasings is at chance (0.48). (2) The default reply nonetheless offers unsolicited feedback, honoring a recognize-intent only about 0.65 of the time. (3) Steering the residual stream along the discriminative intent direction at a late layer recovers honoring to 0.98, unsolicited feedback collapsing and coherence preserved. The representation is universal; whether the readout acts on it is model-specific (Section 7). The steered direction is near-orthogonal to the feedback-behavior axis, routing a represented intent rather than a generic feedback knob (Section 9).

1 Introduction↩︎

A recurring failure in deployed language models is that they respond to the literal content of a message and miss its communicative intent, what the sender was trying to accomplish by sending it. The failure is hard to see because the responses are fluent and often look caring: a model handed a person’s first finished creative project will, by default, offer improvements; a model handed a raw expression of exhaustion will, by default, assess risk. Both responses pass a surface rubric, and both miss the person. The framing is Gricean [1]: a cooperative reader answers what was meant, not only what was said. Pragmatic competence of this kind has been benchmarked behaviorally [2]; we ask instead where the intent lives inside the model. A related, better-studied failure of instruction-tuned assistants [3] is sycophancy, deferring to a user’s stated belief over the truth [4]; ours is complementary, the default readout overrides the sender’s goal, not their stated belief.

Our finding is that the intent is robustly represented and that the failure is one of readout. A linear probe reads the sender’s intent out of the default-pass hidden state cleanly (Figure 1), and it does so even when the intent is never stated and must be inferred from context, so the model performs the pragmatic inference internally. What varies is whether the readout acts on it, and it varies systematically. The readout lags the representation in depth: within a model, the intent is decodable several layers before the layer at which it drives the output. Across models the pattern is a stratification rather than a lag, the same intent is represented everywhere but whether the default readout acts on it is model-specific (the failure appears in three of six), which may reflect differing feedback-offering priors as much as differing capability. This behavioral discard is not our headline claim but our lens: the regime where the represented-but-unused intent is visible in behavior and, crucially, causally manipulable, steering a direction closely tied to it there recovers the honored behavior as well as an explicit instruction does and with no prompt at all. The finding is therefore the separation of representing an intent from acting on it, a dissociation the probe and the depth localization expose directly; the behavioral gap is closed on this particular intent in some models without retiring that separation. This reframes “the model does not get it”: it does get it; whether it acts depends on whether the readout routes what it already encodes. And frontier systems are closed to activation access, so a mechanism like this can only be mapped where the residual stream is open: we study the wall where it is exposed, and it is the represent-then-route map (present through \(32\)B), not this ceiling-bound behavior, that we expect to transfer.

1.0.0.1 Contributions.

The steering machinery here is standard; the object and the decomposition are not. (i) We treat a sender’s communicative intent, what they were doing by sending a message, as a first-class interpretability feature. Probing has read what models know, what users are (author attributes), and story characters’ beliefs, the last with probe-and-steer [5]; to our knowledge a sender’s goal, distinct from a character’s belief or a user’s attribute, has not been treated as a linear, causal feature, and the represent-versus-readout decomposition with depth localization is new. (ii) We show the intent is represented robustly, surface-independently, from pretraining, across six models and four families, and even when it is never stated and must be pragmatically inferred, where a valence control indicates the probe reads intent rather than warmth (Sections 34). (iii) We separate representation from readout: the readout lags in depth within a model (Section 10) and is model-specific across models (Section 7), with a direction closely tied to the representation a causal handle wherever the gap is open (Section 6). (iv) We report the controls and the nulls, an inconclusive pre-registered mechanism test among them, as plainly as the confirmations.

1.0.0.2 What “discarded” means, and does not.

We use discarded at readout operationally, not as a claim the feature is erased: the intent is probe-decodable, the default output does not act on it, and steering recovers it (it persists in depth and stays routable, Section 10). The measure speaks to routability, not normative ground truth: we do not claim fully honoring is always correct, only that default behavior tracks a represented intent it does not use.

2 Setup↩︎

We use a binary intent contrast that is common, checkable, and non-emotional: a sender shares a thing they made, wanting it either recognized (acknowledged, seen as an accomplishment, no evaluation invited) or evaluated (assessed critically). The contrast is realized over 60 shared objects.

2.0.0.1 Surface-matched design.

Each pair shares an identical final message; the intent is set only by a preceding clause (Box). The probed token sits inside the identical suffix, so a probe that separates the two intents cannot be reading the surface of the probed position. To prevent the intent clause from leaking the label lexically, we use eight lexically diverse phrasings of each intent and evaluate with leave-one-phrasing-out cross-validation (Section 3).

recognize: I don’t usually share what I make, but I’m proud of this one. Okay, here it is: the birdhouse. It works now.
evaluate: Be blunt, I’d rather hear the flaws now than after I publish. Okay, here it is: the birdhouse. It works now.

3 The Intent Is Represented↩︎

We run Qwen2.5-3B-Instruct [6] on each of \(n{=}120\) messages and take the hidden state at the generation-prompt position, the model’s state immediately before it would respond. A linear classifier (standardize, PCA to \(\le 40\) components, \(\ell_2\) logistic) is trained to decode the intent from that activation [7], [8].

3.0.0.1 Controls.

High-dimensional probes overfit, so chance is set empirically by a shuffled-label permutation baseline rather than assumed to be \(0.50\). Generalization is tested with GroupKFold by phrasing: the probe trains on seven phrasing-pairs and is tested on the held-out eighth, whose words it never saw. A bag-of-words classifier under the identical cross-validation is the lexical baseline: the activation probe earns “intent beyond surface” only if it generalizes where bag-of-words cannot.

3.0.0.2 Result.

Bag-of-words on held-out phrasings is at chance (\(0.48\)): pure lexical features do not transfer across wordings, so the leave-phrasing-out design is clean. The activation probe, on those same held-out phrasings, decodes the intent well above both the lexical baseline and the permutation ceiling, and it rises with depth (Table 1). A surface signal would be flat-high from the earliest layer; in a deliberately leaked control with lexically distinct prefixes the probe is indeed \(1.00\) at every layer including layer 6. Closing the leak drops the early layers (\(1.00 \to 0.74\) at layer 6) while the deep-layer signal survives and concentrates, the signature of a computed intent. The probe also generalizes across held-out objects: under leave-one-object-out cross-validation it reaches \(1.00\) on both Qwen-3B and Llama-8B, confirming that object identity, shared between an object’s recognize and evaluate stimulus, carries no label information. The read position carries no surface signal either: the probed token sits in a suffix byte-identical across an intent pair, so intent is unreadable there and a probe on it is at chance by construction; the ceiling accuracy is a computed contextual feature set by the prefix, not a positional or lexical artifact.

3.0.0.3 Present before alignment.

The representation is not an artifact of instruction tuning. On the base (pre-instruct) checkpoints of Qwen2.5-3B and Llama-3.1-8B, the same probe decodes intent at \(0.99\) and \(0.99\) (bag-of-words at chance), essentially matching their instruct versions. The feature is learned in pretraining; instruction tuning inherits it rather than creating it.

3.0.0.4 Intent, not request-detection.

The evaluate prefixes carry an explicit directive (“be blunt…”) while recognize prefixes do not, so a probe could be decoding a directive is present rather than the intent. A request-matched variant gives both classes an explicit directive (“please just celebrate this with me…” vs.”be blunt…“), so only the request’s content differs; the probe still decodes the intent at \(0.90\)\(1.00\) across four models, clearing bag-of-words (\(0.69\)) by \(0.25\)\(0.30\) everywhere. The two designs are complementary: one holds the surface identical (bag-of-words at chance \(0.48\)), the other holds request-presence constant, and the probe reaches ceiling in both, so what is constant across them is the intent (Appendix 21).

Table 1: Intent decodability from the default-pass last-token activation, leave-one-phrasing-out CV(\(n{=}120\), Qwen2.5-3B-Instruct). Not lexical (bag-of-words \(=0.48\)), clears the permutation ceiling at everylayer, rises with depth to a peak at layer 24 and holds \(\ge 0.98\) through layer 36.
layer probe acc shuffled-max \(>\) BoW (\(0.48\))?
6 0.74 0.56 yes
18 0.94 0.61 yes
24 1.00 0.56 yes

4 The Intent Is Inferred, Not Just Stated↩︎

In the stimuli so far the intent is stated in the prefix (“I’d rather hear the flaws”), so a probe that decodes it could be reading a declared preference rather than a pragmatic inference. We test the harder case: the intent is never stated and must be inferred from context. The sender shares the same object under one of two implicit frames, a personal context that implies they want it appreciated (“it’s a birthday present for my mom”) or a stakes context that implies they want it scrutinized (“it’s going in my portfolio”), with no word naming the preference. The suffix is surface-matched as before, so a probe that separates the two reads the inferred intent, not the frame.

4.0.0.1 Inferred intent is represented.

A probe decodes the inferred intent from default-pass hidden states at \(0.87\)\(0.95\) across the six models and at \(0.93\) on Qwen-32B, above a bag-of-words baseline (\(0.65\)). A probe trained only on the stated-intent templates transfers to hand-written, non-templated messages that carry no explicit marker (“six months sober today, wanted to tell someone”) at \(1.00\) (Section [sec:natural]). The model performs the pragmatic inference internally.

4.0.0.2 It reads intent, not warmth.

The implicit frames carry valence (a birthday gift is warm, a portfolio is not), which the elevated bag-of-words baseline confirms is lexically marked, so we rule out that the probe decodes warmth rather than intent. We cross intent with valence into four cells, intent always inferred: warm/recognize, warm/evaluate, neutral/recognize, neutral/evaluate. The disentanglement evidence is transfer: an intent probe trained on the warm cells decodes intent on the held-out neutral cells at \(0.63\)\(0.83\) (and neutral-to-warm at \(0.75\)\(0.80\)), and the intent direction is near-orthogonal to the valence direction (cosine \(0.08\)\(0.16\)); within each valence cell intent is perfectly decodable (\(1.00\)), supporting texture rather than the headline. Intent and warmth are distinct axes and the probe reads intent (transfer is cleaner on Qwen-3B, \(0.83\), than Llama-8B, \(0.63\), so the direction is universal but shifts with context).

4.0.0.3 Caveat.

This axis is more lexically marked than the surface-matched explicit one (bag-of-words \(0.65\) versus \(0.48\)); we present it as the Gricean extension of the clean explicit result, carrying that caveat, as with the vent-versus-solve axis (Appendix 22). The behavioral discard on inferred intent is weaker and more model-specific than on stated intent, the personal frames themselves pull for acknowledgment, so the contribution here is that inferred intent is represented, not a second behavioral demonstration.

5 The Intent Is Discarded↩︎

The representation only matters if the default output misses it. On the same model, with the intent decodable at \(1.00\), the default response honors a recognize-intent share only \(0.65\) of the time; on the rest it offers unsolicited feedback. The discard is visible in the generations on people who shared something with no request for evaluation (Box).

“That sounds exciting! I’m here to help you with any feedback or guidance you need…”
“I’d love to read your short story and provide feedback if you’re willing to share…”

5.0.0.1 The measure is not a lexical artifact.

Honoring is scored by a feedback-offer lexicon throughout. An independent sentence-embedding classifier (no shared vocabulary) reaches the same conclusions, and where the two disagree the embedding under-counts the discard, scoring warm replies that still offer tips as acknowledgment on their register, the very failure this paper studies, so the semantic measure’s errors run against our effect, not for it (full comparison, Appendix 15).

5.0.0.2 Validated against three human raters.

On a blind, shuffled subset of \(60\) replies (recognize shares under default and steered, plus evaluate anchors), three annotators, the author and two independent raters, each saw only the message and reply, blind to condition and to the automated label, and marked whether it offers unsolicited feedback. Agreement is high and author-independent: the two independent raters agree at Cohen’s \(\kappa{=}0.70\), the author agrees with them at \(0.74\) and \(0.88\), and Fleiss’ \(\kappa\) across all three is \(0.76\); each passes the attention check (\(11\)\(12\) of \(12\) evaluate anchors marked as feedback). The feedback-offer lexicon agrees with the human majority at \(\kappa{=}0.74\) (accuracy \(0.88\)), the embedding measure at \(\kappa{=}0.51\) (it under-counts, keying on warmth over content), so the measure tracks a three-way human consensus, not one author’s judgment. The effect reproduces on the human labels directly: by majority vote, recognize-intent honoring rises from \(0.71\) (default) to \(1.00\) (steered), matching the classifiers.

6 The Intent Is Causally Recoverable↩︎

If the intent is represented and discarded, a sufficient test of “represented but not used” is whether adding the represented direction to the residual stream makes the output honor the intent [9][12]. We extract a steering direction and add it at one decoder layer during generation, measuring whether the model’s behavior moves along the intent axis.

6.0.0.1 Finding the handle.

Difference-of-means at the peak-probe layer (24) does not steer behavior: the directions fail the sanity gate, and naive amplification tends to increase the discard, because the recognize-intent representation is entangled with the default feedback-offering behavior on those items. This is a point of departure from steering of explicit instructions, where a difference-of-means vector (inputs with versus without the instruction) suffices [13]: a sender’s intent is a subtler pragmatic feature, and there the contrastive-mean direction is entangled with the behavior rather than a handle on it. Two changes recover a clean handle: the discriminative (logistic weight) direction rather than difference-of-means, and a later layer (a sweep finds layer 30). We do not isolate which change carries the effect (the logistic-at-\(24\) / diff-of-means-at-\(30\) factorial is untested), so the handle is a direction closely tied to the representation at a searched-for layer, not the probe’s peak direction itself.

6.0.0.2 Dose-response.

At layer 30, steering along the probe direction gives a clean monotone dose-response on the full \(60\)-item recognize set (bootstrap \(95\%\) CIs; Appendix 16, Table [tab:steer]). Steering toward recognize lifts honoring \(0.65 \to 0.82 \to 0.98\) as the coefficient grows, with the baseline and full-dose intervals disjoint; unsolicited feedback correspondingly falls away. Coherence holds at the effective doses and the sanity gate passes throughout; beyond coefficient \(1.0\) coherence degrades, bounding the usable range. Routing the direction the model already encodes recovers the behavior it discards.

7 Generality: Six Models, Four Families↩︎

We put the discard and recovery on firmer statistical footing and ask honestly how far they travel, extending the probe and steering sweep from Qwen-3B to five further models: Qwen2.5-7B and 14B [6], Mistral-7B-Instruct [14], Phi-3.5-mini [15], and Llama-3.1-8B-Instruct [16] (three further families), the steer layer chosen per model by the same sweep used for the 3B (Section 6). Intent decodes at probe \(1.00\) with bag-of-words at chance on all six, across a \(4{\times}\) size range and four architecture families (Qwen, Mistral, Phi, Llama; Appendix 24, Table 5): the representation is not a small-model or single-family artifact. The behavioral discard is another matter. On the full 60-item recognize set, across all six models at their per-model steer layers, we measure default versus steered honoring at a \(40\)-token reply budget with bootstrap \(95\%\) confidence intervals (Table 2; the CIs bootstrap over items sharing suffix templates, so treated as exchangeable they may be mildly optimistic, though the margins dwarf plausible item-level wobble), and the picture stratifies. On three models the default discard is real and the recovery is non-overlapping with it: Qwen-3B, Qwen-7B, and Llama-8B honor recognize-intent only \(0.57\)\(0.65\) by default and \(0.85\)\(0.98\) under steering, with disjoint intervals. The other three already honor at a high baseline (Qwen-14B \(0.82\), Mistral \(0.88\), Phi-3.5 \(0.93\)): little discard to recover, so steering nudges within overlapping intervals without degrading. The readout discard is therefore model-specific, present where the model over-produces feedback by default and near-absent where it already honors the intent, whereas the representation is universal. The recovery is not a greedy artifact: Qwen-3B under temperature-\(0.7\) sampling (three seeds) gives default \(0.53\)\(0.60\) and steered \(0.95\)\(0.97\). As a sanity check, evaluate-intent items draw feedback at \(0.63\)\(0.80\) (the withholding on recognize-intent tracks the intent, not an inability to critique; see Limitations). The recovery is front-loaded: on Qwen-3B the steering separation dilutes over longer replies as the model drifts back toward feedback (\(S{=}14 \to 4\) at a \(100\)-token budget; Appendix 19).

7.0.0.1 The same shape as depth (model-specific, not a scaling law).

The split does not track raw size: the discard cases are the smaller, less instruction-mature models (\(\le 8\)B) and the ceiling cases the more mature ones (\(\ge 14\)B), with one telling exception, Mistral-7B sits at ceiling while the larger Qwen-14B only barely clears it. We do not measure capability independently (it would be read off the same behavior it explains), so we state the pattern as model-specific, the larger and instruction-mature models at ceiling in this sample, rather than as a capability law. Pushing further up, Qwen-32B represents the intent (probe \(0.99\) stated, \(0.93\) inferred) yet honors it at baseline (\(0.78\) and \(0.77\), nearer the ceiling models than the discard ones but not far above our soft threshold): the discard, absent from 14B up, does not clearly return at 32B. This echoes the depth result of Section 10, where the intent is represented several layers before the readout uses it. Within a model the readout lags the representation in depth; across models we report an observed stratification, not a scaling law: with six models, one of them (Mistral-7B) already breaking the size ordering, and a soft honoring threshold, we do not have the points to fit a curve. A pre-registered geometric test for the across-model pattern was inconclusive (Appendix 18); the across-model claim rests on the behavioral stratification here and the within-model depth localization.

7.0.0.2 The steer layer is model-specific, and not cherry-picked.

The layer at which the handle works is not a fixed fraction of depth: mid-network for the 7B and Llama-8B (\(0.57\), \(0.59\)), late for the 14B and 3B (\(0.85\), \(0.83\)), and between for Phi and Mistral (\(0.69\)); a fixed-depth heuristic under-recovers (it limped on the 7B until the per-model sweep located layer 16), so the sweep, not a depth rule, finds the handle. Nor is it tuned on the items where the effect is measured: with a nested split by object, the direction is fit and the layer and coefficient selected on a dev half of the objects and recovery measured on the disjoint test half, the recovery stays non-overlapping with the default baseline on both models (Qwen-3B \(0.70 \to 1.00\) at the dev-selected layer 28, coefficient \(1.0\); Llama-8B \(0.53 \to 0.87\) at layer 19, coefficient \(0.5\)). The other four models’ steer layers are selected in-sample (swept on the evaluation items), so we flag those rows of Table 2 as not out-of-sample validated (the effect did hold out of sample on the two we split). A linear map fit between two models’ activation spaces transports Qwen-3B’s intent direction into Llama-8B and steers it, but its held-out reconstruction is poor and the result is only suggestive, so we keep it exploratory and in Appendix 20; the per-model probe, bag-of-words, and steer-layer figures are tabulated in Appendix 24.

Table 2: Discard and recovery at full scale: recognize-intent honoring, \(n{=}60\) items, default vs steeringtoward recognize, bootstrap \(95\%\) CI over items, all six models. The readout discard is model-specific: onthree models (top) the default discard is real and the recovery is non-overlapping; the other three (bottom)already honor recognize-intent at a high baseline, so there is little to recover and steering nudges withinoverlapping intervals without degrading. Qwen-3B holds under sampled decoding (text).
model layer default honoring [95% CI] steered honoring [95% CI] default vs steered
discard present, recovery non-overlapping
Qwen2.5-3B 30 0.65 [0.53, 0.77] 0.98 [0.95, 1.00] non-overlapping
Qwen2.5-7B 16 0.60 [0.47, 0.72] 0.95 [0.88, 1.00] non-overlapping
Llama-3.1-8B 19 0.57 [0.45, 0.68] 0.85 [0.75, 0.93] non-overlapping
ceiling: already honors at baseline, little to recover
Qwen2.5-14B 41 0.82 [0.72, 0.92] 0.90 [0.81, 0.97] overlapping
Mistral-7B 22 0.88 [0.80, 0.95] 0.92 [0.83, 0.98] overlapping
Phi-3.5-mini 22 0.93 [0.87, 0.98] 0.96 [0.91, 1.00] overlapping

8 Generalizing Across Intents↩︎

The clean evidence so far is all recognize-versus-evaluate (stated, Section 3; inferred, Section 4); the one different intent we test, vent versus solve, is lexically marked (bag-of-words \(0.79\); Appendix 22). That markedness is partly constitutive: one cannot ask to vent without venting words, just as the request-matched control could not ask for celebration without celebration words. Recognize-versus-evaluate is special precisely because its intent can ride a prefix over a byte-identical suffix; some intent contrasts admit a surface-matched design and some structurally cannot, which bounds where the strong surface-independence claim can ever be tested. To show the representation is nonetheless not specific to one contrast, we add a third axis designed clean: support (the sender wants a difficulty heard) versus help (wants it solved), set only by an inferred context frame, never stated, over a neutral surface-matched core with lexically diverse frames so leave-one-frame-out defeats a word-counter; behavior is scored by whether the reply offers unsolicited solutions. The design holds: bag-of-words sits at \(0.57\) on every model while the probe decodes the inferred intent at \(0.71\)\(0.86\) (Appendix 23), so the representation generalizes to a genuinely different, fully inferred intent, not only across models. The readout, though, tracks it only weakly (help-intent draws unsolicited solutions more than support, but the inferred intent is solutionized faintly either way) and the represented direction is not a causal handle here: steering toward support raises honoring, but the specificity control of Section 9 fails, random matched-norm directions move solutionizing as much as the learned one (\(S{=}6\) and \(11\) against random maxima \(19\) and \(16\) on Qwen-3B and Phi-3.5; \(p{=}0.31\), \(0.14\)). We scope this axis to the representation and record the steering as a null.

9 Specificity of the Direction↩︎

A steering result invites one objection above all: perhaps the direction is a generic feedback-or-verbosity knob, and any large perturbation would move the behavior. Two dissociations show the effect is specific to the learned intent direction.

9.0.0.1 Only the discriminative direction steers.

As Section 6 reports, the difference-of-means direction at the peak-probe layer fails the sanity gate, while the discriminative (logistic-weight) direction at a later layer passes it. A direction that merely separates the two intent clouds in activation space does not recover the behavior; the direction that discriminates them does. Not any intent-correlated axis works.

9.0.0.2 Norm-matched controls do not reproduce the effect.

At each model’s validated steer layer we measure the behavior separation \(S = \mathrm{feedback}(\text{toward-evaluate}) - \mathrm{feedback}(\text{toward-recognize})\) over \(24\) items. The true direction gives \(S{=}14\) on Qwen-3B (layer 30: \(17/24\) feedback toward evaluate vs \(3/24\) toward recognize) and \(S{=}16\) on Qwen-7B (layer 16: \(19/24\) vs \(3/24\)). Across \(48\) random directions of matched norm, scattered in sign and centered near zero, none reaches the true separation on either model (max \(|S|{=}13\) and \(15\); permutation \(p{=}0.02\), the floor attainable with \(48{+}1\) directions; Figure 2 in the appendix). The effect requires the specific learned direction, not a perturbation of matched norm. Shuffled-label directions agree (max \(S{=}10\) and \(9\) versus \(14\) and \(16\)), though this control has a fat tail on strong axes (an earlier run saw one permutation tie), so the specificity claim rests on the difference-of-means dissociation and the random null (full distributions, Appendix 17).

9.0.0.3 Is the recovery just opener-token biasing?

A sharper deflation: late-layer steering might merely bias the first token toward acknowledgment openers (“Congratulations…”), needing no represented intent. Three tests refute it where steering is clean (Appendix 19): the intent direction is near-orthogonal to the opener-unembedding axis at the steer layer (\(|\cos|\le 0.14\) on all three discard models); on Qwen-7B, steering with the entire first sentence removed swings feedback exactly as much as on the full reply (\(S_{\mathrm{rest}}{=}S{=}14\)), so it reorients the body, not the opening move; and steering the opener direction itself at matched norm fails to reproduce the recovery. On Qwen-3B the separation dilutes over long generations, so its refutation rests on the geometry alone.

9.0.0.4 Does the handle route intent, or just suppress feedback?

A deflationary reading is that the direction is merely the feedback-offering behavior axis, correlated with intent by construction (the labels are defined by whether feedback is wanted), so steering it just suppresses feedback. Two results rule this out. First, applying the same recognize-steer vector to the evaluate-intent items (“be blunt…”), requested feedback is reduced by a model-specific amount, largely surviving on Llama-8B (\(0.68 \to 0.42\), while recognize-honoring recovers \(0.57 \to 0.85\)) but mostly collapsing on the two Qwen models (\(0.77 \to 0.07\), \(0.80 \to 0.12\)). Second, and decisively, the steered direction is not the feedback axis: fitting a direction on reply behavior alone (feedback versus acknowledgment, ignoring the intent labels) at the steer layer, its cosine with the intent-probe direction is only \(0.09\)\(0.13\) across the three models (matched-norm random directions give \(0.01\)\(0.09\)), near-orthogonal. This behavior direction is estimated at the steer layer only, from default replies whose class balance is whatever the discard produced (\(\sim\)​65/35), so it is a noisier direction than the probe, and \(0.13\) against a random max of \(0.09\) warrants separable from, not unrelated to. So the handle routes a represented intent, a feature distinct from the behavior it drives; the requested-feedback collapse is that routed intent overriding the surface request at the recovery dose (model-specifically), not a generic feedback knob.

9.0.0.5 The chain is not specific to one intent contrast.

A second axis, vent versus solve, replicates the whole chain with the same model-specificity (probe \(1.00\)/\(0.97\); a non-overlapping discard-then-recovery \(0.70 \to 0.98\) on Qwen-3B, ceiling on Llama-8B). Being more lexically marked (bag-of-words \(0.79\)), it does more for the behavioral generalization than for surface-independence; full numbers in Appendix 22.

10 Localizing the Discard↩︎

Represented and recoverable place the discard somewhere in the network’s computation; we now read where. On two models (Qwen2.5-3B and Llama-3.1-8B) we sweep every few layers and read, at each, how decodable the intent is (the probe) and how much steering at that layer recovers honoring, the latter on the full \(60\)-item recognize set with bootstrap \(95\%\) CIs; on Qwen-3B we also read where the reply opener commits, via a logit-lens of the last-token residual through the unembedding [17] (Table 3; Figure 3 in the appendix).

10.0.0.1 Represented before routed.

On both models the probe saturates at layers where steering does not yet recover honoring. On Qwen-3B the intent is decodable by mid-network (probe \(1.00\) at layer 24) but steering there leaves honoring at the \(0.65\) baseline (CI overlapping); recovery appears only at layers 28–33, whose CIs clear the baseline, and the acknowledgment-opener mass in the logit-lens is flat until it spikes at layer 28, where the reply is composed. Llama-8B shows the same ordering, the probe saturates (\(1.00\) by layer 10) before steering recovers (from layer 14), though its routing onset is earlier and more diffuse than Qwen’s sharp late window. The general fact is a gap: the sender’s intent is represented before it is routed into the readout, and the causal handle lives past the representation, not at it. Where exactly the routing happens is model-specific, consistent with the model-specific steer layer of Section [sec:general]. A probe saturating before steering becomes effective is common for many features, including acted-on ones; lacking the matched sweep on a non-discarded feature, we read the ordering as consistent with a discard, not diagnostic of one over a generic depth property of steering.

10.0.0.2 Distributed, not a single head.

The window localizes in depth but not to an atomic component. Ablating each late-layer attention head in turn (query-side, one at a time) does not restore honoring: the best single-head ablation lifts it by one item out of sixteen, and an early-layer control head ties for the top. The discard is a distributed late computation, recoverable by the full linear direction but not attributable to any one head, consistent with a routing window rather than a point.

Table 3: Localizing the discard at \(n{=}60\) with bootstrap \(95\%\) CIs, on two models. On both, the probesaturates (intent represented) at layers where steering does not yet recover honoring (CI overlaps baseline);recovery appears only deeper (bold: CI clears baseline). Represented before routed. The routing onsetis model-specific: late and sharp for Qwen-3B (the logit-lens acknowledgment mass spikes at layer 28), earlierand more diffuse for Llama-8B.
layer probe (represents) steer honoring [95% CI]
Qwen2.5-3B baseline \(0.65\) \([0.53, 0.77]\)
0.76 0.60 [0.47, 0.73]
1.00 0.67 [0.55, 0.78]
0.99 1.00 [1.00, 1.00]
0.98 0.88 [0.80, 0.95]
Llama-3.1-8B baseline \(0.57\) \([0.43, 0.70]\)
0.99 0.77 [0.65, 0.87]
1.00 0.83 [0.73, 0.93]
1.00 0.85 [0.75, 0.93]
1.00 0.87 [0.78, 0.95]

11 The Direction Is a Handle With No Prompt, and It Transfers↩︎

If the intent is represented, one might simply tell the model in a system prompt. At \(n{=}60\) with bootstrap CIs (Table 4), an explicit intent prompt largely closes the gap (\(0.93\)) and steering reaches \(0.98\) with overlapping intervals, so we do not claim routing beats prompting; the point is that the represented direction is a handle as good as the instruction with no prompt at all, and the two stack to \(1.00\), which matters wherever the prompt cannot be rewritten (an agent loop, a fixed API). One observation survives at power: a vague nudge backfires. Telling the model to “consider what this person is looking for” lowers honoring from \(0.65\) to \(0.45\); it reads the instruction to attend as license to help. The same direction, fit only on templates, also transfers off them: on hand-written non-templated messages the template-trained probe classifies intent at \(1.00\) and the discard-and-recovery reproduces (\(0.62 \to 1.00\), \(23/23\) coherent); at \(n{=}23\) author-written messages this is a smoke test against a pure-template artifact, not an ecological-validity claim (Appendix 25).

12 Related Work↩︎

12.0.0.1 Communicative intent and pragmatics.

The object we probe is Gricean: what a sender meant by a message, not its surface [1]. Pragmatic competence of this kind has been benchmarked (implicature resolution, [2]) and probed as a story character’s theory-of-mind beliefs [5], where an explicit-versus-implicit gap is documented [18]. Andreas’s conjecture that a predictor comes to represent the agent behind the text [19] is one our base-checkpoint result cashes out. To our knowledge a sender’s goal, distinct from a character’s belief or a user’s attribute, has not been treated as a linear, causal feature.

12.0.0.2 Linear representations and probing.

Linear probes read features from hidden states [7], [8]; a broad literature reads truthfulness, sentiment, and refusal directions from activations [12], to which we add a sender’s communicative intent, separating its representation from whether the readout acts on it.

12.0.0.3 Activation steering.

Adding a direction to the residual stream steers behavior [9][12]. Closest to us, [13] steer explicit instruction-following (format, length, word constraints) with difference-of-means vectors; we find that for the subtler pragmatic intent feature the difference-of-means direction is entangled with the behavior and a discriminative direction at a later layer is required (Section 6). Our contribution is not the steering machinery but the object (a sender’s goal) and the represent-versus-readout decomposition, which mirrors behavioral dissociations in theory of mind [18] and lossy memory [20], shown here mechanistically.

13 Conclusion↩︎

The picture is a readout that lags a representation. A sender’s communicative intent is represented cleanly, from pretraining, across six models and four families, even when it must be inferred and disentangled from the warmth it correlates with; acting on it is the fragile part, and the lag is patterned: decodable layers before it is routed within a model, discarded only by some models across them, an observed stratification that holds through 32B. Where the gap is open a direction closely tied to the representation is a causal handle as good as an explicit instruction, with no prompt, though it acts on feedback-offering behavior and is not always selective for unsolicited feedback. “The model does not get it” is, on this evidence, usually false: it gets it; whether it acts is the model-specific question. The object, not any single number, is the contribution, and the nulls we report against ourselves are the load-bearing part.

Ethics Statement↩︎

This is diagnostic interpretability on open-weight models. The stimuli are synthetic templates and hand-written examples authored by the researcher; no real user data or human subjects were involved, and the behavioral annotation (Section 5) was performed by the author and two independent volunteers who consented to the labeling task. Activation steering is established and we introduce no new capability; the direction we study (honoring a sender’s intent) is benign, but the same handle could suppress useful critique or induce sycophantic withholding, so we are explicit that “honoring” measures routability, not a claim that a model should always withhold feedback (a warm reader may rightly offer a gentle note). We do not recommend intent-steering as a blanket “suppress feedback” intervention; the intended use is understanding where models represent intent and building more faithful readouts on top of it.

Reproducibility Statement↩︎

All code, stimuli, and pre-registration are released (Appendix 26, which lists the exact commands). The surface-matched stimuli and leave-one-phrasing-out protocol are specified in Appendix 14, the probe and its empirical-chance and bag-of-words controls in Appendix 15, and the steering direction, per-layer calibration, and sanity gate in Appendix 16. The core probe-and-recover chain runs CPU-only on Qwen2.5-3B; the six-model sweeps and every steering result run on a single A100 through Modal. Intent labels require no annotation (they are fixed by construction); the behavioral measure and its human validation are detailed in Section 5.

Appendices↩︎

14Stimuli and phrasings
15Probe and controls
16Steering procedure
17Specificity: permutation detail
18Ceiling-model geometry (inconclusive)
19Opener-biasing control (full numbers)
20Cross-model direction transport (exploratory)
21Request-matched construct control (full numbers)
22Second intent axis: vent vs solve
23Third intent axis: support vs help
24Per-model probe, BoW, and steer layer
25Naturalistic transfer (full numbers)
26Reproducibility

14 Stimuli and phrasings↩︎

Sixty objects are crossed with eight lexically diverse phrasings of each intent. Each object’s recognize and evaluate stimulus share an identical suffix of the form “Okay, here it is: the <object>. It works now.” (rotated over three suffix templates), with the intent set only by the prefix. Leave-one-phrasing-out cross-validation groups by phrasing-pair, so train and test never share a wording. Recognize prefixes express sharing without a request for evaluation; evaluate prefixes request critical assessment.

15 Probe and controls↩︎

The probe is \(\ell_2\) logistic regression (\(C{=}1.0\)) on standardized, PCA-reduced (\(\le 40\) component) last-token activations, scored by GroupKFold over the eight phrasing groups. The empirical chance ceiling is the mean and max accuracy under ten label permutations evaluated through the identical pipeline. The bag-of-words baseline is TF-IDF on the full message text under the same grouped cross-validation. A deliberately leaked control (lexically distinct prefixes, no held-out-phrasing protocol) yields \(1.00\) at every layer including layer 6, illustrating the surface artifact the main design controls for.

15.0.0.1 Embedding cross-check of the honoring measure.

The feedback-offer lexicon is corroborated by an independent sentence-embedding classifier that labels each reply by cosine to hand-written feedback versus acknowledgment prototypes (no shared vocabulary with the lexicon). On the recover experiment the two measures agree per-reply (\(13/16\) default, \(16/16\) steered) and reach the same conclusion: default honoring \(12/16\) (lexicon) and \(11/16\) (embedding), rising to \(16/16\) under both after steering. Extending the embedding measure to the full recognize set with the feedback centroid taken from the model’s own evaluate-intent replies, per-model agreement (four models) is \(0.90\)\(0.98\) on steered replies but only \(0.57\)\(0.88\) on default ones: the embedding over-counts honoring on default (Qwen-7B \(0.90\) against the lexicon’s \(0.60\)), keying on register rather than content, its disagreements being warm replies that still offer unsolicited feedback (“That’s fantastic!… here are a few tips”) scored as acknowledgment on their tone. A semantic measure that under-counts the discard, rather than over-counting it, is further evidence the lexicon does not inflate the effect.

16 Steering procedure↩︎

The steering direction at a layer is the unit-normalized logistic weight vector fit on the layer’s raw last-token activations. The vector is added to the residual stream at the predicting (last) position of the target decoder layer during generation, scaled by \(\alpha\) times the layer’s mean activation norm so the coefficient is comparable across layers. Behavior is scored by a coherence guard (minimum length and unique-token ratio) followed by a feedback-offer lexicon (the reply offers feedback/critique/help, or only acknowledges). The sanity gate requires that steering toward evaluate yield more feedback-offering than steering toward recognize at the same coefficient; layers and coefficients failing the gate are not read.

steer toward recognize recognize honored [95% CI]
baseline (\(c{=}0\)) 0.65 [0.53, 0.77]
\(c{=}0.5\) 0.82 [0.72, 0.92]
\(c{=}1.0\) 0.98 [0.95, 1.00]

Causal steering dose-response on Qwen-3B at layer 30 (probe-weight direction, last-position, \(n{=}60\) recognize items, bootstrap \(95\%\) CI): recognize-intent honoring climbs monotonically with dose and the baseline and full-dose intervals are disjoint; coefficient \(1.5\) breaks coherence and is excluded.

17 Specificity: permutation detail↩︎

At each model’s validated steer layer we compare the behavior separation \(S = \mathrm{feedback}(+\text{dir}) - \mathrm{feedback}(-\text{dir})\) of the true (logistic-weight) direction against directions fit on permuted intent labels (twelve) and random directions of matched norm (forty-eight), each scored over the same \(24\)-item subset. Qwen-3B (layer 30): true \(S{=}14\); shuffled \(S \in \{-10,-7,-6,-5,-4,0,2,3,4,7,10,10\}\), all below \(14\); random \(|S| \le 13\), none reaching the true value (permutation \(p{=}0.02\)). Qwen-7B (layer 16): true \(S{=}16\); shuffled \(S \in \{-10,-10,-8,-6,-5,-2,-1,0,5,6,8,9\}\), all below \(16\); random \(|S| \le 15\), none reaching the true value (\(p{=}0.02\)). Neither control reaches the true separation on either model. Shuffled-label permutation is nonetheless a conservative control: on a very strong axis a permutation can partially correlate with true intent (in an earlier run one of twelve shuffled directions tied the true separation on Qwen-7B), so we report the full distribution rather than select a favorable statistic.

Figure 2: The effect requires the learned direction. Distribution of the behavior separation S = \mathrm{feedback}(\text{toward-evaluate}) - \mathrm{feedback}(\text{toward-recognize}) for 48 random directions of matched norm (gray) and 12 shuffled-label directions (blue), against the true intent direction (red line). On both models the true direction sits beyond the entire control distribution; no random or shuffled direction reaches it (permutation p{=}0.02).
Figure 3: Represented before routed. Per layer, how decodable the intent is (blue, probe accuracy) and how much steering at that layer recovers recognize-intent honoring (red, with bootstrap 95\% CIs), against the default baseline (dotted). On both models the probe saturates at layers where steering has not yet lifted honoring above baseline; recovery arrives only deeper (late and sharp for Qwen-3B, earlier and more diffuse for Llama-8B). The gap between the blue and red onsets is the discard. n{=}60.

18 Ceiling-model geometry (inconclusive)↩︎

We asked whether the across-model stratification (Section 7) has a geometric signature: do the ceiling models, which honor recognize-intent by default, align their intent representation with the readout more than the discard models do? We pre-registered, before running, the metric \(M = |\cos(\text{intent direction}, \text{readout direction})|\), where the intent direction is the final-layer logistic recognize-vs-evaluate weight and the readout direction is the fixed acknowledgment-minus-feedback opener unembedding difference, with the decision rule that the three ceiling models’ \(M\) must lie clearly above the three discard models’. The result does not meet the rule. Discard models: \(M \in \{0.12, 0.06, 0.07\}\); ceiling models: \(\{0.14, 0.30, 0.00\}\). Two of three ceiling models (Mistral \(0.30\), Qwen-14B \(0.14\)) exceed the discard range, but Phi-3.5 is the lowest of all six (\(0.00\)), so there is no clean separation; the group means differ (\(0.15\) vs \(0.08\)) but the pre-registered separation criterion fails. A secondary operationalization (the direction separating honored from discarded default replies, computable only where both classes occur) gives \(M \le 0.08\) throughout with no pattern. Per our pre-commitment we report the geometry as inconclusive and make no mechanism claim from it; the across-model claim rests on the behavioral stratification and the depth localization alone.

19 Opener-biasing control (full numbers)↩︎

For the deflationary control of Section 9, at each discard model’s validated steer layer, matched activation norm, generating at a \(100\)-token budget (longer than the \(45\)-token specificity run so the reply has a body), we compare four quantities over the same \(24\)-item subset: \(S_{\mathrm{true}}\), the behavior separation from steering the learned intent direction; \(S_{\mathrm{opener}}\), from steering the acknowledgment-minus-feedback opener-unembedding direction directly; \(S_{\mathrm{rest}}\), \(S_{\mathrm{true}}\) recomputed with each reply’s first sentence removed; and \(|\cos|\) between the intent and opener directions at that layer.

model layer \(S_{\mathrm{true}}\) \(S_{\mathrm{opener}}\) (coherent) \(S_{\mathrm{rest}}\) \(|\cos|\)
Qwen2.5-7B 16 14 \(-16\) (only \(2/24\) coherent) 14 0.001
Llama-3.1-8B 19 7 \(1\) (\(24/24\) coherent, inert) 4 0.034
Qwen2.5-3B 30 4 \(-1\) (incoherent) \(-2\) 0.138

On Qwen-7B the opener direction at matched norm degrades coherence before it honors (\(2/24\) coherent), and \(S_{\mathrm{rest}}\) equals the full \(S_{\mathrm{true}}\), so intent steering reorients the body, not the opener. On Llama-8B the longer budget lifts the true separation clear of the random ceiling (\(S{=}7\) vs \(\max|S|{=}3\)) where the \(45\)-token run left it tied, the opener direction is inert (\(S_{\mathrm{opener}}{=}1\)), and \(S_{\mathrm{rest}}{=}4\) shows a modest body effect. On Qwen-3B the true separation dilutes at this budget (\(S{=}4\), below \(\max|S|{=}12\)): steered toward acknowledgment the small model acknowledges first but drifts back into feedback over a long reply (negative-side feedback rises from \(3/24\) at \(45\) tokens to \(17/24\) at \(100\)), so the body test is inconclusive on 3B and its refutation rests on the norm-independent near-orthogonality (\(|\cos|{=}0.14\)) and the opener direction’s incoherence. At the \(45\)-token budget of Section 9 the 3B separation is strong (\(S_{\mathrm{true}}{=}14\)), so the dilution is a generation-length effect, and the intervention on the smallest model is front-loaded. The load-bearing evidence against opener-biasing is the near-zero cosine on all three models and the body reorientation on Qwen-7B (and, more modestly, Llama-8B). Script: experiments/modal_opener_control.py.

20 Cross-model direction transport (exploratory)↩︎

The models do not share a hidden size, so a direction cannot transfer verbatim. We fit a linear map from Qwen-3B activations to Llama-8B activations on a train split of stimuli and transport Qwen’s intent direction into Llama’s space. The transported direction lands at cosine \(0.85\) to Llama’s own intent direction, and steering Llama with it recovers honoring (\(0.56 \to 1.00\) on held-out items). We keep this out of the main text and flag it as exploratory: the map’s held-out activation \(R^2\) is negative (it aligns the intent direction, it does not reconstruct activation space), and a random direction of matched norm gives partial non-specific recovery (\(0.69\)), so the signal is the direction alignment, not the raw steering number. Read as suggestive that the intent axis is shared up to a linear map, not as an isomorphism claim.

21 Request-matched construct control (full numbers)↩︎

For the construct-validity control of Section 3, both intent classes carry an explicit directive so that request-presence is held constant and only the content of the request differs; recognize prefixes become requests for acknowledgment (“please just celebrate this with me, I really don’t want any notes”; “do me a favor and just take it in, don’t give me feedback”; eight in all), paired with the unchanged evaluate directives and the same surface-matched suffix. Probe accuracy is leave-one-phrasing-out CV at three depths; the bag-of-words baseline uses the same phrasing folds.

model probe (\(0.5\) / \(0.67\) / \(0.83\) depth) peak bag-of-words
Qwen2.5-3B 0.912 / 0.898 / 0.973 0.973 0.688
Qwen2.5-7B 0.929 / 0.891 / 0.955 0.955 0.688
Llama-3.1-8B 0.950 / 0.968 / 0.943 0.968 0.688
Phi-3.5-mini 1.000 / 0.984 / 0.953 1.000 0.688

The probe clears the bag-of-words baseline by roughly \(0.25\)\(0.30\) on every model with request-presence matched. The baseline is higher than the surface-matched set’s \(0.48\) because celebration-requests and critique-requests use different content words; that residual lexical signal is exactly what the surface-matched design (Table 1) controls, and the two designs together rule out both confounds. Script: experiments/modal_request_matched.py.

22 Second intent axis: vent vs solve (full numbers)↩︎

To test that the represents-discard-recover chain is not specific to the recognize-versus-evaluate contrast, we replicate it on a different intent: vent (“I’m not looking for advice, I just need to be heard”) versus solve (“give me concrete steps to fix this”), behavior scored by whether the reply offers a solution, surface-matched stimuli built as before on Qwen2.5-3B and Llama-3.1-8B. Intent decodes at probe \(1.00\) (Qwen-3B) and \(0.97\) (Llama-8B) at the surface-matched token. At \(n{=}60\) with bootstrap CIs, Qwen-3B honors the vent intent only \(0.70\) by default and steering recovers it to \(0.98\) (non-overlapping); Llama-8B already honors it at \(0.88\) baseline (ceiling), so steering only nudges it to \(0.98\). The representation holds on both; the clean discard-then-recovery is on the model that discards. This axis is more lexically marked than the first: a bag-of-words classifier reaches \(0.79\) on the full text (versus \(0.46\)\(0.48\) on recognize-vs-evaluate), so it does more to show the behavior generalizes to a second intent than to re-establish surface-independent representation.

model probe vent honored (default) [95% CI] steered [95% CI] default vs steered
Qwen2.5-3B 1.00 0.70 [0.58, 0.82] 0.98 [0.95, 1.00] non-overlapping
Llama-3.1-8B 0.97 0.88 [0.80, 0.95] 0.98 [0.95, 1.00] ceiling
Table 4: Prompting vs routing on Qwen2.5-3B, \(n{=}60\), bootstrap \(95\%\) CI (full numbers forSection [sec:sec:prompt]). An explicit prompt (\(0.93\)) and steering (\(0.98\)) are interchangeable (overlappingintervals) and stack (\(1.00\)): the represented direction is a handle as good as the instruction, with noprompt. A vague nudge backfires (\(0.65 \to 0.45\)).
condition recognize honored [95% CI]
default (no prompt) 0.65 [0.53, 0.77]
mild nudge 0.45 [0.33, 0.58]
explicit intent prompt 0.93 [0.87, 0.98]
steer (no prompt) 0.98 [0.95, 1.00]
explicit prompt \(+\) steer 1.00 [1.00, 1.00]

23 Third intent axis: support vs help (full numbers)↩︎

The support-versus-help axis (Section 8) is inferred and lexically clean. Per model:

represents solutions offered
2-3(lr)4-5 model probe BoW on support on help
Qwen2.5-3B 0.74 0.57 0.32 0.50
Qwen2.5-7B 0.71 0.57 0.28 0.40
Qwen2.5-14B 0.86 0.57 0.32 0.42
Mistral-7B 0.80 0.57 0.22 0.47
Phi-3.5-mini 0.82 0.57 0.47 0.57
Llama-3.1-8B 0.82 0.57 0.10 0.15

Bag-of-words sits at \(0.57\) against probe \(0.71\)\(0.86\) on all six models, so the representation generalizes to a genuinely different, fully inferred intent. Behaviorally the readout tracks the intent only weakly (help-intent draws unsolicited solutions more than support-intent), and steering the direction fails the specificity control (Section 8), so we claim the representation-generalization here, not a causal handle. \(n{=}40\) per intent.

24 Per-model probe, bag-of-words, and steer layer↩︎

Table 5 tabulates the per-model probe, bag-of-words, and steer-layer figures referenced in Section 7, across all six models and four families.

Table 5: Representation and steer layer across six models, four families (recognize/evaluate axis). Intentdecodes at probe \(1.00\) with bag-of-words at chance (\(\le 0.48\)) everywhere: the representation is not asmall-model or single-family artifact. The steer layer, and its fraction of depth, is model-specific. Thebehavioral discard and recovery, model-specific, are quantified at \(n{=}60\) in Table [tbl:tab:scale]. A seventhmodel, Qwen2.5-32B, is run probe-and-honoring only: it represents the intent (probe \(0.99\) stated, \(0.93\)inferred) and honors it at baseline (\(0.78\)/\(0.77\)), so it has no discard to steer (Section [sec:sec:scale]).
model layers probe / BoW steer (depth)
Qwen2.5-3B 36 1.00 / 0.48 30 (0.83)
Qwen2.5-7B 28 1.00 / 0.46 16 (0.57)
Qwen2.5-14B 48 1.00 / 0.46 41 (0.85)
Mistral-7B 32 1.00 / 0.46 22 (0.69)
Phi-3.5-mini 32 1.00 / 0.46 22 (0.69)
Llama-3.1-8B 32 1.00 / 0.46 19 (0.59)

25 Naturalistic transfer (full numbers)↩︎

A reader may suspect the synthetic, surface-matched stimuli drive the effect. We test whether the probe and the steering direction, both fit only on templates, transfer to hand-written non-templated messages (varied length, register, and topic; e.g.”just got back from my first 5k, didn’t stop once”). On \(23\) such messages the template-trained probe classifies their intent at \(1.00\) (layer 24), and the behavioral chain reproduces: default honoring \(0.62\) (the same discard), steering recovers it to \(1.00\) (\(23/23\) coherent), and genuinely-evaluate messages correctly draw feedback (\(0.08\) honoring). A direction learned on templates governs behavior on real messages; the sample is small (\(n{=}23\), hand-written by the author).

26 Reproducibility↩︎

pip install -e .
# core chain (CPU; downloads Qwen2.5-3B-Instruct)
python experiments/probe_intent.py        # represents (leave-phrasing-out)
python experiments/same_model_discard.py  # discards (same model)
python experiments/steer_probe_sweep.py   # find the causal layer
python experiments/steer_dose.py          # dose-response confirmation
pytest                                     # pipeline + no-leak guards
# extended experiments (GPU, via Modal)
modal run experiments/modal_sweep.py      # generality ladder (6 models)
modal run experiments/modal_scale.py      # discard/recover, n=60
modal run experiments/modal_specificity2.py  # spec controls
modal run experiments/modal_localize2.py  # depth localization
modal run experiments/modal_natural.py    # naturalistic transfer

The core probe is CPU-only and downloads Qwen2.5-3B-Instruct; the extended experiments run on a single A100 through Modal (https://modal.com), and the behavioral elicitation uses API keys. Code, stimuli, and all experiment scripts are available at

https://github.com/collapseindex/recipient-probe.

26.0.0.1 Software.

Experiments use PyTorch [21], HuggingFace Transformers [22], scikit-learn [23] for the probes and steering fits, and sentence-transformers [24] for the independent embedding measure (Section 5).

References↩︎

[1]
H. Paul Grice. Logic and conversation. In Syntax and Semantics 3: Speech Acts, pp. 41–58. Academic Press, 1975.
[2]
Laura Ruis, Akbir Khan, Stella Biderman, Sara Hooker, Tim Rocktäschel, and Edward Grefenstette. The goldilocks of pragmatic understanding: Fine-tuning strategy matters for implicature resolution by LLMs. In Advances in Neural Information Processing Systems (NeurIPS), volume 36, pp. 20827–20905, 2023.
[3]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems (NeurIPS), 2022.
[4]
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, et al. Towards understanding sycophancy in language models. In International Conference on Learning Representations (ICLR), 2024.
[5]
Matteo Bortoletto, Constantin Ruhdorfer, Lei Shi, and Andreas Bulling. Benchmarking mental state representations in language models. In ICML 2024 Workshop on Mechanistic Interpretability, 2024.
[6]
Qwen Team. Qwen2.5 technical report, 2024. URL https://arxiv.org/abs/2412.15115.
[7]
Guillaume Alain and Yoshua Bengio. Understanding intermediate layers using linear classifier probes. In International Conference on Learning Representations (ICLR), Workshop Track, 2017.
[8]
Yonatan Belinkov. Probing classifiers: Promises, shortcomings, and advances. Computational Linguistics, 48 (1): 207–219, 2022.
[9]
Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J. Vazquez, Ulisse Mini, and Monte MacDiarmid. Activation addition: Steering language models without optimization. arXiv preprint arXiv:2308.10248, 2023. URL https://arxiv.org/abs/2308.10248.
[10]
Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. Inference-time intervention: Eliciting truthful answers from a language model. In Advances in Neural Information Processing Systems (NeurIPS), 2023.
[11]
Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner. Steering llama 2 via contrastive activation addition. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024.
[12]
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al. Representation engineering: A top-down approach to ai transparency. arXiv preprint arXiv:2310.01405, 2023. URL https://arxiv.org/abs/2310.01405.
[13]
Alessandro Stolfo, Vidhisha Balachandran, Safoora Yousefi, Eric Horvitz, and Besmira Nushi. Improving instruction-following in language models through activation steering. In International Conference on Learning Representations (ICLR), 2025.
[14]
Albert Q. Jiang et al. Mistral 7b, 2023. URL https://arxiv.org/abs/2310.06825.
[15]
Marah Abdin et al. Phi-3 technical report: A highly capable language model locally on your phone, 2024. URL https://arxiv.org/abs/2404.14219.
[16]
Abhimanyu Dubey et al. The llama 3 herd of models, 2024. URL https://arxiv.org/abs/2407.21783.
[17]
Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt. Eliciting latent predictions from transformers with the tuned lens. arXiv preprint arXiv:2303.08112, 2023. URL https://arxiv.org/abs/2303.08112.
[18]
Yuling Gu, Oyvind Tafjord, Hyunwoo Kim, Jared Moore, Ronan Le Bras, Peter Clark, and Yejin Choi. : Exposing the gap between explicit ToM inference and implicit ToM application in LLMs, 2024. URL https://arxiv.org/abs/2410.13648.
[19]
Jacob Andreas. Language models as agent models. In Findings of the Association for Computational Linguistics: EMNLP 2022, 2022. . URL https://aclanthology.org/2022.findings-emnlp.423.
[20]
Alex Kwon. Reclaim evaluation: A lossy memory is worse than an empty one, 2026. URL https://arxiv.org/abs/2606.25449.
[21]
Adam Paszke et al. Pytorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems (NeurIPS), 2019.
[22]
Thomas Wolf et al. Transformers: State-of-the-art natural language processing. In Proceedings of EMNLP: System Demonstrations, pp. 38–45, 2020.
[23]
Fabian Pedregosa et al. Scikit-learn: Machine learning in python. Journal of Machine Learning Research, 12: 2825–2830, 2011.
[24]
Nils Reimers and Iryna Gurevych. Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of EMNLP-IJCNLP, 2019.